Business AI terminology

Terms that matter when AI enters a workflow.

Plain-language definitions, with the business decision each term changes. Start here when a model, vendor, or project brief uses unfamiliar language.

RAG: retrieval-augmented generation

RAG is a pattern in which a system finds relevant, authorized information before the model writes an answer. Instead of asking the model to remember your company policy, the system retrieves the policy, passes the relevant excerpts to the model, and asks it to answer from that evidence.

A grounded answer flow
1. QuestionA user asks about a case.
2. RetrieveSearch approved records and documents.
3. EvidenceFilter for relevance, freshness, and permissions.
4. GenerateModel drafts an answer using those sources.
5. ReviewUser checks citations, edits, or escalates.

Business use: policy support, service responses, research briefs, contract review, and technical troubleshooting. Common misconception: RAG does not make an answer true. It makes evidence visible and gives you a system to improve when the answer is wrong.

Retrieval, embeddings, and vector search

Retrieval means finding useful evidence before an answer is produced. An embedding is a numerical representation of meaning; similar pieces of text sit near one another in this representation. Vector search uses embeddings to find semantically related records, even when they do not share the same words. In a business system, keyword filters, metadata, permissions, and freshness rules should work alongside vector search.

Context, prompts, and context windows

Context is the information available when the model makes an output: instructions, user role, customer record, policy excerpts, previous steps, and tool results. A prompt is the instruction and content sent to the model. The context window is the maximum amount of information a model can consider in one request.

Business implication: a larger context window is not automatically better. Stale, duplicated, irrelevant, or unauthorized information can make a model less reliable. Good context is selected for one decision, traceable to a source, and permission-aware.

Grounding and hallucination

Grounding means anchoring an output in specific, inspectable evidence. A hallucination is a plausible-sounding but unsupported or incorrect output. Treat it as an operational failure mode: define when the system must cite, abstain, ask a question, or route work to a person.

Copilots, workflows, agents, and orchestration

CopilotDrafts, summarizes, or recommends; a person owns the next decision.
Workflow automationFollows a known sequence of rules and steps, such as routing an approved form.
AgentPursues a goal across multiple steps, selecting among permitted tools and checking results.
OrchestrationCoordinates models, tools, data sources, approvals, retries, and logs into one dependable system.

Autonomy is how much the system can decide or do without a person’s approval. Start at the lowest level that creates value: summarize, recommend, draft, then act. Higher autonomy needs tighter permissions, stronger evaluation, and a reversible failure path.

Tools, APIs, and function calling

A tool lets an AI system do something outside the model: search a knowledge base, look up a CRM record, calculate a price, create a ticket, or call an internal service. An API is a defined way for software systems to exchange data or request an action. Function calling is a structured format through which a model asks an application to invoke a tool with named inputs.

The model should not receive unrestricted access merely because it can call a tool. Define the fields it may read or change, the user role required, a validation rule, an approval point for consequential actions, and an audit log.

Evaluation, benchmarks, and observability

An evaluation (or eval) is a repeatable test of whether an AI system produces the desired outcome. A benchmark is a shared standardized test; it can help form a hypothesis but cannot prove business fit. A private evaluation set contains representative cases from your own workflow: normal work, edge cases, ambiguous requests, outdated records, and permission failures.

An operational learning loop
BaselineMeasure the current process.
Test casesScore realistic work and failures.
MonitorTrack quality, cost, latency, and overrides.
ReviewInspect errors with process owners.
ImproveChange context, model, controls, or workflow.

Observability means being able to see what happened in the system: retrieved sources, model output, tool calls, latency, cost, errors, approvals, and outcomes. Without it, a team cannot diagnose a failure or learn whether the pilot is creating value.

Guardrails, human-in-the-loop, and governance

Guardrails are technical and process boundaries that prevent unwanted behavior: permission checks, data filters, output validation, rate limits, approval gates, and rules about prohibited actions. Human-in-the-loop means a person reviews or approves an output or action at a defined point. It is most useful when the review is specific, informed, and connected to an accountable owner.

Governance is the operating system around AI: named owners, intended-use boundaries, data access, evaluation, monitoring, incident response, vendor review, and decisions about when to scale or stop. Use the AI Implementation Canvas ↗ to turn these terms into a concrete pilot plan.