
This AI glossary defines the 50 terms business leaders hear most often in 2026, from agentic AI and RAG to tokens, evals and the Model Context Protocol. Each entry is a plain-English definition of one to three sentences, so you can follow a vendor pitch, scope a project or read a technical proposal with confidence.
Terms are listed alphabetically. Where a term connects to work Aaga does, we link to a deeper guide or service page.
A
Agentic AI
Agentic AI is an approach where AI systems pursue a goal by planning steps, using tools, checking results and adjusting, rather than answering a single prompt. It sits between fixed automation and a human worker: it can handle variation, but it needs clear boundaries. Read our full guide: What is agentic AI?
AI agent
An AI agent is software that uses a language model to decide what to do next and calls tools (APIs, databases, other systems) to complete a task, such as resolving a support ticket or updating a CRM. Agents differ from chatbots because they take actions, not just give answers. Aaga builds these through its AI agent development practice.
AI bias
AI bias is when a model produces systematically unfair or skewed results for certain groups, usually because of patterns in its training data or how a problem was framed. It is managed through representative data, testing across user groups and human review of high-impact decisions.
AI governance
AI governance is the set of policies, roles and controls that decide how an organization selects, approves, monitors and retires AI systems. It covers data use, risk assessment, accountability, documentation and compliance with laws such as the EU AI Act or India's Digital Personal Data Protection Act.
C
Chatbot
A chatbot is software that holds a text conversation with users, typically to answer questions or guide them through a simple flow. Modern chatbots use LLMs and a knowledge base; older ones followed decision trees. See AI agents vs chatbots for where one ends and the other begins.
Chunking
Chunking is splitting long documents into smaller passages before indexing them for search or RAG. Good chunking keeps related ideas together (for example, by section or heading) so the retriever finds complete, useful context.
Computer vision
Computer vision is the field of AI that interprets images and video, for tasks like defect detection, document reading, people counting or robot guidance. It powers quality inspection lines and AI for robotics.
Context window
The context window is the maximum amount of text (measured in tokens) a model can consider at once, including your instructions, the conversation, retrieved documents and its own reply. Larger windows let a model read more, but they cost more per request and do not guarantee the model uses every detail well.
Copilot
A copilot is an AI assistant embedded in a tool people already use, such as an email client, code editor or CRM, that suggests drafts, answers and next steps while a human stays in control. The human approves the output, which is the main difference from an autonomous agent.
D
Data labeling
Data labeling (or annotation) is adding the correct answers to raw data, such as tagging objects in images or marking the intent of a support message, so a model can learn from it or be evaluated against it. Label quality often matters more than data volume. Aaga provides dataset collection and labeling for this.
Deep learning
Deep learning is a type of machine learning that uses neural networks with many layers to learn complex patterns from large amounts of data. It underpins modern speech recognition, computer vision and language models.
E
Embeddings
Embeddings are lists of numbers (vectors) that represent the meaning of text, images or audio, so that similar items end up close together. They make semantic search, recommendations, clustering and RAG possible.
Evals
Evals (evaluations) are repeatable tests that measure how well an AI system performs on the tasks you care about, using a fixed set of inputs and expected outcomes or scoring rules. They are how teams catch regressions when they change a prompt, model or data source. Evals belong in the same pipeline as other automated tests; see our quality engineering service.
F
Few-shot prompting
Few-shot prompting is including a handful of worked examples in a prompt so the model copies the pattern, format or tone. Zero-shot prompting is the same request with no examples.
Fine-tuning
Fine-tuning is further training an existing model on your own examples so it adopts a specific style, format or specialized behavior. It is useful for consistent outputs and narrow tasks, but it is not the best way to give a model up-to-date facts; RAG usually handles that better.
Foundation model
A foundation model is a large model trained on broad data that can be adapted to many tasks, such as GPT, Claude, Gemini or Llama. Most business AI products are built on top of a foundation model rather than trained from scratch.
Function calling
Function calling (also called tool use) is a model's ability to return a structured request to run a specific function with specific arguments, for example get_order_status(order_id). The application runs the function and passes the result back. It is the mechanism that lets AI agents take actions.
G
Generative AI
Generative AI is AI that creates new content such as text, images, audio, video or code in response to a prompt. Business uses include drafting, summarization, document processing and assistants; see our generative AI development services.
Grounding
Grounding is tying a model's answer to a trusted source, such as retrieved documents, a database record or a tool result, instead of relying on what the model memorized during training. Grounded answers can usually cite where the information came from.
Guardrails
Guardrails are the checks around an AI system that keep it within safe and allowed behavior: input filters, output validation, topic restrictions, permission limits on tools and approval steps for risky actions. Good guardrails are layered, so no single check is the only line of defense.
H
Hallucination
A hallucination is a confident but false or unsupported output from a language model, such as an invented policy, figure or citation. It is reduced (not eliminated) through grounding, clear instructions, constrained outputs and evals.
Human-in-the-loop
Human-in-the-loop means a person reviews, approves or corrects AI output at defined points, especially before high-impact actions like refunds, contract changes or medical advice. It is a design choice, not an admission of failure.
I
Inference
Inference is running a trained model to get an output, for example generating a reply or classifying a document. Inference cost and speed, not training, usually dominate the running cost of a business AI system.
L
Large language model (LLM)
A large language model is a neural network trained on very large amounts of text to predict the next token, which lets it understand and generate language, follow instructions and write code. LLMs are the reasoning engine inside most chatbots, copilots and AI agents.
LoRA
LoRA (Low-Rank Adaptation) is a parameter-efficient fine-tuning method that trains small add-on weight matrices instead of the whole model. It makes fine-tuning cheaper and lets one base model serve several specialized adapters.
M
Machine learning
Machine learning is the practice of building systems that learn patterns from data to make predictions or decisions, such as forecasting demand, scoring leads or detecting fraud. It covers classic models as well as deep learning. See machine learning services.
MLOps
MLOps is the set of practices and tooling for deploying, monitoring, versioning and retraining machine learning models in production, similar to DevOps for software. LLMOps applies the same ideas to prompts, LLM versions, evals and cost tracking.
Model Context Protocol (MCP)
The Model Context Protocol is an open standard, introduced by Anthropic in November 2024, for connecting AI applications to external tools and data through a common interface. A system exposes its capabilities once as an MCP server, and any compatible AI client can use them. Read MCP explained for business teams.
Multi-agent system
A multi-agent system is a setup where several AI agents with different roles (for example, a researcher, a writer and a reviewer) work together on a task, coordinated by an orchestrator. It can improve quality on complex work, but adds cost and makes debugging harder.
Multimodal AI
Multimodal AI handles more than one type of input or output, such as text, images, audio and video, in the same model. A multimodal model can, for example, read a photographed invoice and answer questions about it.
N
Natural language processing (NLP)
Natural language processing is the field of AI concerned with understanding and generating human language, including classification, entity extraction, translation and summarization. LLMs are now the dominant NLP technology.
O
Open-weight model
An open-weight model is one whose trained parameters are published, so you can download it and run it on your own infrastructure, for example Llama, Mistral, Qwen or Gemma. It offers more control over data and cost, but you take on hosting, scaling and safety work. Licenses vary, so check them before commercial use.
Orchestration
Orchestration is the logic that coordinates the steps of an AI workflow: which model to call, which tools to use, in what order, with what retries and approvals. It is often where most of the engineering in an agent lives.
P
Prompt engineering
Prompt engineering is designing the instructions, examples and context given to a model so it produces reliable, useful output. In production it is closer to software engineering than wordsmithing: prompts are versioned and tested with evals.
Prompt injection
Prompt injection is an attack where hidden or malicious instructions in user input, a web page, an email or a document try to override an AI system's rules, for example to leak data or trigger an unwanted action. It is a top security risk for agents and is mitigated through least-privilege tools, input isolation and approvals. Our application security team includes LLM guardrails in reviews.
Q
Quantization
Quantization is storing a model's numbers at lower precision (for example 8-bit or 4-bit instead of 16-bit) so it uses less memory and runs faster, with a small accuracy trade-off. It is common when running open-weight models on limited hardware.
R
Reasoning model
A reasoning model is an LLM trained to work through a problem in intermediate steps before answering, which improves results on math, coding, planning and multi-step analysis. It is usually slower and more expensive per answer, so it suits hard tasks rather than simple ones.
Reinforcement learning from human feedback (RLHF)
RLHF is a training method where people rate or rank model outputs, and those preferences are used to adjust the model toward more helpful and safer behavior. It is one of the main reasons modern assistants follow instructions well.
Retrieval-augmented generation (RAG)
RAG is a technique where the system first retrieves relevant passages from your own documents or data, then gives them to the language model to generate an answer. It keeps answers current, specific to your business and citable. See RAG explained for business.
S
Semantic search
Semantic search finds results by meaning rather than exact keywords, usually by comparing embeddings. A query for "cancel my plan" can match a document titled "How to end your subscription." Many systems combine it with keyword search (hybrid search).
Small language model (SLM)
A small language model is a compact language model, typically with a few billion parameters or fewer, that is cheaper and faster to run and can operate on a laptop, phone or edge device. SLMs work well for focused tasks like classification, extraction or routing.
Speech-to-text (STT)
Speech-to-text, also called automatic speech recognition (ASR), converts spoken audio into written text. It is the first stage of most voice AI systems, and its accuracy on accents, names and numbers largely decides call quality. See voice AI in Indian languages.
Synthetic data
Synthetic data is artificially generated data that mimics the statistical patterns of real data, used to train or test models when real data is scarce, sensitive or expensive to label. It must be validated against real-world samples. Aaga offers synthetic data creation.
System prompt
The system prompt is the standing set of instructions given to a model before the user's input: its role, rules, tone, tools and limits. It shapes every reply, so it is treated as a versioned, tested asset.
T
Temperature
Temperature is a setting that controls how random a model's output is. Low values give consistent, predictable answers (good for extraction and support); higher values give more varied output (useful for brainstorming).
Text-to-speech (TTS)
Text-to-speech converts written text into spoken audio. Modern neural TTS sounds close to a human voice and can stream audio as it is generated, which keeps voice conversations responsive.
Tokens
Tokens are the small chunks of text a model reads and writes, often a word or part of a word. In English a token averages roughly three-quarters of a word, while many other languages and scripts use more tokens per word. Model pricing, speed and context limits are all measured in tokens.
Transformer
The transformer is the neural network architecture behind modern LLMs, introduced by Google researchers in the 2017 paper "Attention Is All You Need." Its attention mechanism lets the model weigh how every token relates to every other token in the context.
V
Vector database
A vector database stores embeddings and finds the most similar ones quickly, which is the retrieval engine behind semantic search and RAG. Examples include pgvector (a PostgreSQL extension), Pinecone, Weaviate, Qdrant and Milvus.
Voice AI
Voice AI is technology that lets software hold spoken conversations, combining speech-to-text, a language model and text-to-speech (or a single speech-to-speech model). Businesses use it for AI receptionists, appointment booking, reminders and lead qualification; see AI voice agents and what a voice AI agent is.
How to Use This Glossary
You don't need to master every term to make good AI decisions. Most business conversations come down to four questions, and this glossary maps onto them:
- What should the system do? Chatbot, copilot, AI agent or agentic workflow.
- What does it know? RAG, grounding, embeddings, knowledge base, fine-tuning.
- How is it kept safe and accurate? Guardrails, evals, human-in-the-loop, prompt injection defenses.
- What does it cost to run? Tokens, context window, inference, model size.
If a proposal can't answer those four questions in plain language, ask for a clearer one.
Talk to Engineers, Not Jargon
Aaga is an AI-native engineering team: the people who explain these terms to you are the same people who build the system. If you're weighing where AI fits in your business, book a free consultation and we'll map your use case, data and risks in plain English, then propose a scoped pilot.

