AI GLOSSARY

AI terms, in plain English.

Clear, authoritative definitions of the AI concepts that matter for building organisational capability. Answer-first, jargon-checked, and written in British English.

Agentic AI

Agentic AI is artificial intelligence that can plan and carry out multi-step tasks towards a goal — choosing actions, using tools and reacting to results — rather than only responding to a single prompt.

A standard chatbot answers one question at a time. An agentic system is given an objective and works towards it: breaking it into steps, calling tools or APIs, checking its own progress, and adjusting. This makes it more capable but also raises the stakes for oversight, because it acts rather than just suggests.

AI capability

AI capability is an organisation's ability to use AI well, safely and at scale — the practical skill to turn AI access into changed outcomes, as distinct from simply having AI tools available.

Capability is what sits between owning AI tools and getting value from them. It combines individual skill, shared practice, sanctioned tooling and governance. The Forge AI Capability Maturity Model measures it across five stages.

AI governance

AI governance is the set of policies, oversight and controls that ensure an organisation's use of AI is safe, accountable, compliant and aligned with its values.

Governance covers what AI may be used for, who is accountable for AI-assisted decisions, how data is handled, what competence people need, and how use is evidenced. The Forge AI Governance Scorecard breaks this into five assessable dimensions.

Alignment

Alignment is the work of making an AI system behave in line with human intentions and values — helpful, honest and safe. As models become more capable, keeping their goals and outputs aligned with ours becomes more important.

Artificial general intelligence (AGI)

Artificial general intelligence refers to a hypothetical AI that can match or exceed human ability across essentially any intellectual task, rather than being good at one narrow thing. Today's systems are not AGI, and there is no agreed date for when — or whether — it arrives.

Artificial intelligence (AI)

Artificial intelligence is the field of building computer systems that perform tasks we normally associate with human intelligence — understanding language, recognising images, making predictions or decisions — by learning patterns from data rather than following fixed rules.

Benchmark

A benchmark is a standard test used to compare models on a task such as reasoning, coding or maths. Benchmarks are useful for rough comparison, but a high score does not always translate into being better for your particular job.

Chain-of-thought

Chain-of-thought is prompting or training a model to work through a problem step by step before giving its answer. Reasoning in the open tends to improve accuracy on maths, logic and multi-step tasks.

Context window

The context window is how much text a model can consider at once — your prompt plus its answer — measured in tokens. Once a conversation or document exceeds it, the earliest content is dropped and effectively forgotten.

Deep learning

Deep learning is machine learning that uses neural networks with many layers. The extra depth lets a model learn increasingly abstract patterns — from edges, to shapes, to whole objects — and is what powers most modern language and image AI.

Diffusion model

A diffusion model generates images (or other media) by starting from random noise and repeatedly refining it into a coherent result. It is the approach behind many popular image generators.

Distillation

Distillation trains a smaller "student" model to imitate a larger "teacher" model, capturing much of its ability at a fraction of the size and cost. It is a route to fast, cheap models that stay surprisingly capable.

Embedding

An embedding is a list of numbers that represents the meaning of a piece of text (or an image) so that similar things sit close together. Embeddings are what let AI systems search and compare by meaning rather than by exact words.

Few-shot learning

Few-shot means including a small number of worked examples in the prompt to show the model the pattern you want. It often improves accuracy and format without any retraining of the model.

Fine-tuning

Fine-tuning is the process of further training an existing AI model on a focused set of examples so it adapts to a specific style, domain or task.

Rather than building a model from scratch, fine-tuning takes a capable base model and adjusts it with your own examples — your tone, your domain language, your task patterns. It is powerful for consistency and specialisation, but it is not always the right tool: for answering from a body of documents, retrieval-augmented generation is often simpler and easier to maintain.

Guardrails

Guardrails are the rules and checks placed around an AI system to keep its behaviour within safe, acceptable bounds — filtering harmful requests, constraining what it can access, and catching bad outputs before they reach a user.

Hallucination

A hallucination is when a model states something false or invented as if it were fact. Because a language model predicts plausible text rather than checking a source, confident-sounding errors are a built-in risk — which is why grounding and verification matter.

Inference

Inference is using a trained model to produce an answer — the moment you send a prompt and get a response. Unlike training, it happens every time the model is used, so its speed and cost matter for real-world deployment.

Large language model (LLM)

A large language model is an AI trained on vast amounts of text to predict likely continuations of language. That simple objective, at scale, produces systems that can answer questions, summarise, translate and write — the technology behind tools like ChatGPT and Claude.

Machine learning (ML)

Machine learning is the branch of AI where a system improves at a task by learning from examples instead of being explicitly programmed. You supply data and desired outcomes, and the system works out the patterns that connect them.

Model Context Protocol (MCP)

The Model Context Protocol (MCP) is an open standard for connecting AI models to external tools and data sources through a consistent interface, so a model can access systems without bespoke integration for each one.

MCP standardises how an AI application exposes tools, files and data to a model. Instead of writing a custom connector for every system, a service implements the protocol once and any MCP-aware model can use it. It is becoming a common way to give AI assistants safe, structured access to real systems.

Multimodal

A multimodal model can work with more than one kind of input or output — for example text and images together — rather than text alone. This lets it describe a photo, read a chart, or answer questions about a document's layout.

Neural network

A neural network is a model made of many simple connected units, loosely inspired by the brain. Each connection has a weight that is adjusted during training, and together they let the network map inputs to outputs for tasks like recognition or prediction.

Open-weight model

An open-weight model is one whose trained weights are published, so anyone can download, run, inspect or fine-tune it on their own hardware. This supports privacy and control, and is central to keeping AI capability portable rather than locked to one vendor.

Overfitting

Overfitting is when a model learns its training examples too closely — including their noise and quirks — and then performs poorly on new data. It is the classic sign a model has memorised rather than genuinely generalised.

Parameters

Parameters are the internal values a model adjusts during training to capture what it has learned. A model's size is often quoted by its parameter count (billions of them); more parameters can mean more capability but also more cost to run.

Prompt

A prompt is the instruction or question you give an AI model. Because a model responds to what it is asked, the wording, context and examples in a prompt strongly shape the quality of the answer.

Prompt engineering

Prompt engineering is the practice of designing prompts to get reliable, useful results from a model — being specific, giving context and examples, and setting the format you want. It is often the cheapest way to improve output before reaching for fine-tuning.

Quantisation

Quantisation shrinks a model by storing its weights at lower numerical precision, so it uses less memory and runs faster — often with only a small loss of quality. It is a common way to run capable models on modest hardware.

Reinforcement learning from human feedback (RLHF)

RLHF is a training method where people rate a model's responses and those ratings are used to steer it towards more helpful, honest and harmless behaviour. It is a big part of why modern assistants feel usable rather than merely fluent.

Retrieval-Augmented Generation (RAG)

Retrieval-augmented generation (RAG) is a technique that lets an AI model answer using specific documents or data retrieved at the moment of the query, rather than relying only on what it learned during training.

RAG connects a model to a searchable knowledge source — your documents, a database, a wiki. When asked a question, the system retrieves the most relevant material and gives it to the model as context, so the answer is grounded in your data and easier to keep current. It is often a better first step than fine-tuning when the goal is to answer from a body of known information.

Temperature

Temperature is a setting that controls how varied a model's output is. Low temperature makes responses focused and repeatable; higher temperature makes them more diverse and creative, at the cost of consistency.

Token

A token is the small chunk of text a language model actually reads and writes — often a word or part of a word. Models measure input and output length in tokens, and pricing and context limits are usually counted in them.

Tokenisation

Tokenisation is the step of breaking text into tokens before a model can process it. How text is split affects how much fits in the context window and how the model interprets unusual words, code or other languages.

Training

Training is the process of adjusting a model's weights by showing it many examples and correcting its mistakes, until it reliably produces good outputs. It is computationally expensive and happens before the model is deployed.

Transformer

A transformer is the neural-network design behind most modern language models. Its key idea, "attention", lets the model weigh how much each word relates to every other word, so it can handle long-range context far better than earlier approaches.

Vector database

A vector database stores embeddings and finds the closest matches to a query quickly. It is the retrieval engine behind most RAG systems, letting an application pull the most relevant passages to feed a model as context.

Weights

Weights are the specific learned numbers on a neural network's connections — the concrete form a model's parameters take after training. Releasing a model's weights is what lets others run or fine-tune it themselves.

Zero-shot learning

Zero-shot means asking a model to do a task with no worked examples in the prompt — relying entirely on what it already learned in training. It is fast to try but less reliable than showing the model a few examples.