Generative AI terms, explained so anyone can understand

"Wait, what's an agent?"

You keep hearing these AI terms. Here's what they actually mean.

LLM Agent MCP RAG Tokens Reasoning Skill HITL
See it all in action
Showing connections, click here or press Esc to clear
No terms matched. Try a different keyword.
Models

LLMLarge Language Model

A massive AI model trained on huge amounts of text. It learns patterns in language so it can generate human-like text, answer questions, write code, and more. Well-known LLMs include Claude (Anthropic), Gemini (Google), GPT (OpenAI), and Llama (Meta).

Example: When you ask ChatGPT a question, an LLM reads your input as tokens, uses reasoning to think through a response, and generates an answer word by word.
Models

SLMSmall Language Model

A compact language model designed to be fast, efficient, and run on smaller hardware, even your phone. SLMs trade some general knowledge for speed and privacy, ideal when you need a model to work locally or on a specific task.

Example: Microsoft's Phi-4, Google's Gemma, and Meta's Llama 4 Scout are SLMs that run directly on a laptop or phone, using far fewer tokens per second than a full LLM.
Models

GPTGenerative Pre-trained Transformer

The architecture created by OpenAI that kicked off the generative AI revolution. "Generative" means it creates new text, "Pre-trained" means it learned from huge datasets beforehand, and "Transformer" is the neural network design that makes it work. The Transformer architecture is also used by Gemini, Claude, and Llama.

Example: ChatGPT runs on GPT models like GPT-5.4. It's a specific LLM from OpenAI, and its newer versions are multimodal. They understand images and audio too, not just text.
Models

Token

The basic unit of text that LLMs work with. Usually a word or word-piece. "explaining" becomes "explain" + "ing". There are two types: prompt tokens (your input: the question, system instructions, and any context) and completion tokens (the model's output). Pricing, speed, and context limits are all measured in tokens.

Example: You send a 200-token question to Claude. It generates a 500-token reply. That's 200 prompt (input) tokens + 500 completion (output) tokens = 700 total. Input tokens are cheaper, typically $3/M vs $15/M for output. Gemini 3.1 Pro and Claude Opus 4.6 both support 1M tokens of context.
Models

Reasoning

An LLM's ability to think step-by-step before answering, rather than just pattern-matching. Some models (GPT-5.4 Thinking, Claude extended thinking) have reasoning built in. But agent frameworks can also add reasoning to any model through structured chain-of-thought prompting, no special model needed.

Example: GPT-5.4 Thinking, Gemini 3.1 Pro, and Claude's "extended thinking" have built-in reasoning. But even non-reasoning models can reason. Agent frameworks like LangChain and Agno add chain-of-thought prompting to any LLM, no special model required.
Models

Multimodal

A model that understands and generates more than just text: images, audio, video, and code too. Instead of being limited to one "mode," it handles several at once, like a person who can read, listen, and see.

Example: Gemini 3.1 Pro, Claude Sonnet 4.6, and GPT-5.4 are all multimodal. Send a photo of a menu, an audio clip, or a screenshot. Each input type is converted into tokens the model can process.
Models

Multilingual

A model that understands and responds in multiple human languages. Modern LLMs like Claude, Gemini, and GPT are trained on text from dozens of languages, so they can translate, answer, and converse without needing separate models for each language.

Example: Ask ChatGPT a question in Japanese, get a Japanese answer, then switch to Spanish mid-conversation. Claude, Gemini, and GPT all handle this because they learned multilingual patterns from training data across 100+ languages.
Infrastructure

Agent

A system built around an LLM that can plan and take action, not just answer questions. The LLM is the brain; the agent is the whole package: it calls tools, browses the web, writes files, and loops until a task is done. Autonomy is a spectrum: some agents run end-to-end on their own, others pause at key moments for human approval.

Example: Cursor, Claude Code, Devin, and GitHub Copilot are coding agents. For building your own, platforms like Elastic Agent Builder, LangChain, and Agno let you create custom agents that use tools, connect to services via MCP, and use reasoning to plan each step. Some run on autopilot; others pause for human confirmation before high-stakes actions.
Infrastructure

Tool

Any external capability an LLM or agent can call. On its own, an LLM can only generate text. Give it tools like web search, a code runner, file access, or a database, and it can interact with the real world. This is called "function calling" or "tool use."

Example: Claude, Gemini, and GPT all support tool use natively. When an agent needs today's weather, it calls a weather API tool, gets real data back, and the LLM weaves it into a natural response.
Infrastructure

MCPModel Context Protocol

An open standard (created by Anthropic) that gives LLMs and agents a universal way to connect to tools, apps, and data sources. Think USB-C for AI: one protocol that works everywhere. MCP servers expose tools for the model to call, resources for it to read, and prompts for it to use.

Example: Cursor, Claude Desktop, ChatGPT, Windsurf, and VS Code all support MCP. An agent in any of them can connect to GitHub, Slack, Postgres, or any service that runs an MCP server, all through the same protocol. See the full spec.
Infrastructure

A2AAgent-to-Agent Protocol

An open protocol (created by Google, now under the Linux Foundation) that lets AI agents communicate with each other, even across different companies. Where MCP connects an agent to tools, A2A connects agents to other agents.

Example: Backed by Salesforce, SAP, and others. A travel agent uses A2A to ask a flights agent for tickets, a hotels agent for rooms, and a calendar agent for your schedule, all coordinating automatically. See the full spec.
Infrastructure

Skill

A packaged, reusable set of instructions that teaches an agent how to do something specific. Instead of re-explaining a task every time, you give the agent a skill file and it knows how to perform it whenever needed.

Example: Cursor and Claude Code use SKILL.md files, ChatGPT has custom instructions, and Microsoft's Copilot Studio has "Topics." Create a "deploy to production" skill with steps, tools, and safety checks. Now just say "deploy" and the agent knows what to do.
Infrastructure

Plugin

A pre-built extension that adds capabilities to an LLM application. Anthropic's Claude uses MCP-based integrations (connecting to GitHub, Slack, databases, etc.) while OpenAI offers custom GPTs with "Actions" and a plugin marketplace. The industry is converging on open protocols like MCP rather than vendor-locked plugin systems.

Example: In Claude Desktop, you install an MCP server plugin for GitHub. Now Claude can read repos, create PRs, and review code. In ChatGPT, a custom GPT with an "Action" does the same thing but only works inside ChatGPT. MCP makes plugins portable across any AI client.
Techniques

Human in the LoopHITL

A design pattern where a human reviews, approves, or corrects an agent's actions before they take effect. Instead of letting an agent act fully autonomously, HITL adds checkpoints, critical for high-stakes decisions like deployments, financial transactions, or sending emails on your behalf.

Example: A coding agent in Cursor proposes a refactor across 12 files. Before applying changes, it pauses and shows you a diff. You review, approve, or reject. That approval step is human in the loop. Without it, the agent could introduce bugs unchecked.
Techniques

RAGRetrieval-Augmented Generation

A technique where an LLM retrieves relevant documents from a knowledge base before generating a response. Instead of relying solely on what it memorized during training, it fetches real, up-to-date information first, dramatically reducing hallucinations.

Example: A company chatbot using RAG runs a semantic search on Elasticsearch across internal docs using vectors, then feeds relevant paragraphs to the LLM as context. Frameworks like LangChain and LlamaIndex make this easy to wire up.
Models

EmbeddingsEmbedding Models

Specialized models that convert text, images, or other data into vectors (lists of numbers that capture meaning). They power semantic search and RAG by making it possible to compare content by meaning rather than keywords. Modern embedding models can be multilingual and multimodal.

Example: Jina AI (now part of Elastic) builds embedding models that handle 89+ languages and can encode both text and images into the same vector space. Elasticsearch uses these models under the hood when you index a semantic_text field, turning your documents into searchable vectors automatically.
Techniques

Vectors

A list of numbers that represents the meaning of text, images, or other data. Similar meanings produce similar vectors. Embedding models convert content into vectors so computers can compare meaning mathematically.

Example: "I love dogs" and "I adore puppies" have nearly identical vectors despite different words. Stored in vector-capable databases like Elasticsearch, Pinecone, or Weaviate, this powers semantic search and RAG.
See it in action
AI Assistant
Online
Press Play to watch an AI agent (Claude, ChatGPT, Gemini, etc.) handle a real request, step by step.
How everything connects
Click a node to explore
Models
Infrastructure
Techniques