Token
The smallest unit of text a model processes. It can be a word, part of a word, or even a symbol.
The
dragon
rests
in
ag
ony
.
In context: Text is split into small units the model can read. Each token is linked to a number, like its own ID. Notice how "agony" becomes two tokens: "ag" and "ony".
Tokenization
The process of breaking text into small units (tokens) a model can understand. Each token can be a word, subword, or character.
"Hello, how are you?"
→
Hello
,
how
are
you
?
In context: When you send text to a model, this is the very first step before anything else happens.
Embedding
The model turns each token into numbers that represent its meaning. Tokens with similar meanings have vectors (points positioned in space) that are close together.
cat
dog
pet
car
truck
vehicle
happy
joyful
king
queen
In context: Turns tokens into points in space, grouped by meaning. "Cat" and "dog" end up close together; "car" and "truck" form another cluster.
Context Window
The limit of how much text a model can consider at once. It reads and reasons only within this window, measured in tokens.
The quick brown fox jumps over the lazy dog. This sentence contains
all the letters of the alphabet. It's often used for typing practice...
Context Window (visible to model)
In context: The model can process a limited number of tokens at once. Modern models range from 4K to 200K+ tokens.
Latent Space
An internal map where the model organizes what it has learned. Each point represents a concept, and similar ideas group close together.
In context: Each dot is an embedding, placed near others with similar meaning. It's how the model organizes concepts to relate them efficiently.
Neural Network
A network of connected layers that learn from examples. Each layer refines the data, and together they learn patterns used to recognize images, understand language, or process sounds.
In context: Each layer transforms information a bit, finding patterns. By the end, the network can turn a question into the right answer.
Parameter
Values the model learns during training that determine how strongly different parts of the network connect and respond. Together, they define how the model understands and generates information.
0.42
-0.17
0.89
0.03
-0.56
0.71
-0.28
0.94
0.15
-0.63
0.37
-0.82
7B
70B
405B
In context: More parameters typically mean a smarter, more flexible model—but not always. Quality of training matters too.
Model
A system that has learned from data and can now use that knowledge to predict, generate, or understand new information.
Why is the sky blue?
Light scatters...
Translate: Hello
Bonjour
Summarize this...
Key points are...
In context: A neural network trained on tons of examples so it can predict or generate new things.
Transformer
A type of neural network that looks at every word in a sequence at once. Unlike earlier models that read step by step, it learns how words relate across the whole text.
In context: Understands relationships between words across a whole sentence. The model sees that "sky" and "orange" are connected.
Attention
A mechanism inside Transformers that decides which words to focus on when processing a sentence. Each word looks at others and assigns more weight to the ones that matter most.
In context: Helps the model decide which words to focus on for meaning. When processing "The", it pays most attention to "sat" and "mat".
Pre-training
The first learning stage where a model trains on vast text data to learn patterns, context, and general knowledge.
books
websites
articles
code
papers
forums
↓
Base Model
In context: Like going to school before specializing. The model learns language fundamentals first.
Fine-tuning
Training a pre-trained model on new, specific data so it adapts to a particular task or tone. It keeps what it already knows but learns to apply it in a focused way.
Base Model
+
Medical Data
→
Medical Assistant
In context: Teaches the model a new skill without forgetting what it already knows.
Reinforcement Learning
A training method where the model improves through feedback. It tries actions, receives rewards or penalties, and learns to make better decisions over time.
Model
→
Response
→
Feedback
↩
In context: Improves the model through trial, error, and feedback until it gets better results. RLHF (from human feedback) is commonly used.
Chain of Thought
Step-by-step reasoning the model writes to reach an answer. It helps break complex problems into smaller, more manageable steps.
1
Read the problem
2
Identify what we know
3
Apply the formula
4
Calculate result
In context: Shows how the model thinks through a problem before answering. Often improves accuracy on complex tasks.
Inference
The stage where a trained model uses what it has learned to generate a response. It predicts the next token step by step until the answer is complete.
The capital of France is
Paris
|
In context: This is what happens behind the scenes when you use an AI product. Training teaches; inference applies.
RAG
Retrieval-Augmented Generation. A method that lets a model look up information before answering. It retrieves relevant data from external sources, then uses that context to write a better answer.
Query
→
📄
📄
📄
→
Model
→
Answer
In context: When AI searches the web or your documents before answering, that's RAG in action.
Agent
An autonomous system that uses tools and feedback loops to accomplish tasks. It can plan, execute actions, observe results, and adjust its approach.
🤖
Agent
Search
Code
Browse
In context: Unlike a simple chatbot, an agent can take actions—like searching, writing code, or booking appointments.
Workflow
A predefined sequence of steps where each stage uses the previous result to move the task forward toward a final outcome.
Collect
→
Process
→
Generate
→
Deliver
In context: Connects steps into a clear, predictable path. Unlike agents, workflows follow a fixed sequence.
LLM
Large Language Model. A very large neural network trained on vast text data to understand, predict, and generate human language.
Hello
Bonjour
Hola
你好
مرحبا
こんにちは
In context: ChatGPT, Claude, and Llama are all LLMs. They power most AI chat applications today.
Prompt
The text input you give to an AI model to tell it what you want. It's your question, instruction, or the context that guides the model's response.
Prompt
"Explain quantum computing like I'm 5"
↓
Imagine you have a magic coin...
In context: The better your prompt, the better the response. Prompt engineering is the art of crafting effective instructions.
Hallucination
When an AI model generates information that sounds confident but is actually false or made up. The model doesn't "know" it's wrong.
Who wrote "The Azure Gardens"?
✗
"The Azure Gardens" was written by Margaret Chen in 1987...
(This book doesn't exist)
In context: Models predict likely text, not truth. Always verify important facts, especially for names, dates, and citations.
Temperature
A setting that controls how random or creative the model's outputs are. Low temperature = more focused and predictable. High temperature = more varied and creative.
0.0
Focused
"The sky is blue."
0.7
Balanced
1.0+
Creative
"The sky dances in azure whispers..."
In context: Use low temperature for factual tasks, higher for creative writing or brainstorming.
Few-shot Learning
Teaching a model what you want by showing it a few examples in your prompt. The model learns the pattern and applies it to new inputs.
happy →
sad
hot →
cold
fast →
slow
In context: Instead of explaining rules, you show examples. The model figures out you want opposites.
Zero-shot
Asking a model to do a task without giving any examples. The model relies entirely on its pre-trained knowledge to understand and complete the task.
"Classify this review as positive or negative:"
"The food was amazing!"
→ Positive
In context: Modern LLMs are good at zero-shot tasks because they've seen similar patterns during training.
Multimodal
AI that can understand and work with multiple types of input: text, images, audio, video, or any combination of these.
📝 Text
🖼️ Image
🎵 Audio
🎬 Video
↓
One Model
In context: You can show an image and ask questions about it. The model "sees" and understands both text and visuals.
Diffusion
A technique used by image generators. It starts with random noise and gradually removes it step by step, guided by your text prompt, until a clear image emerges.
In context: DALL-E, Midjourney, and Stable Diffusion all use this technique to generate images from text.
Vector Database
A database designed to store and search embeddings (numerical representations of data). It finds similar items by measuring distance between vectors.
Query: [0.2, 0.8, 0.5]
[0.2, 0.7, 0.6]
[0.9, 0.1, 0.3]
[0.5, 0.5, 0.5]
In context: Powers semantic search and RAG systems. When you search "happy dog photos", it finds images with similar meaning, not just matching words.
Semantic Search
Search that understands meaning, not just keywords. It finds results that are conceptually similar to your query, even if they use different words.
Keyword
"car repair"
❌ "auto mechanic"
Semantic
"car repair"
✓ "auto mechanic"
In context: Searching for "how to fix a broken heart" returns emotional advice, not cardiology articles.
GPU
Graphics Processing Unit. Originally designed for rendering graphics, GPUs excel at the parallel math operations that AI models need, making them essential for training and running AI.
In context: AI training requires massive parallel computation. A single GPU can do thousands of calculations simultaneously.
Quantization
Compressing a model by reducing the precision of its numbers. Instead of using 32-bit numbers, it might use 8-bit or 4-bit, making the model smaller and faster with minimal quality loss.
32-bit
0.7823456789
Large
→
4-bit
0.78
8x smaller
In context: Lets you run large models on smaller devices. A 70B model quantized to 4-bit can run on a laptop.