Foundations

Training

The learning phase: the AI makes guesses, gets told how wrong it was, and adjusts itself a little. Repeated millions of times.

In everyday terms

Training is slow and expensive: weeks on thousands of chips for big models. It happens once (or occasionally), before you ever use the model.

For professionals

Iterative optimisation, typically stochastic gradient descent, that updates parameters to reduce a loss function. For LLMs: pre-training on next-token prediction, then post-training (instruction tuning, RLHF).

Think of it like…

Practising darts: throw, see how far off you were, adjust your aim, repeat.

You've already seen it

When people say a model has a "knowledge cut-off date", that's when its training data ended.

Myth vs reality

Myth: ChatGPT learns from every conversation as it happens.

Reality: Models don't update themselves while you chat. Changes come from separate, later training runs.

Quick check

During training, the model…

Show answer

Repeatedly adjusts its numbers to reduce mistakes: Training = guess, measure error, adjust.

Builds on

Machine Learning (ML)

Related

Inference · Parameters · Fine-Tuning

🔎esc