In plain English
If training is studying, inference is sitting the exam. Every time you ask a chatbot a question and it answers, that's inference.
In practice
Inference is where ongoing AI costs sit: per-token API pricing, GPU servers and response times all relate to inference. For high-volume tasks, a smaller, faster model often beats a large one on cost without losing much quality.
Under the hood
Inference runs a forward pass of the model on new inputs with fixed weights. For language models it is autoregressive, generating one token at a time, so latency grows with output length. Quantisation, batching and caching reduce its cost.
Example
"Inference costs dropped after we moved routine questions to a smaller model."