Also known as: context length
In plain English
The context window is the model's working memory. Everything in the conversation, plus any documents you paste in, has to fit. If it's too long, the oldest parts get cut off or the request fails.
In practice
Larger context windows let you analyse whole contracts or codebases in one go, but bigger isn't always better: cost rises with every token, and models can miss details buried in the middle of very long inputs. For large document sets, RAG is often cheaper and more accurate.
Under the hood
The context window is the maximum sequence length a model was trained or extended to handle. Frontier models now support hundreds of thousands to over a million tokens. Attention compute and key-value cache memory both grow with context length.
Example
"The whole policy manual fits in the model's context window."