Do Long Context Windows Actually Get Used?
Models advertise windows of a million tokens. A new benchmark asks how much of that window they can really use when it counts.
In one paragraph
A preprint introduces LongSpan, a benchmark that places answer-bearing passages at controlled positions inside synthetic documents from 8,000 up to 512,000 tokens and tests 14 language models on six tasks. Average accuracy fell from 91 percent at 8,000 tokens to 64 percent at 128,000, and evidence in the middle third of a document was on average 14 points harder than evidence at the start or end. The authors put the effective window at roughly 30 to 40 percent of the advertised one. The work is not peer reviewed.
What did the researchers set out to measure?
Sellers report the longest input a model accepts, which is not the same as the length over which it reasons well. The authors of a preprint called LongSpan wanted to measure the second thing, which they call the effective context length.
They built synthetic and semi-synthetic documents from 8,000 up to 512,000 tokens. The passage holding the answer was placed at the start, the middle or the end. The semi-synthetic versions came from public-domain text, and the misleading passages were chosen to look like the evidence, so a model could not simply spot the sentence that did not fit. Answers had to cite the supporting passage as well as state the answer.
Six task types covered simple retrieval, multi-hop reasoning and aggregation across several passages. Each condition had 3,600 test items per context length and was run with three prompt phrasings. Fourteen publicly described models took part, and the answers were scored automatically against reference answers.
How much of the window do models actually use?
Across the 14 models, average accuracy fell from 91 percent at 8,000 tokens to 64 percent at 128,000. The decline was gradual, not a cliff, which makes the loss easy to miss in ordinary use.
Position mattered as much as length. When the evidence sat in the middle third of the document, accuracy was on average 14 points lower than when it sat at the start or the end.
The gap between models was wide. The best model lost only 9 points between 8,000 and 128,000 tokens, while the weakest lost 41. But the authors caution that the ranking changed with the task: no single model led everywhere.
They then defined the effective window as the length at which accuracy stays within 10 points of the short-context score. By that measure it was roughly 30 to 40 percent of the advertised window. Task type mattered too. Simple retrieval held up much better than jobs that required combining several facts.
Why does the effective context length matter?
If the usable window is smaller than the advertised one, then loading everything into a prompt can quietly lower quality, and the model still answers fluently, which makes the errors hard to notice.
A fair reading is that long documents deserve a placement decision. Put the material that matters most at the start or the end rather than burying it. Trim what you can. The paper also suggests that retrieval tools, which fetch just the relevant passages, are not made obsolete by long windows; the two approaches look complementary.
The result fits what many people who work with long inputs report anecdotally. What it changes is the default assumption that a larger advertised window means more usable material. What it does not change is that some tasks stayed accurate at long lengths, so the right question is not whether long context works in general but whether it works for the job at hand.
What does the paper not show?
Read before you quote it
It does not show that long context windows are useless. Several tasks stayed accurate at long lengths, and some models degraded much more slowly than others.
It also does not show why accuracy drops. The pattern is documented; the underlying mechanism is not tested, whether that is attention, training data or something else. Nor can the study rule out that the gap would shrink with different training or prompting, since it captures systems as they were when the tests ran.
The authors list several limits. The documents are synthetic or only semi-synthetic, so they may not resemble real long documents, which carry structure, repetition and cross references. Automatic scoring may mark some correct paraphrased answers as wrong; a manual check of 200 items found this affected a small share of results. The models were queried through public interfaces whose versions can change without notice, so a repeat run may not produce identical numbers.
And because the effective window is measured against short-context accuracy, a model that is mediocre at short lengths can look robust simply because it had little to lose. The paper is a preprint and has not been through peer review.
- It does not show that long windows are useless; some tasks held up well to long lengths.
- It does not show why accuracy drops, only that it does under these conditions.
- The documents are synthetic or semi-synthetic, so they may not resemble real long documents.
- Models were tested through public interfaces whose versions can change without notice, which limits exact replication.
Takeaways
- saysAverage accuracy fell from 91 percent at 8,000 tokens to 64 percent at 128,000 tokens across the 14 models.
- saysEvidence in the middle third of a document was on average 14 points harder for these models to use than evidence at the start or end.
- explainsA context window is the length of text a model can take in at once; accepting a long input is not the same as reasoning well over it.
- interpretsA practical reading is to place the most important material at the start or end, trim what can be trimmed, and test your own task at your own length.
- interpretsRetrieval pipelines and long windows look complementary rather than substitutes, so fetching the relevant passages still has a role.
Ask the paper
In PaperSays you can ask this paper anything, by text or by saying “Hey Aiden”. Three questions listeners asked:
How much of an advertised context window is actually usable?
Roughly 30 to 40 percent, by the paper's measure of staying within 10 points of short-context accuracy. Hear the full answer in the app →
Where should evidence sit in a long document?
At the start or the end; the middle third scored on average 14 points lower. Ask in the app →
Why does accuracy drop as inputs grow?
The paper documents the drop but does not test the cause. Ask in the app →
Briefing reviewed by an editor before publication. The paper belongs to its authors; this page summarises it and links to the original. Reviewed by Dr. H. Lindqvist. Paper licence: CC BY 4.0. Corrections: support@papersays.com.