Input: pick where the true answer sits in a long (50+ sentence) maintenance log — near the start, buried in the middle, or near the end — then run the real retrieval prompt yourself.
Outcome: no invented score — the full test document is shown openly, with its exact word count, so you can judge for yourself how substantial the test really is. Two "lookalike" facts are planted elsewhere in the passage, same format, different values, so a wrong answer means something real: the AI grabbed a similar-but-wrong fact, not that it simply missed the needle.
Useful for: seeing a genuine, verifiable retrieval test — and appreciating, either way, what the result actually shows about how reliable long-context retrieval is at this scale.
The question that will be asked: "What part number was used to replace the fuel filter for Generator Unit 3, and on what date?" Only one sentence in the whole passage answers this exactly. Two other sentences mention similar-looking part numbers and dates for different generator units and components — planted deliberately as lookalikes, not as the answer.
—sentences in this test document
—words the AI has to read before answering
Scale vs. published research tests (which use 10,000–100,000+ words)
This is the full test document, shown openly below — nothing hidden. Read it yourself before or after running the prompt.
the true needle (answers the question)lookalike distractor (similar format, wrong answer)
This is a genuinely fixed, deliberately long passage — not something you type in — so the test is consistent every time you run it. The correct answer is: FX-3387, March 14th. If an AI tool answers with FX-2210 or FX-4051 instead, or a wrong date, that's a real, observable retrieval error — not a claimed one.