A million-token context window sounds like the end of your retrieval problems. I get the appeal. If you can fit the whole manual, repo, or incident history into one prompt, it feels like the model should finally have everything it needs. But bigger context is not the same as better answers. This post walks through what recent research actually shows, where long-context models still break down, and why a smaller, better-curated prompt often works better.










