Library

I don't collect papers. I collect ideas that change how I build systems.

Why it's in my library

This is the paper that convinced me retrieval hit-rate isn't enough to evaluate a RAG system. A model can retrieve the right passage and still fail to use it, just because of where it landed in the context window. That's why eval suites need a faithfulness metric separate from retrieval accuracy.

Applied in

Key ideas

Long ContextEvaluationRAGFaithfulness

Read paper ↗arxiv.org