Samuel Birhanu
software engineer — agentic AI, Addis Ababa, Ethiopia
ragbench-lite, in plain English
A lot of AI products answer questions by reading your own documents. They sound convincing even when they're wrong, and that's the problem: you can't tell a good system from a lucky one without testing it.
ragbench-lite is that test. I wrote ten questions about a set of real documents, where I already know the right answer and which document it lives in. The harness asks each question, then checks two things: did the system actually find the right document, and does the answer stick to what that document says? The report is the result — every question, every score, out in the open.
If a client hands me a chatbot that "works", this is how I'd prove it works.
Under the hood
The corpus is real markdown — Agent Barn's architecture decision records and Klikt's docs — split into overlapping chunks with LangChain's RecursiveCharacterTextSplitter.
- Retrieval: a BM25 index (rank_bm25) over the chunks; the gold source document must appear in the top-k retrieved chunks for recall@k to score.
- Answers: generated through LiteLLM (OpenRouter, a cheap flash-class model), grounded in the retrieved context.
- Judging: an LLM-as-judge scores faithfulness (every claim checked against the retrieved context, 0–1) and relevance (answer vs a human-written reference answer).
- Current scores on the live report: 10 golden questions, recall@5 = 0.9, faithfulness = 1.0, relevance = 0.95. The one recall miss is known and traced to a document missing from the corpus.
- Shape: five small modules — ingest → retrieve → answer → evaluate → report —
and one
make demothat reproduces everything end-to-end, including the self-contained HTML report.