Samuel Birhanu

Samuel Birhanu

software engineer — agentic AI, Addis Ababa, Ethiopia

ragbench-lite, in plain English

A lot of AI products answer questions by reading your own documents. They sound convincing even when they're wrong, and that's the problem: you can't tell a good system from a lucky one without testing it.

ragbench-lite is that test. I wrote ten questions about a set of real documents, where I already know the right answer and which document it lives in. The harness asks each question, then checks two things: did the system actually find the right document, and does the answer stick to what that document says? The report is the result — every question, every score, out in the open.

If a client hands me a chatbot that "works", this is how I'd prove it works.

Under the hood

The corpus is real markdown — Agent Barn's architecture decision records and Klikt's docs — split into overlapping chunks with LangChain's RecursiveCharacterTextSplitter.

Open the live eval report →
README · source tarball · GitHub