ragbench-lite evaluation report

10 golden questions · BM25 top-5 retrieval · answer model openrouter/google/gemini-3-flash-preview · judge openrouter/google/gemini-3-flash-preview

90%
recall@5
1.00
faithfulness (LLM judge)
0.95
answer relevance (LLM judge)
QuestionRecallFaith.Relev.
How are Agent restore points captured and restored, and why tar jobs instead of CSI snapshots? hit
runtime-and-deployment.md, CHANGELOG.md, INDEX.md
1.00 1.00
answer
Agent restore points are captured by running a Kubernetes Job that archives an Agent's persistent volume onto a per-restore-point PVC. This mechanism is used instead of CSI `VolumeSnapshot` or PVC `dataSource` cloning because the default `local-path` storage class used in k3s and k3d supports neither. Using tar jobs ensures the design works across every storage class, including local development environments. Additionally, tar jobs allow for the deliberate exclusion of credential material and files regenerated on boot, which a byte-exact CSI snapshot could not distinguish.

faithfulness: The answer accurately reflects all claims made in the context regarding the use of Kubernetes Jobs for restore points and the reasons for avoiding CSI snapshots.

relevance: The answer covers all key points of the reference while providing additional accurate technical context regarding storage class compatibility and exclusion sets.

How does the transactional domain-event outbox work? hit
CHANGELOG.md, 2026-07-25-transactional-domain-event-outbox.md, domain-events.md
1.00 1.00
answer
The transactional domain-event outbox persists internal Domain Events using a PostgreSQL database to ensure atomicity between business mutations and event intent. When a domain-specific repository operation produces an event, it writes the business state, one Outbox Message, and one Event Delivery for each registered handler using a single SQLModel session and commit. This approach keeps the domain code transport-neutral before introducing Dramatiq/Redis delivery. By avoiding separate sessions or worker-based delivery row creation, PostgreSQL serves as the durable source for both publication and handler-delivery state.

faithfulness: The answer accurately summarizes the transactional outbox mechanism, its components, and its purpose as described in the context.

relevance: The answer correctly identifies the core mechanism of writing events to a database table within the same transaction as the state change to ensure atomicity, while providing additional implementation details.

Why are organization roles and agent access roles separate? hit
2026-07-18-permission-backed-organization-roles.md, 2026-07-18-assigned-agent-access-boundary.md, INDEX.md
1.00 1.00
answer
Organization roles and agent access roles are separated to decouple resource visibility from organization-level permissions. This separation allows ordinary members to see and act only on specific Agents assigned to their membership, rather than all organization-owned resources. Additionally, authorizing named permissions instead of role names avoids coupling every authorization site to a fixed enum, facilitating the future introduction of organization-defined roles. This structure also ensures that Agent creation grants immutable provenance and explicit access that is not erased by subsequent organization-level role changes.

faithfulness: The answer accurately reflects the context's points regarding the separation of roles, the use of named permissions for future flexibility, and the handling of agent creator provenance.

relevance: The answer correctly identifies that the separation decouples organization-level permissions from specific agent visibility and allows for independent assignment, covering all key points of the reference answer with additional detail.

How does cost tracking and spend attribution work for organizations and agents? hit
INDEX.md, AF-281-cost-tracking-findings.md, costs.md
1.00 1.00
answer
Cost tracking is managed through the `cost_record` table, a cost sync/healing CronJob, and specific spend limits for Organizations and Agents. Spend attribution is derived from provider-reported costs rather than a local price map, as the system relies on the provider as the source of truth. All cost history survives agent deletion, as cost rows do not cascade and the platform must render spend for non-existent agents. Current limitations include a lack of per-user or per-conversation attribution and the inability to track cache-read token costs. Infrastructure costs are explicitly excluded from the tracking scope.

faithfulness: All claims in the answer, including the management mechanisms, data retention policies, and specific limitations, are directly supported by the provided context.

relevance: The answer covers all key points from the reference, including per-agent ingestion, the CronJob reconciliation, and organization spend limits, while providing additional accurate technical context.

How do Agent Webhooks deliver signed machine triggers? hit
INDEX.md, system-map.md, agent-webhooks.md
1.00 1.00
answer
Agent Webhooks deliver signed machine triggers through an authenticated HTTP ingress at `/agent-hooks/v1/{webhook_id}`. The Agent Webhooks domain manages admission, caller idempotency, and durable invocation history. Once admitted, the system performs a native job submission to a runtime (Hermes or OpenClaw), which executes the resulting one-shot Agent Trigger Job. Delivery to the agent occurs through a runtime-owned Communication Connection. These webhooks are agent-subordinate resources that remain separate from Platform or Communication Connections.

faithfulness: Every claim in the answer, including the specific ingress path, domain responsibilities, and the relationship between webhooks and connections, is directly supported by the provided context.

relevance: The answer covers all key points of the reference, including signed HTTP triggers, job submission, and decoupling from the communications platform, while providing additional technical detail.

What is the runtime configuration digest used for agent update prompts? hit
agents.md, runtime-and-deployment.md, CHANGELOG.md
1.00 1.00
answer
The `running_config_digest` is a record of the platform-authored code and runtime images a specific Agent's pod was built from. It is derived from the static closure of the Agent assembly code plus the runtime image references. The system compares this recorded digest against what the API would currently build to determine if an update is available. When these values differ, `AgentRead.update_available` becomes true, triggering an advisory update signal. This signal is purely advisory, blocks no operations, and clears once the Agent is restarted. It specifically excludes the Agent's own configuration, such as `pending_model` or Skill-level updates.

faithfulness: Every claim in the answer is directly supported by the provided context, including the definition of the digest, its derivation, the comparison logic, and its advisory nature.

relevance: The answer provides a highly detailed and accurate explanation that covers all key points of the reference answer, including the comparison mechanism and the update_available flag.

How do organization templates move from draft to published? hit
templates-and-skills.md, templates-and-skills.md, 2026-09-04-organization-templates-use-draft-publish.md
1.00 1.00
answer
Organization templates move from draft to published through an explicit publishing action, which creates the next immutable version of the template. This process adopts the Platform Template model, where a single mutable draft exists in `agent_template_draft` for each lineage. Content changes are draft-gated, and the previous `PATCH` method that created a version on every save has been removed. Once published, the new version is immutable and does not mutate the lineage's history.

faithfulness: The answer accurately reflects the transition from version-per-save to a draft-publish model where publishing creates an immutable version, as described in the context.

relevance: The answer covers all key points of the draft/publish lifecycle, including the creation of immutable versions and the transition from mutable drafts, while providing additional technical context.

How do required skill versions constrain agent configuration? MISS
INDEX.md, CHANGELOG.md, CHANGELOG.md
1.00 0.90
answer
Required skill versions constrain agent configuration by enforcing exact version matches during the `select_agent_template` process. While Override drafts and published versions only validate the presence, visibility, and provider requirements of a skill, the exact version pin is not activated until selection. An agent cannot be selected if its current pins do not match the required skill versions specified in the configuration. Additionally, skill assignments pin explicit versions at apply time, and publishing a newer skill version does not automatically update an existing pin. Recovery from a bad skill version requires manually re-pinning the agent to an older version.

faithfulness: The answer accurately reflects the context regarding skill version enforcement during selection, the validation rules for drafts, and the manual re-pinning process for skill versions.

relevance: The answer correctly identifies that agent selection enforces version requirements and mentions the specific selection process, though it omits the specific UI sections mentioned in the reference.

What does the platform Event Delivery Monitor track? hit
domain-events.md, domain-events.md, domain-events.md
1.00 0.70
answer
The Platform Event Delivery Monitor tracks **Event Deliveries**, specifically monitoring the health of the Event Delivery pipeline. It does not track Outbox Messages, as these represent immutable publication intent without a lifecycle status. The monitor tracks the `PENDING → ENQUEUED → PROCESSING → SUCCEEDED | DEAD_LETTERED` lifecycle of mutable delivery rows. Additionally, it surfaces curated Actor/Subject display strings when included in the event payload. The tool is strictly read-only and intended for Platform Administrators to inspect global delivery health.

faithfulness: Every claim in the answer, including the specific lifecycle states and the distinction between Event Deliveries and Outbox Messages, is directly supported by the provided context.

relevance: The answer correctly identifies event deliveries and their lifecycle but explicitly contradicts the reference by stating it does not track outbox messages.

How does interactive command approval work for agents? hit
INDEX.md, runtime-and-deployment.md, CHANGELOG.md
1.00 0.90
answer
Interactive command approval is currently limited to Hermes-only agents, as OpenClaw agents do not yet have user-configurable controls and default to `AUTO`. The system maps the `approval_mode` field to the runtime's approval policy, pinning settings like `timeout` and `cron_mode` to prevent silent policy shifts during upgrades. Headless or cron-driven runs default to a `deny` policy, refusing dangerous commands immediately since no human is present. The exception is `BOOT.md` runs, which the runtime treats as interactive sessions, causing flagged commands to park until the timeout expires. For supported agents, the delivery state involves the runtime adapter, Web Chat, and specific platform slices.

faithfulness: The answer accurately reflects the context regarding Hermes-only support, the pinning of approval settings, the handling of headless/BOOT.md runs, and the delivery state components.

relevance: The answer covers the same key points as the reference, including the role of the runtime adapter and Web Chat, while providing additional technical context about specific agent types and policies.