BranchPilot

Evaluation

A deterministic regression set, with its limits stated explicitly.

The offline evaluation is designed to catch retrieval and permission regressions during development. It is not a production-quality benchmark and does not justify an accuracy percentage claim for real customer data.

12

Source retrieval

Membership points, returns, register support and branch-local procedures.

5

Permission isolation

Branch A/B negative leakage plus support/admin positive visibility.

3

Abstention / boundary

Unrelated medical/payroll questions and a grounded “safe code is not stored” case.

Latest local offline run

20 / 20 cases passed with the deterministic test-only hash embedding provider and a 0.30 minimum-similarity threshold during repository verification.

What is measured

  • Expected source is top-ranked for in-scope synthetic questions.
  • Forbidden branch documents are never eligible for the actor.
  • Questions with no sufficiently similar approved source abstain.
  • Citation IDs must be members of the retrieved set.

What is not measured

Real-world document diversity, OCR quality, adversarial multilingual prompts, production latency, model-provider drift and real customer error rates are not established by this synthetic set. Those need a larger representative corpus before any production claim.