Evaluation
A deterministic regression set, with its limits stated explicitly.
The offline evaluation is designed to catch retrieval and permission regressions during development. It is not a production-quality benchmark and does not justify an accuracy percentage claim for real customer data.
12
Source retrieval
Membership points, returns, register support and branch-local procedures.
5
Permission isolation
Branch A/B negative leakage plus support/admin positive visibility.
3
Abstention / boundary
Unrelated medical/payroll questions and a grounded “safe code is not stored” case.
Latest local offline run
20 / 20 cases passed with the deterministic test-only hash embedding provider and a 0.30 minimum-similarity threshold during repository verification.
What is measured
- Expected source is top-ranked for in-scope synthetic questions.
- Forbidden branch documents are never eligible for the actor.
- Questions with no sufficiently similar approved source abstain.
- Citation IDs must be members of the retrieved set.
What is not measured
Real-world document diversity, OCR quality, adversarial multilingual prompts, production latency, model-provider drift and real customer error rates are not established by this synthetic set. Those need a larger representative corpus before any production claim.