Results
Establishing a retrieval baseline before scaling an agent rollout
An assessment engagement that stopped a rollout, measured the actual failure, and restarted it on a defensible baseline.
Client not named — engagement covered by a confidentiality term.
Problem
An internal support agent was performing well in demonstration and poorly in production. The team had assumed a model limitation and had begun scoping an upgrade. No evaluation harness existed, so nobody could say which answers were wrong or why.
Approach
We assessed the seven stages against the live estate. Sources and structure placed at L0 and L1: the knowledge base held superseded policy documents with no supersession marking, so retrieval was returning them as readily as current ones. We built an evaluation set from known-correct answers and measured the baseline before proposing any change.
-
7 stages
Assessed against live systems
Scope of the assessment engagement as delivered.
-
L0–L1
Sources and structure placement
Maturity placement recorded during assessment, evidence attached in the report.
-
3 weeks
Kickoff to blueprint
Actual elapsed engagement time.
The finding was not that the model was weak. It was that no stage of the supply chain distinguished a current policy document from a superseded one, so the retrieval layer had no signal to rank on.