Read a small sample deeply
The first few live conversations matter more than a large transcript pile. Read them end to end so you see how the agent opens, resolves, and escalates.
Look for friction: slow starts, unclear policy language, or turns where the customer asked something the knowledge base did not cover.
Label what went right and wrong
- Correct answer with clear supporting context.
- Correct escalation when the agent should stop.
- Missing knowledge or a vague answer that needs cleanup.
Set a review rhythm for the first week
Daily review is enough for phase one. The important thing is consistency: identify what to fix, update the right Agent Builder section, and then verify the change in the next transcript sample. For recurring behavior gaps, capture the fix as a memory in Train.
