Skip to main content
Duckie agents improve through a loop: test, deploy carefully, observe real runs, then update the agent. This is an operating model, not just a debugging process. Validate before public replies, watch early performance, and keep improving based on real conversations.

Key Concepts

The Testing and Rollout Loop

  1. Configure the agent with knowledge, guidelines, guardrails, runbooks, workflows, and tools.
  2. Use Test > Playground for fast, interactive scenario testing.
  3. Use Test > Replay Chats to compare Duckie against real historical conversations.
  4. Turn important scenarios into Test > Batch Test suites.
  5. Run Batch Tests with a Rubric and optional Agent test instructions.
  6. Create a deployment in Testing mode with Internal notes only and No write actions when needed.
  7. Review Analyze > Runs and fix gaps in knowledge, guidelines, guardrails, runbooks, workflows, or tool access.
  8. Switch the deployment to Live only after quality is consistent.
  9. Monitor Performance, Breakdown, Runs, and Alerts after launch.
  10. Repeat the loop after major product, policy, or workflow changes.

What to Observe

Examples

Signs You Should Iterate

  • The agent escalates too often or too rarely.
  • Runs show repeated failed tool calls.
  • Customers ask questions that knowledge does not answer.
  • Batch Test scores drop after a policy, product, or prompt change.
  • Resolution rates vary sharply by category or attribute.
  • Reviewers frequently reject approval requests.

Testing

Learn about Playground, Replay Testing, and Batch Testing.

Runs

Inspect execution details and outcomes.

Performance Metrics

Track volume, resolution, deflection, escalation, and timing.

Knowledge Gaps

Find and close unanswered questions.