background shape
Case Study · AI-First Platform

How Progress Support uses Agent Engineering to catch escalations earlier and keep AI agents on track

Two internal AI agents review customer conversations for security and customer-experience escalation signals. Progress Agent Engineering lets the Support team see how those agents behave in production, check whether they made the right call, track what each interaction costs and test changes against real support cases.

100% of in-scope customer messages

Checked by AI agents for escalation signals.

24/7 automated monitoring

Potential security and customer-experience risks can be surfaced even outside staffed support hours.

95%+ escalation accuracy

Reviewed support cases show that the agents make the expected escalation decision in more than 95% of cases.

50% faster investigation

Trace-level context cuts the time Support needs to understand and verify a flagged interaction by half.

Story in Detail

AI agents that help Support catch issues before they grow

Progress Support uses two internal AI agents to watch customer conversations for signals that may need attention.

Every new customer reply passes through them. One agent looks for potential security escalations. The other watches for customer-experience risks: a customer becoming frustrated, a problem dragging on or an interaction that may be heading toward a poor outcome.

The team does not staff Support 24 hours a day. The agents give the team another set of eyes on incoming conversations and can surface cases that deserve attention sooner.

The agents were already working well. Once they were running across real customer conversations, the harder question became clear:

 

Can the team see why an agent made each call, and know whether it made the right one?

 

That is where Progress Agent Engineering comes in.

 

1 · Runtime 2 · Observe 3 · Improve Customer reply Inbound support message Security agent eval: prompt-injection CX agent service.name: cx-agent Escalation decision Escalate or auto-resolve Agent Engineering trace agent.run tool_history eval.verdict Human review Flagged traces only Dataset input / expected Experiment variant · scored Updated agent shipped to prod Redeploy to live traffic Your agents Progress Agent Engineering Process step

From runtime to improvement: every escalation decision is traced, reviewed and reused as a test case.

Want to get started?

Progress Agent Engineering gives you full visibility into your AI agents and LLM-powered applications. It captures traces, measures costs, and evaluates output quality — so you can release AI features with confidence.

Let’s get started

The challenge: a correct-looking label is not enough

An AI agent can return a plausible decision and still get the case wrong.

A routine issue might be treated as a security escalation. A frustrated customer might slip through. A model change could affect behavior that worked yesterday. And a model that performs well may cost far more than another model that would make the same decision.

At first, the Support team checked escalations manually. Each one generated an email, and a team member reviewed whether the agent had made the right call.

That worked as an early quality check. But the final escalation label only tells you the outcome.

Progress Agent Engineering gives the team the run behind the decision.

 

progress support

The agent’s escalation call, with the reasoning it recorded on the ticket.

observability traces

The same interaction in Progress Agent Engineering — latency, tokens, cost and the full output behind the decision.

The agents help us keep an eye on customer conversations and flag cases that may need attention. With Progress Agent Engineering, we can check how they’re behaving and whether the decisions match what we’d expect.

Dragan Grigorov

Progress

See the run behind the escalation

When an agent flags an interaction, Support can inspect what happened instead of stopping at the final label.

The trace helps the team answer questions they already need to answer:
  • Why did the agent treat this as a security issue?
  • Did it interpret the customer’s message correctly?
  • Was this a real escalation or an overly cautious call?
  • What did this interaction cost to process?
  • Are we seeing the same behavior in other conversations?
This has already exposed some of the judgment calls inside the workflow. Licensing-related conversations, for example, can sometimes be flagged more aggressively than a human reviewer would choose.

Those cases are useful. They show where the agent’s judgment differs from the team’s and give Support something concrete to test.

 

Turn reviewed cases into evals

The team’s first quality system was simple: check the escalations by hand.

Those reviewed cases can now become the evaluation set.

Real production traces can be added to datasets with known security escalations, customer-experience risks, correct non-escalations and difficult edge cases. The team can rerun those cases after changing a prompt, model or part of the agent workflow.

The evaluation question can start very simply:

 

Should this customer interaction have been escalated? Yes or no?

Test model changes against real support cases

The Support team wanted to answer a practical question: could another model make the same escalation decisions, or better ones, at a lower cost?

Using Datasets & Experiments, the team replayed real support tickets through the same prompt and scored the results with the same evaluators. For escalation accuracy, the comparison uses the seven tickets that appeared in both experiment runs.

 

The same production support cases, tested against two models

 Current modelCandidate modelDifference
 Azure OpenAI GPT-5.4-miniOpenAI GPT-4.1-mini 
Escalation accuracy71.4% (5 of 7)85.7% (6 of 7)+14.3 pts
Avg. cost / interaction$0.1235$0.109711.1% lower
Avg. latency1.38 s1.14 s17.5% faster
Avg. tokens / interaction31027611.1% fewer

Escalation accuracy measures whether the model routed a ticket to a human when policy required it, rather than closing the ticket itself. The seven tickets shown for accuracy are the cases common to both experiment runs. Cost, latency and token figures are the per-item averages recorded for each full run.

 

How the team ran the comparison

The experiment can be reproduced directly in Progress Agent Engineering:
  1. Open the support-ticket dataset in Improve → Datasets & Experiments.
  2. Select Run Experiment, choose the production model and keep the existing prompt template and evaluators.
  3. Run the experiment again with the candidate model, using the same prompt and evaluators.
  4. Select both experiment runs and choose Compare Selected.
The comparison view shows average latency and token usage in the summary, evaluator results in Metrics Overview, and each ticket side by side when the team wants to inspect a disagreement.

 

11.1% lower cost, with higher escalation accuracy in the test

On the seven support tickets common to both runs, the candidate model made the expected escalation decision in six cases. The production model did so in five.

The misses matter too. Both tickets missed by the production model were closed when support policy required them to be escalated to a human. The candidate model missed one of the seven.

At the same time, the candidate run averaged $0.1097 per interaction compared with $0.1235 for the production model. It also recorded lower average latency and token usage.

Because the team can open every result row by row, a model comparison does not end with an aggregate score. Support can see which tickets changed, where a model made the wrong call and whether the cost difference is worth acting on.

One support issue can become a reusable test

A strange interaction does not have to disappear once somebody reviews it.

Support can keep the trace, label the expected behavior and add it to the cases used to test the agent later.

That creates a simple improvement loop:

 

Observe Review Evaluate Experiment Improve
A flagged conversation becomes a reviewed example. The reviewed example becomes an eval. The eval becomes part of the test set for the next prompt or model change.

The team gets more value from the same production evidence every time the agent changes.

What changed

The Support team now has a clearer operating view of its AI agents:
  • Potential security and customer-experience risks can surface earlier, including outside staffed support hours.
  • When an escalation looks questionable, Support can inspect the trace behind the decision instead of reviewing the final label alone.
  • Reviewed customer interactions can become evaluation cases that are reused after prompt, model or workflow changes.
  • Model experiments can compare quality, latency and cost against the same real support cases.

What’s next

Progress Support is expanding the workflow with larger evaluation datasets and experiments based on real customer interactions. This gives the team a repeatable way to check whether changes improve agent quality, introduce regressions or create an opportunity to lower model cost.

 

 

See how Progress Agent Engineering helps teams understand, evaluate and improve production AI agents.

Explore Progress Agent EngineeringBook a demo

Read Similar Success Stories

Icanpreneur Launches First “Accelerator-as-a-Software” Platform, Leveraging Kendo UI for a Seamless Customer Experience

Read the Whole Story
CaseStudy-Resource_Image

BAYOOTEC Delivers Integrated Digital Transformation UX 15% Faster with Kendo UI

Read the Whole Story
Ascendas Slashes Report

Ascendas Slashes Report Development Time by 50% with Progress

Read the Whole Story