Settingsintermediate

How to Review Outcomes and Guardrail Insights

The outcomes section of the Supervisor Overview classifies every conversation the Agent Stack closed without a handover, with a deterministic taxonomy instead of a single model opinion, and shows how confident that classification is. This article covers the taxonomy, the confidence tiers, the trend chart, the closed-conversation filters, the tenant opt-in for handover knowledge mining, and the superadmin Guardrails page.

8 min read

How to Review Outcomes and Guardrail Insights

The outcomes section of Supervisor’s Overview answers a question the old single resolution-rate number couldn’t be trusted to answer honestly: did the conversation actually get resolved, or did it just end? Instead of asking a model to judge “resolved vs. unresolved” from scratch, the transcript itself now decides the category and how confident that category is. The model (the judge) is only called in for the one thing the transcript can’t tell you on its own: what the customer’s last message was actually doing.

This article covers the outcomes section of the Overview, the confidence tiers behind each row, the trend chart, how to filter the closed-conversation list, the tenant control for mining handover replies into quick drafts, and the superadmin Guardrail Insights page that sits alongside it.

Where to find it

  • Outcomes sit in Supervisor’s Overview view, next to Review.
  • Guardrails is a superadmin-only page, separate from Supervisor, that looks across every tenant’s guardrail activity.
  • The handover mining control lives in the Supervisor settings menu in the Supervisor header, not inside the Overview itself — it’s a tenant-level setting, not a view filter.

How conversations are classified

Every conversation the Agent Stack closed on its own, without a handover, is assigned one outcome category. The category is derived deterministically from transcript facts — how the conversation ended, whether the customer replied again, whether the customer left a rating — not from a model’s overall impression:

  • Confirmed resolved — The customer said, in substance, that their issue was handled.
  • Assumed resolved — The conversation ended without complaint and without a follow-up question — most likely resolved, but nobody said so.
  • Unresolved — The conversation ended with an open question, complaint, or unmet request still on the table.
  • Abandoned — The customer stopped responding mid-conversation, before anything was resolved or explicitly left open.
  • Unassessable — There was nothing to read, or the conversation ended on a platform fallback reply. The row stays visible rather than being silently dropped.

The important design choice here: confirmed and assumed are never added together into one headline resolution rate. Showing a single optimistic number was the exact credibility problem this taxonomy replaces, so the Overview always shows both counts side by side and lets you read them separately.

Rows that get reopened or superseded — a later close on the same conversation, or a reopen after the close — aren’t deleted. They are marked retracted and leave the rates, and the What these rates cover note counts them as having left when the customer came back.

Conversations closed before this feature shipped don’t have a real confidence measurement, so those marked resolved at the time are shown as assumed resolved at an unknown tier rather than being assigned a confidence they were never actually measured for.

Confidence tiers

Alongside the category, each row carries a confidence tier. The tier reflects how strong the evidence behind the classification is — a short, ambiguous last message earns a lower tier than a substantive one that clearly confirms or disputes resolution. Because the tier is computed from the same deterministic transcript signals as the category, it doesn’t change from one run to the next for the same conversation.

The judge is invoked only when there’s a last customer message to interpret at all — silent exits (the majority of abandoned conversations) never need a model call, and a failed judge call does not make a row unassessable: the row is classified from the transcript alone, and a note under the tiles counts how many.

Reading the trend chart

The trend chart plots the resolution picture over time. The most recent 48 hours of the chart is drawn pale rather than solid — conversations that closed that recently can still reopen, so that segment of the trend is provisional and worth re-checking later rather than treated as final.

An info icon sits on the chart (and on every other widget in the tab — the KPI cards, the coverage note, and the trend chart) so you can hover for a short explanation of exactly what that widget is counting, without leaving the page.

Filtering the closed-conversation list

The closed-conversation list below the chart can be filtered by confidence tier, and each row shows:

  • The outcome category (confirmed resolved, assumed resolved, unresolved, abandoned, unassessable)
  • The confidence tier
  • Whether the row is retracted (reopened or superseded after close)
  • The conversation’s current status — closed conversations can live in either the done or archived inbox, and the list reflects which one so a deep link opens the right place
  • The judge model used for that row’s classification, where one was called

Use the confidence filter when you want to sanity-check the lower tiers specifically — that’s where “assumed” resolutions and thin evidence live, and it’s the fastest way to spot conversations worth a manual look.

Handover knowledge mining opt-in

Handover mining reads what your own team wrote to a customer during a human handover and counts the conversation on the waiting quick draft that already holds its question. When Supervisor is available for your tenant and the platform-side rollout flag is enabled, mining is on by default — you don’t need to opt in before eligible handover replies can be screened.

The control in the Supervisor header is a tenant-level opt-out for this mining only. Turning off Draft from handovers stops mining before any further handover replies are read, while leaving the rest of your Supervisor settings unchanged.

A few things worth knowing about the control:

  • It’s a tenant-level mining switch, separate from any other Supervisor setting — toggling it doesn’t touch anything else.
  • Mining only runs while Supervisor is available for your tenant, the platform-side flag for this capability is enabled, and Draft from handovers remains on. If the tenant-level switch is turned off, mining is skipped before anything is read, at no cost.
  • Not every closed handover produces a quick draft. There’s a deliberate screening step before a suggestion is drafted at all, so the queue only ever contains items worth a reviewer’s time; conversations that don’t clear that bar are skipped.

Quick drafts in the review queue

When a handover reply clears the mining screen, it shows up in the review queue as a quick-draft suggestion — the same place quick drafts a person explicitly asked for appear. Mined suggestions carry a badge so a reviewer can tell at a glance that a suggestion came from something a human teammate actually said to a customer, rather than being requested directly.

If the mined answer matches an existing waiting draft, Supervisor doesn’t add a duplicate suggestion. It records the conversation on that pending draft instead, and the draft’s detail view lists them under Also asked in N other conversations. That means a mined draft can carry the badge whether it was newly created from a handover reply or later accumulated repeat matches from other handovers.

If the mined answer instead matches a draft that was discarded or rejected within the last 90 days, Supervisor still does not create a new draft — but it does not throw the signal away either. It records the conversation against that discarded draft, and the row in Supervisor → History → Discarded drafts then reads “N conversations have asked this since you discarded it.” That line is the reviewer’s cue that a discard may have been the wrong call. Past 90 days the match no longer suppresses anything, and the topic can be drafted again.

The review queue is ordered by the most-asked drafts first, so repeated customer needs rise to the top. Both the queue row and the detail view show how many conversations contributed to the draft, giving reviewers a quick read on how broadly useful the answer is before they approve, edit, or reject it.

Guardrail Insights (superadmin)

Separate from Supervisor, Guardrails is a superadmin page that gives a cross-tenant view of every guardrail trigger — regenerations, forced fallbacks, forced handovers, and the other guardrail checks — in one normalized stream instead of scattered across different logs.

The page includes:

  • A KPI row: Total triggers, Forced fallbacks, Recovered by regeneration, and Triggers per 1000 turns.
  • A stacked history chart, bucketed by day — or by hour, automatically, when the selected date range is 48 hours or less.
  • Filters for guardrail, guardrail family, outcome, tenant, and date range. The guardrail filter accepts multiple values at once.
  • A paginated events table underneath the chart and filters.
  • A floating detail pane for a selected event, showing the raw detail payload, with a drill-down link that opens the underlying tenant conversation in a new tab.

Because the drill-down needs to land on the right inbox tab, each event carries the conversation’s real lifecycle status — a closed conversation won’t open into an “active” list and show up empty.

If you toggle the guardrail or date filters and see gaps or ragged bars in the chart, that’s expected for very sparse periods — the chart fills in buckets with no events client-side so the stacked chart stays readable, rather than silently dropping days that had nothing to show.

Summary

  • The outcomes section of the Overview replaces a single model-judged resolution rate with a deterministic taxonomy: confirmed resolved, assumed resolved, unresolved, abandoned, and unassessable — confirmed and assumed are always shown separately, never summed.
  • Each row carries a confidence tier; the trailing 48 hours of the trend chart is pale because those conversations can still reopen.
  • Handover knowledge mining is enabled by default when Supervisor is available and the platform-side rollout flag is on. The Draft from handovers control in the Supervisor header is a tenant-level opt-out: turning it off stops mining before further handover replies are read, without changing other Supervisor settings. Mined quick-draft suggestions carry a badge in the review queue. A matching mined answer is recorded on the existing pending draft (Also asked in N other conversations) rather than drafted twice; if the match is a draft discarded in the last 90 days, it is recorded there instead and surfaces under History → Discarded drafts as renewed demand.
  • Guardrails, under Superadmin, gives a cross-tenant, filterable view of every guardrail trigger with drill-down into the source conversation.

Tags

Ai FeaturesHow To