Skip to main content
Skip to the work
ReplyGaugeTurn failed support conversations into a reviewable test set.

Conversation quality

  1. c1ungrounded好的,我帮您查一下进度。

    Make each failed reply useful for the next review.

    t2 · Cite the matching knowledge source before answering.
  2. c2policy conflict按 kb-refund-7,现在就给您办理退款。t2 · Follow the cited knowledge; do not contradict it.
  3. c3missed handoff按 kb-refund-7,超过30天不能直接退款。t2 · Hand the conversation to a human.

First real result

A failed-conversation review set with issue labels, cited turns, expected behavior, and replayable regression cases

One label per miss, in this order, quoting the assistant turn.

  1. 01ungrounded
  2. 02policy_conflict
  3. 03missed_handoff
  • c1ungroundedsource kb-refund-7

    好的,我帮您查一下进度。

    cited turn t2

    Expected: Cite the matching knowledge source before answering. Replay answer: 我会转接人工处理。

  • c2policy_conflictsource kb-refund-7

    按 kb-refund-7,现在就给您办理退款。

    cited turn t2

    Expected: Follow the cited knowledge; do not contradict it. Replay answer: 我会转接人工处理。

  • c3missed_handoffsource kb-refund-7

    按 kb-refund-7,超过30天不能直接退款。

    cited turn t2

    Expected: Hand the conversation to a human. Replay answer: 我会转接人工处理。

02 / BEFORE A REVISION IS KEPT

The gate runs in this order. All three must pass.

  1. 01critical_failures=0
  2. 02grounded_rate>=0.95
  3. 03handoff_rate>=0.90

03 / WHAT YOU BRING

ReplyGauge is a quality bench for supplied support conversations, built to expose unsupported answers, recurring failure types, and regressions before revised guidance is accepted.

Compare an answer with the supplied knowledge behind it. Reviewers can label the miss, test a revision against saved cases, and decide whether it is ready to keep.

Sign in and you go straight back to the ReplyGauge conversation. Work happens in chat. This page never asks for your card details or email address.

THIS PAGE WILL NOT CLAIM

A review set is not a live support change.

  • Live support integrations
  • Calibrated accuracy
  • Autonomous production changes
  • Improved resolution rates
  • Verified customer outcomes

Before acting

  • Conversation evidence must remain tied to its source turn so reviewers can verify every finding.
  • Quality labels and suggested revisions can be wrong and require human calibration.

STARTER / MONTHLY

One Starter tier, listed in the catalog.

Billing appears inside the signed-in workspace. Card payment runs through Stripe Checkout.

BEFORE YOU UPLOAD A THREAD

Questions

What do I have to supply?

The support conversations and the knowledge version those replies were meant to use. ReplyGauge reviews supplied material; it does not connect to a live support channel.

What labels can a failed reply receive?

The first hit among ungrounded, policy conflict, and missed handoff. The review row quotes the assistant turn so you can check the label against the source.

What is ready after the first review?

A failed-conversation review set with issue labels, cited turns, expected behavior, and replayable regression cases. A required-handoff policy keeps the expected answer as a human handoff.

Does a passing gate change the production bot?

No. The gate checks critical_failures=0, grounded_rate>=0.95, and handoff_rate>=0.90 on the supplied set. Accepting revised guidance remains a human decision.

Can the labels be wrong?

Yes. Quality labels and suggested revisions can be wrong and require human calibration. Every finding stays tied to its source turn so a reviewer can disagree.

ReplyGauge

Signing in opens the ReplyGauge conversation. This page uses PostHog for product analytics (anonymous, optional). See Privacy.