Conversation quality
- c1ungrounded
好的,我帮您查一下进度。
Make each failed reply useful for the next review.
t2 · Cite the matching knowledge source before answering. - c2policy conflict
按 kb-refund-7,现在就给您办理退款。
t2 · Follow the cited knowledge; do not contradict it. - c3missed handoff
按 kb-refund-7,超过30天不能直接退款。
t2 · Hand the conversation to a human.
First real result
A failed-conversation review set with issue labels, cited turns, expected behavior, and replayable regression cases
One label per miss, in this order, quoting the assistant turn.
- 01
ungrounded - 02
policy_conflict - 03
missed_handoff
c1 ungroundedsource kb-refund-7好的,我帮您查一下进度。
cited turn t2Expected: Cite the matching knowledge source before answering. Replay answer: 我会转接人工处理。
c2 policy_conflictsource kb-refund-7按 kb-refund-7,现在就给您办理退款。
cited turn t2Expected: Follow the cited knowledge; do not contradict it. Replay answer: 我会转接人工处理。
c3 missed_handoffsource kb-refund-7按 kb-refund-7,超过30天不能直接退款。
cited turn t2Expected: Hand the conversation to a human. Replay answer: 我会转接人工处理。
02 / BEFORE A REVISION IS KEPT
The gate runs in this order. All three must pass.
- 01
critical_failures=0 - 02
grounded_rate>=0.95 - 03
handoff_rate>=0.90
03 / WHAT YOU BRING
ReplyGauge is a quality bench for supplied support conversations, built to expose unsupported answers, recurring failure types, and regressions before revised guidance is accepted.
Compare an answer with the supplied knowledge behind it. Reviewers can label the miss, test a revision against saved cases, and decide whether it is ready to keep.
Sign in and you go straight back to the ReplyGauge conversation. Work happens in chat. This page never asks for your card details or email address.
THIS PAGE WILL NOT CLAIM
A review set is not a live support change.
- Live support integrations
- Calibrated accuracy
- Autonomous production changes
- Improved resolution rates
- Verified customer outcomes
Before acting
- Conversation evidence must remain tied to its source turn so reviewers can verify every finding.
- Quality labels and suggested revisions can be wrong and require human calibration.
STARTER / MONTHLY
One Starter tier, listed in the catalog.
Billing appears inside the signed-in workspace. Card payment runs through Stripe Checkout.
BEFORE YOU UPLOAD A THREAD
Questions
What do I have to supply?
The support conversations and the knowledge version those replies were meant to use. ReplyGauge reviews supplied material; it does not connect to a live support channel.
What labels can a failed reply receive?
The first hit among ungrounded, policy conflict, and missed handoff. The review row quotes the assistant turn so you can check the label against the source.
What is ready after the first review?
A failed-conversation review set with issue labels, cited turns, expected behavior, and replayable regression cases. A required-handoff policy keeps the expected answer as a human handoff.
Does a passing gate change the production bot?
No. The gate checks critical_failures=0, grounded_rate>=0.95, and handoff_rate>=0.90 on the supplied set. Accepting revised guidance remains a human decision.
Can the labels be wrong?
Yes. Quality labels and suggested revisions can be wrong and require human calibration. Every finding stays tied to its source turn so a reviewer can disagree.