Booking disputes and deposit refunds are the two conversation types most likely to make or break a renter's or buyer's trust in a property platform, yet most QA scorecards still treat them like generic service tickets. They shouldn't be. A missed deadline on a deposit refund carries a statutory clock attached to it; a mishandled booking dispute can trigger a regulatory complaint months later. Scoring these conversations properly means building a QA scorecard around three things: whether the agent followed the correct timeline, whether they cited the actual policy (not a generic script), and whether the tone matched the financial stress the customer was under. RevelirQA, an AI AutoQA platform, was built to score this kind of high-stakes conversation at 100% coverage rather than the 1-5% manual QA teams typically manage to review.
TL;DR
- Deposit refunds and booking disputes involve statutory or contractual timelines (14-30 days in the US, 10 days in the UK for tenancy deposits) that a QA scorecard must check explicitly, not assume.
- Manual QA sampling reviews 1-5% of tickets, which means the riskiest disputes, the ones a regulator or an angry customer escalates, are often the ones QA never sees.
- AutoQA (auto QA) scores every conversation against the platform's own policies, catching missed timelines and inconsistent refund reasoning across 100% of tickets instead of a small sample.
- Sentiment arc (how a customer's tone shifts from the start to the end of a conversation) often reveals dissatisfaction that a "resolved" ticket status hides completely.
- A good scorecard for disputes and refunds separates "was the policy correct" from "was the delivery appropriate", because agents can get both, either, or neither right.
About the Author: This guide is written by the RevelirQA team at Revelir AI, which builds automated quality assurance software used in production by Xendit and Tiket.com to score thousands of customer service conversations a week, including dispute and refund conversations in fintech and travel where policy accuracy and timeline compliance are graded, not assumed.
Why Do Booking Disputes and Deposit Refunds Need a Different QA Approach?
Booking disputes and deposit refunds are structurally different from most service conversations because they involve a deadline, a monetary amount, and often a third party (a landlord, a host, an owner) whose interests may conflict with the renter's. A standard QA scorecard built for "was the agent polite and did they resolve the issue" misses all three of those variables. It cannot tell you whether the agent quoted the correct refund window, whether they applied the platform's actual dispute policy or an approximation of it, or whether the resolution reason logged in the helpdesk matches what was actually said.
This matters because the regulatory and contractual environment around deposits is specific rather than general. In the US, earnest money refunds depend on the contingencies written into the purchase contract, while security deposits are usually governed by statutory limits of 14 to 30 days. In the UK, tenancy deposits must be returned within 10 days of the parties agreeing on the amount owed. EU rules vary by member state under national tenancy and consumer protection law. A QA scorecard that doesn't check the agent's stated timeline against the correct jurisdictional rule isn't really scoring compliance, it's scoring politeness and calling it something else.
What Should a QA Scorecard for Deposit Refunds Actually Measure?
A QA scorecard for deposit refund conversations should measure policy accuracy, timeline communication, and documentation quality as separate criteria, not one blended "resolution quality" score. Blending them hides where the failure actually occurred, and coaching an agent effectively requires knowing exactly which part broke.
- Timeline accuracy: did the agent state the correct refund window for the customer's jurisdiction and contract type, rather than a generic company-wide estimate?
- Policy citation: did the agent reference the specific clause or SOP that applies, or did they improvise an explanation that sounds plausible but isn't documented anywhere?
- Deduction justification: if any amount was withheld, did the agent explain the specific reason (damage, unpaid fees, cancellation terms) rather than a vague "as per policy"?
- Escalation trigger: did the agent correctly identify when a dispute needed to move to mediation, ODR, or a compliance team, instead of trying to resolve something outside their authority?
- Tone under financial stress: did the agent's language match the seriousness of the situation for the customer, who is often waiting on money they need?
Each of these can be scored as a binary (did/did not happen), a multi-option field (correct, partially correct, incorrect), or a scaled criterion, depending on how granular the QA metrics need to be. RevelirQA supports all three formats on the same QA scorecard, which matters here because "timeline accuracy" is naturally binary while "tone under stress" is naturally a scale.
How Should Booking Dispute Escalations Be Scored Differently From Simple Complaints?
Booking dispute escalations should be scored against a decision tree, not a satisfaction score, because the core question is whether the agent made the right call about ownership of the problem, not whether the customer left happy. A dispute over a cancelled reservation, a mismatched listing, or a double-booking usually has a correct next step defined somewhere in the platform's SOPs. This is a different problem from the timeline question in a deposit refund. Here the QA metric that matters most is whether the agent identified the right owner of the resolution (platform, host/agent, or third-party insurer) and routed accordingly.
Industry practice offers a useful reference point. The National Association of REALTORS requires arbitration or ethics complaints to be filed within 180 days of a transaction closing, and best practice across the industry favors resolving disputes through mediation or Online Dispute Resolution platforms within 3 to 30 days. A platform's own SOPs will typically be tighter than that external ceiling, so the QA scorecard should check against the company's internal SLA, not the regulatory outer limit, while flagging any case approaching the regulatory deadline as a distinct risk category.
This is where scoring 100% of conversations instead of a sample changes the risk profile. If only 3% of dispute tickets get manually reviewed, the pattern of agents quietly routing disputes to the wrong team, or not routing them at all, may never surface until a customer files a formal complaint near the 180-day mark. Scoring every conversation catches the pattern while there's still time to retrain the agent, not after the escalation has already reached a regulator.
Why Does Manual QA Sampling Miss the Riskiest Conversations?
Manual QA sampling misses the riskiest conversations because reviewers tend to pull tickets that are easy to evaluate quickly, and dispute and refund conversations are rarely quick. A reviewer working through a sampling quota will gravitate toward short, clearly resolved tickets over a long back-and-forth involving a partial deposit deduction and an unhappy customer, simply because the sample has to get reviewed on a schedule. The result is a QA sample that is statistically biased against exactly the conversations that carry the most compliance and retention risk.
This is the specific gap AutoQA is built to close. Auto QA, meaning automated quality assurance software that scores every conversation rather than a manual sample, removes the incentive to skip the long, messy ticket, because there's no human reviewer choosing which tickets to look at. RevelirQA ingests the platform's own SOPs and refund policies via retrieval-augmented generation into a vector database, then retrieves the relevant policy before scoring each conversation, so a deposit refund conversation gets checked against the platform's actual refund clause, not a generic industry benchmark.
| Dimension | Manual QA sampling | AutoQA (RevelirQA) |
|---|---|---|
| Coverage | Typically 1-5% of tickets | 100% of conversations |
| Basis for scoring | Reviewer judgment, general QA scorecard | Platform's own SOPs and refund policy, retrieved per ticket |
| Bias risk | Skews toward short, easy-to-review tickets | Every ticket scored identically, including long disputes |
| Audit trail | Reviewer notes, often informal | Full reasoning trace: model, documents retrieved, reasoning |
| Agent types scored | Human agents only | Human agents and AI chatbots on the same QA scorecard |
What Does a Good QA Scorecard Look Like in Practice?
A good scorecard for this category of conversation makes the timeline check, the policy citation check, and the tone check independently visible, so a coaching conversation can point to the exact criterion that failed. An agent who gets the refund amount right but delivers it coldly needs different coaching than one who gets the tone right but quotes the wrong deadline. Collapsing both into one "quality score" makes coaching guesswork.
Sentiment arc is a particularly useful QA metric here that most helpdesks don't produce natively. It tracks how a customer's tone shifts between the first and last message in a thread. A dispute ticket can close as "resolved" in the helpdesk status field while the customer's language went from neutral to frustrated over the course of the conversation, which is a retention risk a status field alone will never show. RevelirQA enriches every ticket with this kind of signal, alongside contact reason and recurring issue type, so a QA or ops team can see not just whether an agent followed policy but whether the policy itself is generating repeat friction.
[QUOTE REQUESTED: a line from a Revelir customer describing how ticket enrichment or sentiment arc surfaced a refund or dispute pattern they hadn't seen before]
Frequently Asked Questions
How long should a deposit refund conversation take to resolve?
It depends on jurisdiction and contract type. US security deposits typically fall under 14 to 30 day statutory limits, UK tenancy deposits must be returned within 10 days of agreement on the amount, and EU timelines vary by member state.
What's the difference between a QA scorecard and a customer satisfaction score?
A QA scorecard measures whether an agent followed specific policies and procedures correctly; a satisfaction score measures how the customer felt about the interaction. A conversation can score well on satisfaction while failing the QA scorecard if the agent, for example, quoted an incorrect refund timeline that happened not to matter to that particular customer.
Can AutoQA evaluate chatbot conversations as well as human agents?
Yes. RevelirQA scores AI chatbots and human agents on the same QA scorecard, which matters for property platforms running a chatbot for first-line dispute intake alongside human reps handling escalations.
Does automated QA replace the need for a compliance team on disputes?
No. AutoQA scores conversations against the platform's SOPs and flags policy misses; it doesn't replace the mediation, arbitration, or regulatory process itself. It gives compliance and CX teams visibility into where agents are deviating from policy before those deviations become formal complaints.
How does ticket sampling bias actually happen in practice?
Reviewers working through a fixed review quota tend to select tickets they can evaluate quickly, which statistically favors short, clearly resolved conversations over long, multi-message disputes, even though the latter carry more compliance and retention risk.
What compliance frameworks apply to property platforms handling refunds and disputes?
Depending on jurisdiction and business model, platforms may need to comply with frameworks like the UK FCA's Dispute Resolution Complaints Sourcebook (DISP) and Consumer Duty, the UK Digital Markets, Competition and Consumer Act, and US FTC regulations, all of which require defined timelines, transparent pricing, and fair refund handling.
Is a multilingual scorecard necessary for global or regional property platforms?
Yes, if the platform serves customers across multiple languages, because policy citation and tone need to be evaluated in the language the conversation actually happened in, not translated after the fact. RevelirQA scores conversations in English, Indonesian, Thai, and Tagalog, among others.
About Revelir AI
Revelir AI builds RevelirQA, an AI AutoQA platform that scores 100% of customer service conversations against a company's own policies and QA scorecard, rather than the 1-5% a manual QA team can typically sample. Founded in 2025 by Rasmus Chow and headquartered in Singapore, Revelir is in production at Xendit and Tiket.com, scoring thousands of conversations a week across fintech and travel, two industries where refund timelines and dispute handling carry real regulatory weight. Every RevelirQA score comes with a full reasoning trace, the model used, the policy documents retrieved, and the reasoning behind the score, giving CX and compliance teams an auditable record rather than a black-box number. The platform evaluates human agents and AI chatbots on the same QA scorecard, integrates with any helpdesk via API, and is built for high-volume, digitally-native businesses operating well beyond a single market.
If your platform handles booking disputes or deposit refunds at volume and your QA process still relies on sampling a handful of tickets a week, it's worth seeing what 100% coverage actually surfaces. Learn more at Revelir AI.
