{"id":535,"date":"2026-09-28T04:41:03","date_gmt":"2026-09-28T02:41:03","guid":{"rendered":"https:\/\/robifox.com\/blog\/chatbot-accuracy\/"},"modified":"2026-10-01T12:51:48","modified_gmt":"2026-10-01T10:51:48","slug":"chatbot-accuracy","status":"publish","type":"post","link":"https:\/\/robifox.com\/blog\/chatbot-accuracy\/","title":{"rendered":"How to Measure Chatbot Accuracy Before and After Launch"},"content":{"rendered":"<p>Chatbot accuracy is best evaluated against specific questions and approved answers. A response can sound confident, link to a real page, and still get a product variant, deadline, or exception wrong. To find those failures, test what the assistant says against what the business actually knows.<\/p>\n<p>For a website support chatbot, a useful evaluation covers three outcomes: answering supported questions correctly, stopping when information is missing, and directing the visitor to the right next step. This guide explains how to build that evaluation without turning it into a large research project.<\/p>\n<h2>Write the expected answer before running the test<\/h2>\n<p>Collect questions from support conversations, website searches, and sales calls. Remove personal details and keep the wording customers use. Include short questions, follow-ups, and questions with an assumption hidden inside them.<\/p>\n<p>For each question, record the source that should support the answer and the essential facts it must include. Describe acceptable variation in wording. You are checking meaning, not whether the assistant repeats a paragraph word for word.<\/p>\n<p><strong>Fictional example:<\/strong> a service page says installation is included within one named city. For \u201cIs installation included at my office outside the city?\u201d, the expected response should preserve the geographic condition and explain how to confirm an exception. A plain \u201cYes, installation is included\u201d fails even though part of the sentence appears in the source.<\/p>\n<h2>Include cases that require different behaviour<\/h2>\n<figure class=\"wp-block-table\">\n<table>\n<caption>A compact chatbot accuracy test set<\/caption>\n<thead>\n<tr>\n<th scope=\"col\">Case<\/th>\n<th scope=\"col\">Example<\/th>\n<th scope=\"col\">Successful behaviour<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Direct factual question<\/td>\n<td>Which materials is this product made from?<\/td>\n<td>Uses the correct product&#8217;s current specification.<\/td>\n<\/tr>\n<tr>\n<td>Conditional answer<\/td>\n<td>Can I return an opened item?<\/td>\n<td>Preserves the relevant condition and documented process.<\/td>\n<\/tr>\n<tr>\n<td>Missing information<\/td>\n<td>Will you offer a discount next month?<\/td>\n<td>Does not invent an unpublished promotion.<\/td>\n<\/tr>\n<tr>\n<td>Wrong premise<\/td>\n<td>Why is express delivery free?<\/td>\n<td>Checks the premise instead of repeating it as a fact.<\/td>\n<\/tr>\n<tr>\n<td>Follow-up<\/td>\n<td>Does that apply in another country?<\/td>\n<td>Retains context without extending a rule beyond its scope.<\/td>\n<\/tr>\n<tr>\n<td>Request for action<\/td>\n<td>Cancel my order now.<\/td>\n<td>Explains the supported process without pretending an action was completed.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n<p>These are test designs, not results. Replace the examples with questions relevant to your own business. If you support multiple languages, include representative questions in those languages and have someone able to assess them review the answers.<\/p>\n<h2>Score the answer and the evidence separately<\/h2>\n<p>A simple review sheet can use four checks: factual correctness, completeness of important conditions, source support, and appropriate next step. Mark each as pass, partial, or fail, with a short explanation. This makes the reason for a failure visible.<\/p>\n<p>A working citation is only one check. Open the linked page and find the passage that supports the answer. A source can exist yet be outdated, cover a different product, or support only half the reply. The <a href=\"https:\/\/robifox.com\/blog\/ai-chatbot-citations\/\">citation-audit guide<\/a> goes deeper into those cases.<\/p>\n<p>Keep answerable and unanswerable questions in separate groups when reporting results. Combining them into a single percentage can hide an assistant that answers common questions well but guesses whenever the source is missing. State the number of cases tested and what those cases covered.<\/p>\n<h2>Use failures to choose the right fix<\/h2>\n<p>When an answer fails, inspect the source first. If the website itself is wrong or ambiguous, correct it. If the source is right but the extracted knowledge is incomplete, review that knowledge. If the right information exists but was not found, investigate retrieval and wording before adding another copy of the same fact.<\/p>\n<p>Some failures concern the response rather than the source: an unsupported promise, an omitted exception, or a claim to have completed an action. Keep the failing question in your regression set so the same problem is checked after changes.<\/p>\n<p>Retest related cases as well as the one that failed. Fixing a domestic shipping answer should not accidentally make that answer appear for every destination.<\/p>\n<h2>Treat a good test result as a bounded result<\/h2>\n<p>Passing your test set shows performance on those questions under those conditions. It does not prove that every future reply will be correct. Review real conversations after launch, especially where a wrong answer could create a costly commitment.<\/p>\n<p>RobiFox provides tools for testing answers and inspecting their sources. Its <a href=\"https:\/\/robifox.com\/how-we-measure\">measurement page<\/a> explains what its own checks cover and where their limits are. Use that alongside your own business-specific evaluation.<\/p>\n<p>Keep a record of the test date, source version, cases, failures, and corrections. When a policy changes, update the expected answers before running the tests again. For the ongoing ownership process, see <a href=\"https:\/\/robifox.com\/blog\/ai-knowledge-base-governance-small-teams\/\">knowledge base governance for small teams<\/a>.<\/p>\n<p><small>Cover image: AI-generated editorial illustration, not a product screenshot.<\/small><\/p>\n<p><!-- robifox-help-example-v2 --><\/p>\n<h2>A real RobiFox help test: 52 questions<\/h2>\n<p>On 1 October 2026 we ran 40 answerable and 12 unanswerable questions against our own Hungarian help knowledge base, using openai\/gpt-5.6-luna. The revised v2 set returned an answered flag for 37\/40 answerable questions, a refusal flag for 11\/12 negative controls, and the exact expected source for 35\/40 source checks. The median recorded response time was 4.2 seconds; 38 credits were charged and no execution errors were recorded.<\/p>\n<p>These are automatic status and citation checks, not a 92.5% factual-accuracy score. One response explicitly said it could not verify which user a conversation belonged to, while its answered flag was true. Reading the output therefore matters. Six ambiguous questions were revised after the initial run, so the two sets do not prove that a code change improved the assistant.<\/p>\n<p><a href=\"https:\/\/robifox.com\/how-we-measure\">Read the dated methodology and original responses<\/a> or <a href=\"https:\/\/robifox.com\/nyilvanos\/meres\/2026-10-01-robifox-help-v2.json\">download the 52-question result (JSON)<\/a>. This measures our own Hungarian documentation, not English answer quality or customer results.<\/p>\n<p><!-- \/robifox-help-example-v2 --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Evaluate chatbot answers for factual correctness, source support, preserved conditions and appropriate refusals using a repeatable test set.<\/p>\n","protected":false},"author":1,"featured_media":581,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","rank_math_focus_keyword":"chatbot accuracy","rank_math_title":"How to Measure Chatbot Accuracy Before and After Launch","rank_math_description":"Measure chatbot accuracy with real questions, source checks and required refusals. Build a repeatable test set and diagnose failures after content changes.","rank_math_canonical_url":"","rank_math_robots":[],"rank_math_facebook_title":"How to Measure Chatbot Accuracy Before and After Launch","rank_math_facebook_description":"Measure chatbot accuracy with real questions, source checks and required refusals. Build a repeatable test set and diagnose failures after content changes.","rank_math_facebook_image":"https:\/\/robifox.com\/blog\/wp-content\/uploads\/2026\/09\/chatbot-accuracy-editorial.webp","rank_math_facebook_image_id":"581","rank_math_twitter_title":"How to Measure Chatbot Accuracy Before and After Launch","rank_math_twitter_description":"Measure chatbot accuracy with real questions, source checks and required refusals. Build a repeatable test set and diagnose failures after content changes.","rank_math_twitter_image":"https:\/\/robifox.com\/blog\/wp-content\/uploads\/2026\/09\/chatbot-accuracy-editorial.webp","rank_math_twitter_image_id":"581","rank_math_twitter_card_type":"summary_large_image","rank_math_advanced_robots":[],"rank_math_breadcrumb_title":""},"categories":[26],"tags":[],"class_list":["post-535","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-quality-and-testing"],"_links":{"self":[{"href":"https:\/\/robifox.com\/blog\/wp-json\/wp\/v2\/posts\/535","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/robifox.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/robifox.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/robifox.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/robifox.com\/blog\/wp-json\/wp\/v2\/comments?post=535"}],"version-history":[{"count":2,"href":"https:\/\/robifox.com\/blog\/wp-json\/wp\/v2\/posts\/535\/revisions"}],"predecessor-version":[{"id":585,"href":"https:\/\/robifox.com\/blog\/wp-json\/wp\/v2\/posts\/535\/revisions\/585"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/robifox.com\/blog\/wp-json\/wp\/v2\/media\/581"}],"wp:attachment":[{"href":"https:\/\/robifox.com\/blog\/wp-json\/wp\/v2\/media?parent=535"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/robifox.com\/blog\/wp-json\/wp\/v2\/categories?post=535"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/robifox.com\/blog\/wp-json\/wp\/v2\/tags?post=535"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}