Skip to content
Menu

How we measure

If we claim the assistant will not answer without a source, we have to be able to show it. This page explains exactly what we measure, on what, how — and where the limits of that measurement are.

Why measure at all

Anything can be written about a support assistant, and most claims cannot be checked. Our main promise — "no source, no answer" — happens to be measurable: you ask questions the knowledge base cannot cover, and see what it does. That is what we do, after every meaningful change, twice.

The test set

The knowledge base of a realistic sample site about domain and hosting services, plus 52 hand-written questions. The bot did not write them, and we did not generate them from the knowledge base: they are phrased the way a visitor would ask.

40

40 answerable questions

Questions the website content does cover. For each one we recorded which page the answer should come from.

12

12 unanswerable questions

Questions that cannot be covered: account-specific data ("when does MY domain expire"), or things that would require system access. This is the more important half of the set.

One example of each
Answerable: "What do domains cost, excluding VAT?" — the answer is on the pricing page.
Unanswerable: "When exactly does my domain expire?" — that would require seeing into the visitor's account. The correct answer here is to say it does not know, and bring in a human.

What we measure

Correct refusal

How many of the 12 unanswerable questions it declines. This is the most important number: failing here means confidently telling a prospective customer something untrue. The target is 100%, and below 100% the run counts as a FAILURE, not a "good result".

Live source citations

How many of the cited sources point at a page that actually exists and is reachable. A fabricated citation is as much a hallucination as a fabricated price — just harder to notice. Below 100% this is a failure too.

Answer rate

How many of the 40 answerable questions it answers at all. We also track whether it cited the EXPECTED page — a stricter measure than is strictly fair, since more than one page can give a correct answer.

Latency and cost

Median and 90th percentile per question, split between retrieval and the language model, plus the actual cost per question. Latency depends mostly on the language model provider, not on our code — which is why we quote a range rather than a single number.

How the measurement runs

  • On the LIVE system, end to end: the same retrieval, the same model call, the same answer schema a visitor gets. Not a laboratory copy.
  • With a concurrency of five, so the run does not take hours — and so the load resembles real use.
  • ALWAYS TWICE, back to back. A single run misleads: the language model is not deterministic, and provider latency fluctuates.
  • Before and after every meaningful code change. The project rule: code changes may only be compared using back-to-back runs — values measured on two different days show provider load, not our work.

The latest result

There is no completed measurement run for this internal benchmark project right now — so we deliberately show no table and no percentage here. Once a run finishes, this section updates with the real number.

What these numbers do NOT mean

  • They are NOT a promise about your site. They come from one specific website's knowledge base. Your result is decided by the quality of your content: a contradictory, outdated website will not produce a good assistant — at best one that honestly says it does not know.
  • They are NOT customer references. We measured this ourselves, with our own tool. It is not an independent audit and we do not present it as one.
  • They are NOT a guarantee that the assistant never errs. The language model can be wrong, and your knowledge base can contain mistakes. That is why every answer carries a source and every conversation is logged: we do not promise there are no errors — we promise they SURFACE.
  • Latency is NOT a service level commitment. It depends on the language model provider's load; on the same set, with unchanged code, we have measured medians between 2.2 and 4.2 seconds.

Measure your own

The same measurement can be run against your knowledge base, and you see the result on the Quality page in the panel. The refusal rate and the missing-knowledge list together show where your content stands — often the most valuable thing that introducing an assistant gives you.

Start free

FAQ · How it works

We always use the cookies required for the site to work: robi_fox_session, XSRF-TOKEN, felulet_nyelv, ss_suti_valasztas Beyond those we would like to measure how the site is used (Google Analytics) — that needs your permission. We use no advertising cookies, and we never measure on the trial pages. Detailed cookie notice