← Back to Quality

Answer quality report

Answer quality per question category: Gem 1.0 alongside two variants of Gem 2.0

Smoke test Β· first, small-scale measurement

A first, small-scale smoke test of the exam, with a third participant: alongside Gem 1.0 and the hybrid Gem 2.0, a test variant that generates every answer also took part.

Scores per category

The three variants side by side per question category. Each bar is the percentage of the maximum score in that category.

Scores per variant

The same measurement seen the other way round: per variant the weighted total, built up from the seven categories. A block is the number of points the variant scored in that category β€” where a block is narrower than in the other variants, points are lost.

All figures for this measurement

Scores per category

The share is the portion of the 100 points this category can contribute.
Question category Gem 1.0 Gem 2.0 hybrid Gem 2.0 purely generative Share
Rarely asked 69.2% 88.7% 84.2% share 47%
Frequently asked 96.9% 97.3% 83.3% share 18%
Blocking 91.1% 99.3% 99.3% share 10%
Offering a staff member 100.0% 55.2% 50.5% share 7%
Sensitive 97.1% 97.2% 92.9% share 7%
Referring to another organisation 92.8% 92.5% 93.9% share 6%
Asking follow-up questions 98.3% 88.7% 69.3% share 5%
Weighted total 82.8% 89.6% 83.3%

Scores per variant

Variant Rarely asked Frequently asked Blocking Offering a staff member Sensitive Referring to another organisation Asking follow-up questions Total (out of 100)
Gem 1.0 31.8 17.4 8.7 7.5 6.9 5.3 5.2 82.8
Gem 2.0 hybrid 41.5 17.6 9.5 4.1 6.9 5.2 4.7 89.6
Gem 2.0 purely generative 39.3 15.3 9.5 3.7 6.6 5.3 3.7 83.3
What do these categories mean?

The categories correspond to the behaviour expected of Gem: answering on substance for frequently asked, rarely asked and sensitive questions, and referring, asking follow-up questions or blocking for the others. Higher and more precise requirements apply to the correctness of frequently asked and sensitive questions than to rarely asked ones.

Rarely asked
Questions that come in rarely, but still deserve a good answer.
Frequently asked
The questions that are asked most often.
Blocking
Inappropriate or misleading questions the assistant should refuse to answer.
Offering a staff member
Questions where referring to a staff member is the right answer.
Sensitive
Questions on topics that call for extra care.
Referring to another organisation
Questions that belong to another body rather than the municipality.
Asking follow-up questions
Unclear questions, where the assistant should first ask for clarification.

What the exam shows so far

The first exam rounds confirm the strength of Gem 2.0 and expose weak spots. As expected, Gem 2.0 scores better than Gem 1.0 in most categories β€” most clearly on the less frequently asked questions, where Gem 1.0 had no answer and Gem 2.0 answers well and faithfully to the source. Gem 2.0 also blocked every inappropriate and misleading question, where Gem 1.0 let some through.

Gem 1.0 still scores better on referrals to a staff member: language models are naturally inclined to answer on substance and have to be instructed specifically when not to. They also ask too few follow-up questions when a question is unclear. Both points are on the improvement list.

The exam also helps with design choices. A test variant that generates every answer turned out to miss nuances that are secured in the pre-written answers. The combination β€” fixed answers where they exist, generated answers where they do not β€” comes out strongest. No longer just an assumption, but a measured result.

About this measurement

Gem 1.0
Recognition with non-generative AI and fixed answers based on scripted logic.
Gem 2.0 hybrid
Recognition with generative AI, fixed answers for frequently asked questions and generated answers for less frequently asked ones.
Gem 2.0 purely generative
Recognition with generative AI, and all answering done by generative AI.
Test setup
Tested at the municipality of Utrecht, assessed by a combination of an AI judge and automated script checks.
Share
A category's share indicates how often it occurs in practice; β€œrarely asked” is by far the largest, at almost half.

Small print: method & caveats

This is a first, small-scale test, intended as a smoke test for a quick first impression of answer quality. If it shows sufficient quality, a follow-up test with more questions per category will follow. This result is therefore not automatically representative of all topics or all municipalities.

The round also produced improvements for the exam itself. The difference between β€œsix days” and β€œsix working days” was not always picked up, and in one case a correct hyperlink went unrecognised. We are adjusting the exam so that next round's assessments are sharper.

All measurements

Each measurement has its own page and stays there. Only compare measurements that place the same systems side by side.