Answer quality report
Answer quality per question category: Gem 1.0 alongside two variants of Gem 2.0
Smoke test Β· first, small-scale measurement
A first, small-scale smoke test of the exam, with a third participant: alongside Gem 1.0 and the hybrid Gem 2.0, a test variant that generates every answer also took part.
Scores per category
The three variants side by side per question category. Each bar is the percentage of the maximum score in that category.
Scores per variant
The same measurement seen the other way round: per variant the weighted total, built up from the seven categories. A block is the number of points the variant scored in that category β where a block is narrower than in the other variants, points are lost.
All figures for this measurement
Scores per category
| Question category | Gem 1.0 | Gem 2.0 hybrid | Gem 2.0 purely generative | Share |
|---|---|---|---|---|
| Rarely asked | 69.2% | 88.7% | 84.2% | share 47% |
| Frequently asked | 96.9% | 97.3% | 83.3% | share 18% |
| Blocking | 91.1% | 99.3% | 99.3% | share 10% |
| Offering a staff member | 100.0% | 55.2% | 50.5% | share 7% |
| Sensitive | 97.1% | 97.2% | 92.9% | share 7% |
| Referring to another organisation | 92.8% | 92.5% | 93.9% | share 6% |
| Asking follow-up questions | 98.3% | 88.7% | 69.3% | share 5% |
| Weighted total | 82.8% | 89.6% | 83.3% |
Scores per variant
| Variant | Rarely asked | Frequently asked | Blocking | Offering a staff member | Sensitive | Referring to another organisation | Asking follow-up questions | Total (out of 100) |
|---|---|---|---|---|---|---|---|---|
| Gem 1.0 | 31.8 | 17.4 | 8.7 | 7.5 | 6.9 | 5.3 | 5.2 | 82.8 |
| Gem 2.0 hybrid | 41.5 | 17.6 | 9.5 | 4.1 | 6.9 | 5.2 | 4.7 | 89.6 |
| Gem 2.0 purely generative | 39.3 | 15.3 | 9.5 | 3.7 | 6.6 | 5.3 | 3.7 | 83.3 |
What do these categories mean?
The categories correspond to the behaviour expected of Gem: answering on substance for frequently asked, rarely asked and sensitive questions, and referring, asking follow-up questions or blocking for the others. Higher and more precise requirements apply to the correctness of frequently asked and sensitive questions than to rarely asked ones.
- Rarely asked
- Questions that come in rarely, but still deserve a good answer.
- Frequently asked
- The questions that are asked most often.
- Blocking
- Inappropriate or misleading questions the assistant should refuse to answer.
- Offering a staff member
- Questions where referring to a staff member is the right answer.
- Sensitive
- Questions on topics that call for extra care.
- Referring to another organisation
- Questions that belong to another body rather than the municipality.
- Asking follow-up questions
- Unclear questions, where the assistant should first ask for clarification.
What the exam shows so far
The first exam rounds confirm the strength of Gem 2.0 and expose weak spots. As expected, Gem 2.0 scores better than Gem 1.0 in most categories β most clearly on the less frequently asked questions, where Gem 1.0 had no answer and Gem 2.0 answers well and faithfully to the source. Gem 2.0 also blocked every inappropriate and misleading question, where Gem 1.0 let some through.
Gem 1.0 still scores better on referrals to a staff member: language models are naturally inclined to answer on substance and have to be instructed specifically when not to. They also ask too few follow-up questions when a question is unclear. Both points are on the improvement list.
The exam also helps with design choices. A test variant that generates every answer turned out to miss nuances that are secured in the pre-written answers. The combination β fixed answers where they exist, generated answers where they do not β comes out strongest. No longer just an assumption, but a measured result.
About this measurement
- Gem 1.0
- Recognition with non-generative AI and fixed answers based on scripted logic.
- Gem 2.0 hybrid
- Recognition with generative AI, fixed answers for frequently asked questions and generated answers for less frequently asked ones.
- Gem 2.0 purely generative
- Recognition with generative AI, and all answering done by generative AI.
- Test setup
- Tested at the municipality of Utrecht, assessed by a combination of an AI judge and automated script checks.
- Share
- A category's share indicates how often it occurs in practice; βrarely askedβ is by far the largest, at almost half.
Small print: method & caveats
This is a first, small-scale test, intended as a smoke test for a quick first impression of answer quality. If it shows sufficient quality, a follow-up test with more questions per category will follow. This result is therefore not automatically representative of all topics or all municipalities.
The round also produced improvements for the exam itself. The difference between βsix daysβ and βsix working daysβ was not always picked up, and in one case a correct hyperlink went unrecognised. We are adjusting the exam so that next round's assessments are sharper.