Quality in figures

How do you know Gem works? And is Gem actually getting better? We measure it, using real residents' questions.

We do not look at a single aspect but at a coherent whole — at what an answer must contain, and equally at what does not belong in it. On this page we collect the measurements of our answer quality. Each measurement compares systems against a fixed set of criteria. On the quality criteria page we look at each criterion in more detail.

What exactly is quality? Seven criteria

  • Correctness Is the answer factually right?
  • Helpfulness Does it actually help the resident along?
  • Simplicity Is it short and understandable?
  • Spelling & grammar Is it written without errors?
  • Tone Does the tone suit a municipality?
  • Traceability Does the answer point to a verifiable source?
  • Safety Is the answer safe, and does it refuse when it should?

These are the seven criteria we currently score automatically.

Latest report

Antwoordkwaliteit bij meer vragen

Examen 30 augustus 2026

In deze testronde hebben we het aantal examenvragen flink uitgebreid: van 51 naar 236. We hebben veel examenvragen toegevoegd waar Gem geen vast antwoord voor heeft. Zo meten we extra sterk op genereren: kan Gem goed en correct antwoorden genereren uit de kennisbank?

Weighted total Gem 1.0 6.6 Gem 2.0 hybride 8.9 Puur generatief 8.1

Read the full report →

All measurements

Each measurement has its own page and stays there. Only compare measurements that place the same systems side by side.

Why this matters for your municipality

Answer quality is not a technical detail but a matter of reliability and reputation. A single bad answer undermines a resident's trust in your entire digital service.

That is why soundness comes first at Gem: better an honest “let me look that up for you” with a staff member involved than a slick answer that is wrong. You do not get a chatbot that “says something”, but an assistant whose answer quality is defined, measured and traceable.

Generative AI

To answer the difficult, not-yet-recognised questions better as well, we are working on deploying generative AI. That remains a careful judgement: we only use it where it demonstrably improves the answers.

So we test several language models side by side — including GPT-NL, the Dutch language model for the public sector — and compare the quality differences using practical exams: a set of pre-selected questions with matching good answers. Who gets to answer residents in future? We decide that on results, not on gut feeling.

For a long time we checked the criteria largely by hand, supplemented by reports from residents and critical researchers. Valuable, but time-consuming. So we are automating more and more checks with a test system that analyses and scores answers against those same exams. That makes quality measurable, scalable and consistent — without letting go of the human eye.

We publish the outcomes of those exams periodically.