Who signs off on a government algorithm’s margin of error?

Who signs off on a government algorithm’s margin of error?

by Victor Angelier

 

 

The Dutch frameworks require organisations to determine what margin of error is acceptable, and to explicitly assign tasks and responsibilities. But they leave it to the organisation itself to decide who signs off on that particular trade-off. And in the four public descriptions I reviewed, I did not find that trade-off explicitly documented.

 

The Algorithm Register includes Public Eye, a decommissioned system with which the Municipality of Amsterdam counted people from camera images. Under the heading Performance, it states that the algorithm had to be approximately 70 percent accurate to provide relevant insights for regulating traffic, and that in practice it achieved approximately 90 percent accuracy. This is followed by a single sentence: “We infer this from the training images.” Further down, it states that quality and accuracy were periodically evaluated using those same training data, by a small number of employees who assessed whether the algorithm correctly recognised people as people.

 

The system is no longer in use. The listing remains: the Algorithm Framework still uses Public Eye today as an example of how to evaluate the accuracy of an algorithm.

 

Two things stand out. Performance established on the material from which a model learned says too little about how that same model performs on genuinely unseen images. And nowhere does it say who determined that 70 percent was good enough. The number is there. The trade-off is not, let alone the name.

 

That question extends beyond Amsterdam.

 

I am not raising this because I think Amsterdam works carelessly. I am raising it because I have been occupied with little else for months. For my master’s research, I am building models that predict each day where a new wildfire will start in Sweden, and most of that work is not in the model but in the evaluation: when are you entitled to say that the thing is worth anything? Anyone who carries that question around for six months encounters the same three questions again and again in public-sector AI dossiers.

 

One: a percentage describes a situation, not a system.

 

The Algorithm Framework lists accuracy, precision, recall, F1 and the ROC curve as ways of assessing performance. Useful measures, and they measure different things: some apply at one chosen threshold, while others summarise behaviour across all thresholds. But none of these measures tells you whether a model score of 0.20 really means that roughly one in five of those cases will be positive.

 

That difference is called calibration. A model can rank cases extremely well while simultaneously producing unreliable probability estimates. As long as you use such a score only to sort a work queue, you can live with that. Once you use it to determine who should be investigated further, you cannot — at that point, the question of what the score means is no longer a technical detail but the heart of accountability. I did not find calibration explicitly addressed in the publicly searchable text of the framework.

 

In addition, one of those measures becomes misleading when the cases being sought are rare. If one in a thousand files is genuinely problematic, a model that always says “no” effortlessly achieves 99.9 percent accuracy.

 

Two: performance does not travel.

 

In my research, I develop the models using Swedish data and then test them in the Valencian region, without retraining or adjusting them. The software remains identical. The climate, the landscape and the frequency with which fires start do not.

 

I do not yet know whether the model survives that transition — that is precisely why you test rather than assume. And that is exactly why this step must not be skipped in public-sector systems. A model that has been properly validated at one implementing organisation can arrive at another with the same code and documentation, but there it encounters a different population and different records.

 

Records, because that is where the second half of this problem lies. The Algorithm Framework puts the question sharply itself: does the data actually describe the phenomenon you want to investigate? A model that learns from previous reports, inspections or enforcement files primarily learns the pattern of those records. A neighbourhood with more recorded incidents may have more incidents, or it may be monitored more intensively. The model cannot distinguish between the two. The decision-maker relying on the output must.

 

Whoever shifts the threshold chooses how many people are wrongly selected and how many genuine cases are missed. That is not a neutral software setting.

 

Three: the threshold is policy.

 

Suppose a risk model assesses a thousand files. Lower the threshold and you find more genuine cases, but you also wrongly select more people. Raise it and the picture reverses. Which setting is the right one cannot be derived from the data — not because the statistics are inadequate, but because the question is not a statistical one.

 

Consider a model that sorts reports of loose paving slabs: a miss there mainly costs a wasted inspection visit. With a fraud-risk model, the same kind of error means that someone receives a letter they should never have received. The model does not know the difference. It can only show how many errors of each type occur at each threshold.

 

The Algorithm Framework therefore explicitly requires organisations to determine what margin of error is acceptable, together with the question of which errors are worse to make. That is exactly the right question. But a framework can require that someone answer it; it cannot provide the answer. Particularly in a chain involving supplier, procurement and implementation, it must therefore be clear who makes that trade-off. The framework also insists on this: tasks and responsibilities should be explicitly assigned, for example in a RACI matrix. But it leaves it to the organisation to decide which role receives which task. On the framework page, that measure is associated with the roles of developer and project manager; no separate governance role is listed.

 

Two things that can change on Monday

 

No new assessment framework is needed. Two habits are enough.

 

Before deployment, define what evidence is sufficient to accept the system. Not one overall score, but: which data were used to measure performance, which types of error were distinguished, what the outputs mean, and under which conditions that performance must hold. If the model later moves to another region, target group or implementing organisation, that is not merely an implementation decision but grounds for fresh evidence.

 

And turn the error trade-off into a decision with an owner. Not because the person signing off needs to be able to read an ROC curve, but because they are responsible for what happens to the people who end up in the wrong category. Record which balance between errors was accepted, on the basis of which figures, by whom, and when that choice will be reviewed again.

 

Then a listing in the Algorithm Register like this becomes useful as well. Not simply 90 percent accurate, full stop — but: measured on what, with which errors, and who decided that was good enough.

 

Compliance does not prove reliability. Reliability does not prove acceptability. Between the two sits a decision. And that decision should have a name beneath it.

 

Victor Angelier is an IT Director and technology entrepreneur with 28 years’ experience in software, infrastructure and cybersecurity. For his master’s research, he develops and evaluates machine-learning models for wildfire risk prediction, with an emphasis on validation and generalisability. He is the founder of IamVERA.ai.

Sources:

 

Politie (2025) Threat-to-life model. Algoritmeregister. Available at: https://algoritmes.overheid.nl/nl/algoritme/oorg10264/77616265/threattolife-model.

 

Rijksdienst voor Identiteitsgegevens (2024) Intelligent zoeken in de Beheervoorziening BSN. Algoritmeregister. Available at: https://algoritmes.overheid.nl/nl/algoritme/oorg10103/82973360/intelligent-zoeken-in-de-beheervoorziening-bsn.

 

Gemeente Amsterdam (2026) Lokaliseren van lantaarnpalen. Algoritmeregister. Available at: https://algoritmes.overheid.nl/nl/algoritme/gm0363/17364371/lokaliseren-van-lantaarnpalen.

 

Gemeente Amsterdam (2025) Public Eye. Algoritmeregister. Available at: https://algoritmes.overheid.nl/nl/algoritme/gm0363/38748497/public-eye.

 

Ministerie van Binnenlandse Zaken en Koninkrijksrelaties (no date) Evalueer de nauwkeurigheid van het algoritme (5-VER-02). Algoritmekader. Available at: https://algoritmes.overheid.nl/algoritmekader/voldoen-aan-wetten-en-regels/maatregelen/5-ver-02-evalueer-nauwkeurigheid/.

 

Ministerie van Binnenlandse Zaken en Koninkrijksrelaties (no date) Controleer de datakwaliteit (3-DAT-01). Algoritmekader. Available at: https://algoritmes.overheid.nl/algoritmekader/voldoen-aan-wetten-en-regels/maatregelen/3-dat-01-datakwaliteit/.

 

Ministerie van Binnenlandse Zaken en Koninkrijksrelaties (no date) Taken en verantwoordelijkheden zijn toebedeeld in de algoritmegovernance (0-ORG-10). Algoritmekader. Available at: https://algoritmes.overheid.nl/algoritmekader/voldoen-aan-wetten-en-regels/maatregelen/0-org-10-inrichten-taken-en-verantwoordelijkheden-algoritmegovernance/.

 

Ministerie van Binnenlandse Zaken en Koninkrijksrelaties (no date) MinBZK/Algoritmekader [source code repository]. GitHub. Available at: https://github.com/MinBZK/Algoritmekader.

 

Researched and drafted with AI assistance, edited and fact-checked by the author. Illustration: AI-generated.