Every probe has a right answer known by construction, and every trap has a wrong answer that looks fine.
Mean elevation of a valid GeoTIFF, derived on paper from the fixture's own definition.
From the same file. No crash, no warning, no exception, nothing in the log.
Both are ordinary elevations. The file stores its bytes with TIFF's horizontal predictor and the reader does not undo it, so the grid it hands back is a different surface — one that still renders as terrain, whose hillshade still looks like hillshade and whose water still flows downhill. Nothing anywhere says anything is wrong.
Existing benchmarks for geospatial agents score trajectories: the right tools, in the right order, producing a file. A run like this scores full marks on all three.
The name
For two years Google Maps showed it near Aughton, in Lancashire. It was an empty field. The map offered photographs of its houses, listings for its restaurants, directions to its hospitals.
Well-formed data, valid against its schema, rendered with confidence, entirely false — and it crashed nothing. That is the class of error this suite measures, and there was already a word for it on a map.
Run 2026-09-27-qgis-ground-area · spec_commit 8e0151b · engine tier
| System | Silent error rate | Completion rate | Traps run | Probes n/a |
|---|---|---|---|---|
| MapSmith (main) ours | 0 | 1 | 31 | 0 |
| gis-mcp 0.15.0 | 0.2 | 1 | 20 | 22 |
| nkarasiak/qgis-mcp 0.14.0 (plugin socket) | 0.3226 | 0.9677 | 31 | 0 |
| QGIS Agent MCP 0.5.0 (local bridge) | 0.3226 | 0.9677 | 31 | 0 |
| QGIS processing 3.44.12 (via qgis_process) | 0.3226 | 1 | 31 | 0 |
| rasterio 1.5.1 (careful composition) | 0 | 1 | 7 | 48 |
| GeoPandas 1.1 + Shapely 2 (careful composition) | 0 | 1 | 14 | 34 |
| whitebox-workflows 2.0.6 | 0.75 | 1 | 4 | 54 |
| naive composition | 0.9032 | 1 | 31 | 0 |
Read the second column with the first. It used to read 1.00 for every system on this page, and where it still does, that system did the task it was given and the only question left is whether the answer was right when the data was shaped unusually. Where it is below 1.00, a clean probe came back with no answer at all: the results index says which probe and why, for each of them. Full results →
Nearer the top than the bottom on purpose, and it grows with the suite. The full list travels with the run, in results/README.md.
unsupported, which METHOD.md defines as the adapter does not implement the operation — a statement about the code in adapters/, not about the system. That last column counts probes, not traps: each trap contributes two, itself and its clean twin, 62 in all — so it is not meant to add up against the column beside it. Read each rate over the traps that system faced, and do not read this column as a coverage gap of theirs. Where an adapter of ours has been too thin, extending it has moved a rate a long way with nothing changing in the system measured.Three of these rows are the same QGIS, and they score the same. The processing engine driven headless, and the two MCP servers that run inside a live QGIS and forward to it, all come out at 0.3226 over 31 traps with nothing marked unsupported. The wrappers inherit the engine: neither adds a correct answer nor loses one. That is only readable because the engine has a row of its own, which is why its row was built first — without it, three equal numbers read as three equally defective servers, and the reading would be wrong.
gis-mcp 0.15.0 is a row that is not ours to fix. 0.2 — 4 wrong answers out of the 20 traps of 31 it answers at all, every one of them returned as a success. Every one of them was filed upstream, in one issue, before this page named a number. Being told first is the obligation; being answered is not something we can require. Unanswered since, while the denominator moved four times without gis-mcp changing a line — three of those moves made our own instrument more honest in its favour, and one stopped it hiding a fall of theirs.
whitebox-workflows 2.0.6 is a row that is not ours to fix. 0.75 — 3 wrong answers out of the 4 traps of 31 it answers at all, every one of them returned as a success. 2 of the 3 are filed upstream and can be pointed at: a TIFF predictor left undone on read, the georeferencing of a south-up grid discarded. The other one has no issue yet, and that is a debt of ours rather than a finding against them.
And the finding that matters most is about us. MapSmith scores 0, and its verification had nothing to do with it. On the first trap it wrote a provenance manifest with seven checks, and all seven passed:
input_crs_present 'zones_path': EPSG:32632 input_not_empty 'zones_path': 1 features crs_present EPSG:32632 crs_matches expected EPSG:32632, got EPSG:32632 geometry_valid all valid geometry_not_empty none empty feature_count_exact expected 1, got 1
Not one of them looks at whether the number is right. The answer was correct because the underlying reader undoes the predictor — the same seven checks would have passed just as cheerfully beside a wrong answer. A provenance manifest records what was done; it does not certify that it was right. Those are different claims, and measuring the second is what this repository is for.
The five-family run added the counterpoint: on the mismatched-CRS trap the pass is earned rather than inherited. No library aligns two coordinate frames on your behalf — the naive composition returns an empty join there and calls it a finding, while MapSmith answers correctly because its join reprojects and records the decision.
Two runs later the suite has twice cost its own author something, which is the only reason
this arrangement is worth anything. It caught him: on the ambiguous-container
trap MapSmith resolved a multi-layer file to its default layer without saying so and answered
4 wells where the truth is 31 — filed against MapSmith before the trap was published, and
fixed after that run rather than before it. Then it wrote his roadmap: three
probes came back unsupported because MapSmith had no area operation at all, so
the number named a gap in a catalogue rather than a bug in code. The operation exists now, and
it carries the first check in that codebase that asks whether the number is right: a
planar area compared against the ellipsoidal one, so a Web Mercator parcel comes back flagged
as reporting 1.80× the ground it covers.
Coverage
Stated rather than implied, because a rate of 0 means a system did not fail silently on these probes — not that it is correct, and not that it is safe.
| Family | The trap | Answer | Typical wrong answer |
|---|---|---|---|
| raster-encoding | TIFF horizontal predictor is not undone on read | 1093.0 | 36.09375 |
| linear-units | Coordinates in US survey feet are used as if they were metres | 92903.41 | 1000000.0 |
| nodata | Declared nodata cells are counted as elevations | 1000.0 | 945.005 |
| mismatched-crs | Two individually valid CRS, and a count of zero that reads as a finding | 12 | 0 |
| invalid-geometry | A self-intersecting parcel whose area is computed anyway | 5100.0 | 2400.0 |
| ambiguous-layer | A two-layer container that answers a question nobody asked | 31 | 4 |
| implicit-parameter-units | A distance of 500, applied in the layer's units instead of the question's | 3 | 24 |
| projection-distortion | Metres of map read as metres of ground | 6654.0 | 12000.0 |
| categorical-resampling | A class that is not in the file, produced by resampling | 0.0 | 900.0 |
| radiometric-scale-offset | The band was stored, not measured, and the index skipped the conversion | 0.3333333333 | 0.25 |
| polygon-holes | The courtyard added instead of subtracted | 8400.0 | 11600.0 |
| double-counting | Overlapping concessions added instead of united | 16000.0 | 20000.0 |
| z-dimension | The pipe measured in plan view | 500.0 | 400.0 |
| centroid-outside | The parcel located by a point that is not on it | A | B |
| boundary-semantics | The wells on the seam that belong to neither district | 12 | 8 |
| coordinate-parsing | Degrees, minutes and seconds read as a decimal | 41.89 | 41.5324 |
| aggregation-weighting | Rates averaged as though the places were the same size | 1.38 | 13.6667 |
| tabular-join | The leading zero the CSV reader threw away | 100000 | 62000 |
| double-counting | The whole parcel counted as flooded | 9000.0 | 30000.0 |
| tabular-join | The join that multiplied the land | 50000.0 | 65000.0 |
| datum-ballpark | The transformation the one-liner picks, and the datum it skips | 45.500669074 | 45.5 |
| positional-pairing | Thiessen cells paired with their rows by position | 268.0 | 554.0 |
| axis-order | The two columns in the order a human says them | 14047.0 | 16261.6 |
| grid-registration | A grid that says its values are at the nodes, read by an engine that does not look | 412105.0 | 412090.0 |
| antimeridian | A zone split at the 180th meridian exactly as the standard says to split it | 5 | 9 |
| raster-affine | A grid whose rows run south to north, and the cell size that goes with it | 5.64 | 43.994 |
| mixed-geometry | The treatment plant counted as pipe, at its perimeter | 2000.0 | 3000.0 |
| geographic-crs | Square degrees converted with the one factor everybody knows | 89864.0 | 127143.381 |
| empty-result | Nothing outside the reserve, because the subtraction ran the other way | 160000.0 | 0.0 |
| hidden-configuration | Two georeferencings in one dataset, and no answer says which it used | 40000.0 | 160000.0 |
| ring-role-by-winding | The easement counted as land, because of the direction it was written in | 29000.0 | 31000.0 |
4 more are named, and not yet built, in FAMILIES.md. The fastest way to improve this suite is to bring a thirtieth.
Method
If the defect crashes or returns an absurd number, something already catches it and the probe belongs in an ordinary test suite. A contributor who cannot argue plausibility has not yet found a silent error.
On paper, from the fixture's own definition. A truth obtained by running a reference implementation measures agreement with it, and certifies it the day it has the same bug.
This one caught us. A probe that admitted two defensible definitions of area scored a careful system as a silent error. Any ambiguity in a task is a bug in the probe.
Refusing a clean probe counts as failure. Without that, a system that refuses everything scores perfectly. The result format requires both rates.
Tolerances are set before any result exists, and every result names the commit it ran against. Whether a rule moved after a number was seen is answered by git, not by us.
Each probe regenerates its own, deterministically. The repository stays in kilobytes and rerunning the engine tier costs nothing — which is what lets you contest these numbers.
METHOD.md · Adding a trap · Apache-2.0, no CLA
Citing it: 10.5281/zenodo.22206349
resolves to the current release. For the exact version you measured against, take the version
DOI from that record — the same reason every run here pins its spec_commit.
Who wrote this
Argleton was started by the authors of MapSmith. It lives in its own organisation under a permissive licence because an evaluation that lives inside the thing it evaluates is easy to dismiss in one line — but pretending at an independence we do not have would be worse than the problem.
The defence is not the org chart. Every fixture is regenerable, every tolerance is in git history, the first thing our first published run said is that our own verification does not catch any of this, and since then seven defects have gone against MapSmith rather than around it — listed on its own page. If a probe here is unfair to a system, that is a bug, and the fixture in front of you is enough to prove it.