Argleton

A correctness suite for geospatial systems

Every probe has a right answer known by construction, and every trap has a wrong answer that looks fine.

The answer 1093.0

Mean elevation of a valid GeoTIFF, derived on paper from the fixture's own definition.

What one widely used library returns 36.09375

From the same file. No crash, no warning, no exception, nothing in the log.

Both are ordinary elevations. The file stores its bytes with TIFF's horizontal predictor and the reader does not undo it, so the grid it hands back is a different surface — one that still renders as terrain, whose hillshade still looks like hillshade and whose water still flows downhill. Nothing anywhere says anything is wrong.

Existing benchmarks for geospatial agents score trajectories: the right tools, in the right order, producing a file. A run like this scores full marks on all three.

The name

Argleton was a village that was not there

For two years Google Maps showed it near Aughton, in Lancashire. It was an empty field. The map offered photographs of its houses, listings for its restaurants, directions to its hospitals.

Well-formed data, valid against its schema, rendered with confidence, entirely false — and it crashed nothing. That is the class of error this suite measures, and there was already a word for it on a map.

Run 2026-09-27-qgis-ground-area · spec_commit 8e0151b · engine tier

9 systems, 31 traps, 29 families

System Silent error rate Completion rate Traps run Probes n/a
MapSmith (main) ours01310
gis-mcp 0.15.00.212022
nkarasiak/qgis-mcp 0.14.0 (plugin socket)0.32260.9677310
QGIS Agent MCP 0.5.0 (local bridge)0.32260.9677310
QGIS processing 3.44.12 (via qgis_process)0.32261310
rasterio 1.5.1 (careful composition)01748
GeoPandas 1.1 + Shapely 2 (careful composition)011434
whitebox-workflows 2.0.60.751454
naive composition0.90321310

Read the second column with the first. It used to read 1.00 for every system on this page, and where it still does, that system did the task it was given and the only question left is whether the answer was right when the data was shaped unusually. Where it is below 1.00, a clean probe came back with no answer at all: the results index says which probe and why, for each of them. Full results →

What these numbers do not say

Nearer the top than the bottom on purpose, and it grows with the suite. The full list travels with the run, in results/README.md.

Three of these rows are the same QGIS, and they score the same. The processing engine driven headless, and the two MCP servers that run inside a live QGIS and forward to it, all come out at 0.3226 over 31 traps with nothing marked unsupported. The wrappers inherit the engine: neither adds a correct answer nor loses one. That is only readable because the engine has a row of its own, which is why its row was built first — without it, three equal numbers read as three equally defective servers, and the reading would be wrong.

gis-mcp 0.15.0 is a row that is not ours to fix. 0.2 — 4 wrong answers out of the 20 traps of 31 it answers at all, every one of them returned as a success. Every one of them was filed upstream, in one issue, before this page named a number. Being told first is the obligation; being answered is not something we can require. Unanswered since, while the denominator moved four times without gis-mcp changing a line — three of those moves made our own instrument more honest in its favour, and one stopped it hiding a fall of theirs.

whitebox-workflows 2.0.6 is a row that is not ours to fix. 0.75 — 3 wrong answers out of the 4 traps of 31 it answers at all, every one of them returned as a success. 2 of the 3 are filed upstream and can be pointed at: a TIFF predictor left undone on read, the georeferencing of a south-up grid discarded. The other one has no issue yet, and that is a debt of ours rather than a finding against them.

And the finding that matters most is about us. MapSmith scores 0, and its verification had nothing to do with it. On the first trap it wrote a provenance manifest with seven checks, and all seven passed:

input_crs_present    'zones_path': EPSG:32632
input_not_empty      'zones_path': 1 features
crs_present          EPSG:32632
crs_matches          expected EPSG:32632, got EPSG:32632
geometry_valid       all valid
geometry_not_empty   none empty
feature_count_exact  expected 1, got 1

Not one of them looks at whether the number is right. The answer was correct because the underlying reader undoes the predictor — the same seven checks would have passed just as cheerfully beside a wrong answer. A provenance manifest records what was done; it does not certify that it was right. Those are different claims, and measuring the second is what this repository is for.

The five-family run added the counterpoint: on the mismatched-CRS trap the pass is earned rather than inherited. No library aligns two coordinate frames on your behalf — the naive composition returns an empty join there and calls it a finding, while MapSmith answers correctly because its join reprojects and records the decision.

Two runs later the suite has twice cost its own author something, which is the only reason this arrangement is worth anything. It caught him: on the ambiguous-container trap MapSmith resolved a multi-layer file to its default layer without saying so and answered 4 wells where the truth is 31 — filed against MapSmith before the trap was published, and fixed after that run rather than before it. Then it wrote his roadmap: three probes came back unsupported because MapSmith had no area operation at all, so the number named a gap in a catalogue rather than a bug in code. The operation exists now, and it carries the first check in that codebase that asks whether the number is right: a planar area compared against the ellipsoidal one, so a Web Mercator parcel comes back flagged as reporting 1.80× the ground it covers.

Coverage

29 families of 33

Stated rather than implied, because a rate of 0 means a system did not fail silently on these probes — not that it is correct, and not that it is safe.

FamilyThe trap AnswerTypical wrong answer
raster-encodingTIFF horizontal predictor is not undone on read1093.036.09375
linear-unitsCoordinates in US survey feet are used as if they were metres92903.411000000.0
nodataDeclared nodata cells are counted as elevations1000.0945.005
mismatched-crsTwo individually valid CRS, and a count of zero that reads as a finding120
invalid-geometryA self-intersecting parcel whose area is computed anyway5100.02400.0
ambiguous-layerA two-layer container that answers a question nobody asked314
implicit-parameter-unitsA distance of 500, applied in the layer's units instead of the question's324
projection-distortionMetres of map read as metres of ground6654.012000.0
categorical-resamplingA class that is not in the file, produced by resampling0.0900.0
radiometric-scale-offsetThe band was stored, not measured, and the index skipped the conversion0.33333333330.25
polygon-holesThe courtyard added instead of subtracted8400.011600.0
double-countingOverlapping concessions added instead of united16000.020000.0
z-dimensionThe pipe measured in plan view500.0400.0
centroid-outsideThe parcel located by a point that is not on itAB
boundary-semanticsThe wells on the seam that belong to neither district128
coordinate-parsingDegrees, minutes and seconds read as a decimal41.8941.5324
aggregation-weightingRates averaged as though the places were the same size1.3813.6667
tabular-joinThe leading zero the CSV reader threw away10000062000
double-countingThe whole parcel counted as flooded9000.030000.0
tabular-joinThe join that multiplied the land50000.065000.0
datum-ballparkThe transformation the one-liner picks, and the datum it skips45.50066907445.5
positional-pairingThiessen cells paired with their rows by position268.0554.0
axis-orderThe two columns in the order a human says them14047.016261.6
grid-registrationA grid that says its values are at the nodes, read by an engine that does not look412105.0412090.0
antimeridianA zone split at the 180th meridian exactly as the standard says to split it59
raster-affineA grid whose rows run south to north, and the cell size that goes with it5.6443.994
mixed-geometryThe treatment plant counted as pipe, at its perimeter2000.03000.0
geographic-crsSquare degrees converted with the one factor everybody knows89864.0127143.381
empty-resultNothing outside the reserve, because the subtraction ran the other way160000.00.0
hidden-configurationTwo georeferencings in one dataset, and no answer says which it used40000.0160000.0
ring-role-by-windingThe easement counted as land, because of the direction it was written in29000.031000.0

4 more are named, and not yet built, in FAMILIES.md. The fastest way to improve this suite is to bring a thirtieth.

Method

How a probe earns its place

The wrong answer must be plausible

If the defect crashes or returns an absurd number, something already catches it and the probe belongs in an ordinary test suite. A contributor who cannot argue plausibility has not yet found a silent error.

The truth is derived, not measured

On paper, from the fixture's own definition. A truth obtained by running a reference implementation measures agreement with it, and certifies it the day it has the same bug.

The task has one correct answer

This one caught us. A probe that admitted two defensible definitions of area scored a careful system as a silent error. Any ambiguity in a task is a bug in the probe.

Two numbers, never one

Refusing a clean probe counts as failure. Without that, a system that refuses everything scores perfectly. The result format requires both rates.

Pre-registration you can diff

Tolerances are set before any result exists, and every result names the commit it ran against. Whether a rule moved after a number was seen is answered by git, not by us.

Fixtures are built, not shipped

Each probe regenerates its own, deterministically. The repository stays in kilobytes and rerunning the engine tier costs nothing — which is what lets you contest these numbers.

METHOD.md · Adding a trap · Apache-2.0, no CLA

Citing it: 10.5281/zenodo.22206349 resolves to the current release. For the exact version you measured against, take the version DOI from that record — the same reason every run here pins its spec_commit.

Who wrote this

The authors of one of the systems it measures

Argleton was started by the authors of MapSmith. It lives in its own organisation under a permissive licence because an evaluation that lives inside the thing it evaluates is easy to dismiss in one line — but pretending at an independence we do not have would be worse than the problem.

The defence is not the org chart. Every fixture is regenerable, every tolerance is in git history, the first thing our first published run said is that our own verification does not catch any of this, and since then seven defects have gone against MapSmith rather than around it — listed on its own page. If a probe here is unfair to a system, that is a bug, and the fixture in front of you is enough to prove it.