Contours, Not Grids

a topological token for perception

Interactive companion to the paper (PDF). The token data on this page is the paper's own: the 37-region test scene, its addresses, its sentence. Everything runs in this tab.

The two-minute version. Below, the same argument, touchable.

1Squares are drawn by the sensor, not the scene

A vision transformer dices the plane into fixed 16×16 squares: 4,096 tokens for this scene, whatever is drawn. Flat segments help, but a bag of segments asserts no structure — inside, part-of and not stay statistical. The token this paper proposes is the segment plus its position in the containment tree. Toggle the two tokenizations of the same scene.

4,096 tokens

Panel 1: the paper's unit-test scene. The grid spends the same budget on empty background as on the one structure that defines meaning; the contour tokenization spends tokens only where the scene commits to a boundary — 74 tokens, 865 bytes.

2The eleven states, mechanically

Picasso deleted marks state by state and the bull survived. The classical tokenizer has exactly two deletion knobs: k colour bins (tone → flat colour) and RDP tolerance ε (faithful contour → lo-fi essence). Slide ε and watch the scene walk its own eleven states — every state a complete, valid token tree, reconstructed here from its tokens alone. No learned model appears anywhere in this loop.

THE RASTER (640² samples, ~1.2 MB)
REBUILT FROM ITS TOKENS

Panel 2: each slider stop is a run of the full classical pipeline (k-means bins → Suzuki–Abe tracing → RDP → laminar tree → tokens → reconstruction). The sentence gets shorter; the scene stays itself until, abruptly, it doesn't — that cliff is what the paper calls the Picasso curve.

3The tree and the address

Every pair of regions is disjoint or strictly nested, so the scene is a tree, and every region gets a permanent 64-bit address: eight bits per level, the path from the root. Click two regions below — first the inner, then the outer — and watch is B inside A compile to one shift and one exclusive-or. No geometry is touched.

click a region…

Panel 3: the same move a discrete global grid makes in geographic space, run through part-of space: coarse-to-fine digits, containment as prefix arithmetic, O(1) and branch-free.

4The token: four fields

A region enters the model as (shape noun, placement, address, fill): a 12-bit word from a quantized contour vocabulary, the affine that puts it back, its tree position, and its material. Click any region in Panel 3 and read its token here.

click a region in panel 3…

5The sentence: segmentation is nested NER

The tree flattens to a bracketed stream, and at that moment vision inherits the sequence toolbox, because hierarchical segmentation is nested named-entity recognition: one BIO layer per depth, where the boundary is B, the interior is I, and O is transferred to the parent — laminarity written as a tagging rule. Hover the sentence; the region lights up above.

Panel 5: the scene's actual 74-token sentence (865 bytes gzipped). [Bank of [America]] and [eye inside [head inside [body]]] break the same single-layer tagger and are repaired by the same multi-layer scheme.

6It runs

The numbers behind the panels: the authored scene flattens to 74 tokens, 865 bytes and rebuilds at IoU 0.903 through a 24-word vocabulary with every address surviving exactly; the model-free raster front end rebuilds at 0.953; and the rate–distortion program prices every intermediate state in milliseconds.

Figure: authored tree (left), rebuilt from the token sentence alone (right).

Figure: three proposers priced by the same dynamic program, at full rate (top) and at the rate floor (bottom): the classical tree dissolves into bins, the neural tree stays recognisably the place, the authored tree holds the scene in thirteen tokens.

7The field test

The token survives contact with a real sensor. On hand-labelled 30 m Landsat, a mixed tokenizer — promptable masks propose, classical tracing and the laminar tree dispose — turns three passes into a 94-token queryable archive: green in 2023 and not in 2024 is answered exactly by set subtraction, and the same quantized models run in a browser tab.

Figure: the field pipeline, zero-shot, on two 30 m Landsat sites and two 10 m European field chips.

Both live systems are public: the Earth-observation instrument (tokenize a Landsat scene in your tab, run the flip query on its bit columns) and the reasoning-space companion (the index these tokens land in).