Skip to content
MH
← all work
Ten Thousand Polygons VLM-assisted annotation at volume

10,130 annotations from 850 lines of pipeline

Six annotation projects across 34 label classes, delivered as CVAT-importable COCO. An open-vocabulary VLM proposes boxes, SAM turns them into instance masks, and a renderer produces visually consistent plates regardless of source resolution.

2020 — 2026 · for six annotation projects

Qwen2.5-VLSAM ViT-BCOCO 1.0CVATOpenCVOllama

10,130

Annotations

across 53 images

1,761

Densest frame

larvae in one dish

34

Classes

across six taxonomies

850

Pipeline

lines of Python total

Pipeline

  1. image
  2. Qwen2.5-VL
    grounded
  3. NMS
    IoU 0.5
  4. SAM ViT-B
    best of 3
  5. COCO 1.0
    export
  6. human
    correction

One tool, four modes

The same 308-line script handles a person silhouette, an ornate cigar crest, a masked face in CCTV and a petri dish of larvae. The VLM runs text-prompted and open-vocabulary, which removes the need to train a detector per class — the cigar-crest prompt and the petri-dish prompt are the only difference between two completely different projects.

Knowing when not to use the model

The larvae project needed 10,002 instance polygons, with up to 1,761 in a single image. The VLM is called exactly once per image, to find the dish. Everything inside comes from a classical CV pass: adaptive threshold, morphological open and close, contour extraction, with the region eroded inward first so the bright glass rim does not produce a ring of hundreds of false instances.

Separating the expensive step from the tunable one

The mask-compliance project runs in two passes. An expensive pass caches every VLM box to JSON; an instant second pass applies the precision filters. That let the filter be tuned dozens of times without re-running the model, and is named in his own notes as the single biggest time saving in the project.

Precision rules that are actually specific

A face box is kept only if it sits inside a detected person region, and only if the candidate is at least 12% skin-toned in YCrCb space. Together those reject the COVID poster showing masked faces, the staff photographs, the monitors and the office printer — everything a face detector reliably mistakes for a face at 20 to 60 pixels wide.

The claim is deliberately modest

The report draws a hard line between labels produced by the pipeline and existing human ground truth merely rendered for review, and repeats the distinction on every table. The framing: this is a correction-ready starting point designed to cut human annotation time, not a substitute for it. The value is in the ratio — 10,130 instances placed automatically, of which a human nudges edges, adds misses and splits merges, rather than drawing 10,130 polygons from scratch.

From the archive

“A VLM cannot count 1,761 objects, and it does not need to.”

“The resulting 14:67 split is a finding about the room, not a labelling artefact.”

“The polygon follows the silhouette, not a convex hull.”

The plates

Thirty-eight specimen plates.

Three of the six projects are shown. The person, cigar-logo and mask-compliance sets are client data containing identifiable faces, so their images are held back — their counts still appear in the totals above. Click any plate to enlarge.

Run log

Every run recorded on this project.