10,130 annotations from 850 lines of pipeline
Six annotation projects across 34 label classes, delivered as CVAT-importable COCO. An open-vocabulary VLM proposes boxes, SAM turns them into instance masks, and a renderer produces visually consistent plates regardless of source resolution.
2020 — 2026 · for six annotation projects
10,130
Annotations
across 53 images
1,761
Densest frame
larvae in one dish
34
Classes
across six taxonomies
850
Pipeline
lines of Python total
Pipeline
- image
-
Qwen2.5-VL
grounded -
NMS
IoU 0.5 -
SAM ViT-B
best of 3 -
COCO 1.0
export -
human
correction
One tool, four modes
The same 308-line script handles a person silhouette, an ornate cigar crest, a masked face in CCTV and a petri dish of larvae. The VLM runs text-prompted and open-vocabulary, which removes the need to train a detector per class — the cigar-crest prompt and the petri-dish prompt are the only difference between two completely different projects.
Knowing when not to use the model
The larvae project needed 10,002 instance polygons, with up to 1,761 in a single image. The VLM is called exactly once per image, to find the dish. Everything inside comes from a classical CV pass: adaptive threshold, morphological open and close, contour extraction, with the region eroded inward first so the bright glass rim does not produce a ring of hundreds of false instances.
Separating the expensive step from the tunable one
The mask-compliance project runs in two passes. An expensive pass caches every VLM box to JSON; an instant second pass applies the precision filters. That let the filter be tuned dozens of times without re-running the model, and is named in his own notes as the single biggest time saving in the project.
Precision rules that are actually specific
A face box is kept only if it sits inside a detected person region, and only if the candidate is at least 12% skin-toned in YCrCb space. Together those reject the COVID poster showing masked faces, the staff photographs, the monitors and the office printer — everything a face detector reliably mistakes for a face at 20 to 60 pixels wide.
The claim is deliberately modest
The report draws a hard line between labels produced by the pipeline and existing human ground truth merely rendered for review, and repeats the distinction on every table. The framing: this is a correction-ready starting point designed to cut human annotation time, not a substitute for it. The value is in the ratio — 10,130 instances placed automatically, of which a human nudges edges, adds misses and splits merges, rather than drawing 10,130 polygons from scratch.
From the archive
“A VLM cannot count 1,761 objects, and it does not need to.”
“The resulting 14:67 split is a finding about the room, not a labelling artefact.”
“The polygon follows the silhouette, not a convex hull.”
The plates
Thirty-eight specimen plates.
Three of the six projects are shown. The person, cigar-logo and mask-compliance sets are client data containing identifiable faces, so their images are held back — their counts still appear in the totals above. Click any plate to enlarge.
Run log