Skip to content
MH
← all work
Privacy Face-Blur Frame segmentation conditioned on face content

F1 49.8 → 84.7 across eight model generations

Find every picture frame on a wall that contains a human face, segment it tightly, and blur it. Not face detection — the faces are printed, small and distant, and blurring only the face leaves the photo identifiable.

2026 · for a UK property-listings platform

YOLO11m-segRF-DETR-SegSAM ViT-BDINOv2PyTorchONNX

84.7

Best F1

FrameFusion-CV, new GT

81.2

Shipped

v7 single model, imgsz 1792

~5s

Inference

down from ~30s

9,175

Annotated

images, 4 canonical classes

Pipeline

  1. 4K photo
  2. RF-DETR-Seg
    @1344
  3. v3 + v7
    cross-verify
  4. confirmation
    gate IoU 0.5
  5. frame masks
  6. blur

The problem was the metric, not the model

The inherited pipeline scored 85.6 F1 and looked healthy. Scored against frame polygons instead of face boxes it collapsed to 53.9 — it had been emitting face-shaped SAM masks all along, and the face metric could not see the difference. Every number after that point is reported against two metrics that are allowed to disagree: frame-polygon F1 for task quality, and face coverage for privacy.

Two classes beat more data

The largest single gain in the campaign was not a bigger model or a longer schedule. Splitting one "frame" class into frame-with-face and frame-without-face moved precision from 64% to 86% at equal recall, because a faceless frame stopped being background and became something the model was explicitly taught not to blur.

Five retrains that all lost

The v4 campaign tried to buy precision back. All five variants scored below the v3 baseline. The post-mortem found why: turning v3 up to confidence 0.72 lands on exactly the same operating point as the best v4, obtainable free in one line. But a confidence dial turns both ways and v4 could not reach above 78.3% recall at any threshold, while v3 operated at 79.3%. It was a strictly dominated curve, not a tunable version of the original.

Then the rejected architecture won

The project brief had explicitly ruled out DETR-family models as data-hungry and uncertain. A later survey found RF-DETR-Seg mask AP-small was 28.4 against YOLO11-XL 18.8 — best available on the exact failure mode, since 92% of remaining errors were frames under 0.3% of a 4K image. It won on the first run, at lower resolution and a sixth of the epochs. The rejection was overturned by evidence.

And the ensemble that finally worked

Every previous ensemble had failed because YOLO models make correlated mistakes. Before spending compute, a diagnostic measured whether these three actually differed: v3 and v7 together caught 48 frames RF-DETR missed, and the oracle union reached 92.4% against 86.4% solo. That was the go signal. Twenty-seven fusion variants were then swept on CPU from cached predictions; the winner keeps confident RF-DETR detections outright and recovers borderline ones only when a different architecture confirms them.

The ceiling is honest

Roughly 70% of what remains is tiny, dim frames around 20-40px, where a mask cannot reliably land at IoU 0.5 — a limit no amount of training removes. The reachable target was named as coverage, not strict mask F1: 94.7% of faces already blurred, aiming at 98%. And a ground-truth audit relabelled about 66 frames, dropping false positives from 165 to 125 with no model change at all. The artwork false positives had always been correct; the labels were wrong.

From the archive

“The old pipeline was never doing what we thought it was.”

“v4 = v3 jammed at a high threshold, with the dial broken so it cannot turn back down.”

“A wall of family photos is exactly the worst place to go silent.”

Run log

Every run recorded on this project.