BOCCHI & MSDCT-UNet

A More Realistic and Challenging Benchmark for Local Motion Blur Detection with MSDCT-UNet

A real-captured benchmark whose textured blurred foregrounds defeat gradient shortcuts, paired with a frequency-aware encoder–decoder that reads blur in the DCT domain.

Kuan-Lin Chen1·Yuan-Kang Lee2·Cheng-Yuan Chiang1·Jian-Jiun Ding1

1Graduate Institute of Communication Engineering, National Taiwan University 2MediaTek Inc.

Input Blur annotation

Drag the divider to reveal the blur annotation  ·  arrows change image

0.850

In-domain mIoU on BOCCHI

First among 13 methods, +1.2 pp over the second-best baseline.

633

Training images

Enough to beat every other training source on cross-dataset transfer.

3

of 3 cross-dataset metrics

BOCCHI-trained models lead on mIoU, Dice and Recall alike.

Abstract

Fewer images, better transfer

Local motion blur detection requires pixel-level localization of blurred regions. Existing benchmarks let models rely on gradient shortcuts that fail to transfer. We introduce BOCCHI (Blurred Objects Captured across Cameras with Human-annotated Imagery), a real-captured benchmark of 633 pixel-annotated images whose textured blurred foregrounds against varied backgrounds defeat these shortcuts, and propose MSDCT-UNet (Multi-Scale Discrete Cosine Transform UNet), a frequency-aware encoder-decoder injecting multi-scale DCT priors through DCT Attention and FiLM. MSDCT-UNet ranks first in in-domain mIoU (0.850, +1.2 pp over the second-best baseline) and boundary localization on BOCCHI. BOCCHI-trained models outperform every other training source on cross-dataset transfer with only 633 training images: +2.2 pp mIoU over OMoBlur, +13.8 pp over ReLoBlur, and +34.9 pp over CUHKmotion.

Benchmark

The BOCCHI dataset

633 real-captured images across 5 cameras and diverse indoor/outdoor scenes, each annotated at pixel level.

633
Images
5
Cameras
0.22
Mean blur ratio
68.7
Mean sharp gradient

Two distinguishing properties

Blur-ratio distribution and gradient statistics across datasets

(a) Blur-ratio distribution across 5 datasets: CUHKmotion spans uniformly; ReLoBlur and BOCCHI concentrate at low ratios; OMoBlur occupies the moderate range; the Inference Dataset spans the intermediate regime. (b) Per-dataset distribution of mean gradient in blur vs. sharp regions. BOCCHI shows the largest PR25/μblur ratio (0.68): the sharp-region bottom-25% gradient overlaps the blur mean, defeating gradient-based shortcuts.

Data efficiency

Which training source transfers best?

Same 13 models, same held-out Inference Dataset, no fine-tuning. Only the training source changes.

BOCCHI
633 images
0.000
OMoBlur
994 images
0.000
ReLoBlur
1,200 images
0.000
CUHKmotion
200 images
0.000
Cross-dataset mIoU, averaged over 13 models  (0 – 0.60)

More training images do not buy better transfer: ReLoBlur has almost twice BOCCHI's images and lands 13.8 pp lower. The OMoBlur average aggregates 12 of 13 models (NAFNet × OMoBlur omitted).


Qualitative

Predictions, method by method

Switch methods to compare predictions on the BOCCHI validation set. Each panel shows 4 input–prediction pairs.

Currently showing GT reference

Selected success cases where MSDCT-UNet outperforms baselines. Predictions are binary masks (white = sharp, black = predicted blur).

Quantitative

Results

In-domain performance

Per-cell metrics: mIoU / BdF1 on each dataset's own validation set. Cells shaded 1st 2nd 3rd per column. *NAFNet × OMoBlur omitted (aspect-ratio width collapse).

Method BOCCHI CUHKmotion ReLoBlur OMoBlur
mIoUBdF1mIoUBdF1 mIoUBdF1mIoUBdF1
MSDU-Net0.7960.5610.6400.3300.9230.8200.7720.667
BiSeNetV20.7600.4700.5250.1660.9160.8000.7490.613
STDC10.7780.4570.5500.2600.9180.7950.7480.631
STDC20.7990.4950.5580.1970.9180.8000.7550.627
NAFNet0.7600.4520.5230.2130.9140.821N/A*N/A*
DDRNet-230.7010.3510.5200.1720.9100.7660.7030.583
Cellpose30.8380.6430.6650.4030.9270.8500.7910.675
KDSNet-R500.7830.4690.5490.2190.9180.8030.7480.627
KDSNet-R1010.7730.4460.5440.2430.9210.8120.7440.630
MSDSeg0.7140.3600.5170.2020.9090.7660.7210.607
BEVANet0.7780.5180.5480.2170.9190.7980.7210.622
ESMDL-UNet0.6770.2700.4320.1390.7880.3640.6270.552
MSDCT-UNet (Ours)0.8500.6830.6400.3690.9300.8510.7750.664

Cross-dataset generalization

mIoU / Dice / Recall on the Inference Dataset, no fine-tuning. The average row highlights BOCCHI as the strongest training source. *NAFNet × OMoBlur omitted (OMoBlur average aggregates 12 of 13).

Method Trained on
BOCCHI
Trained on
OMoBlur
Trained on
ReLoBlur
Trained on
CUHKmotion
mIoU / Dice / RecmIoU / Dice / RecmIoU / Dice / RecmIoU / Dice / Rec
MSDU-Net0.609 / 0.830 / 0.8760.616 / 0.805 / 0.8100.556 / 0.758 / 0.7360.142 / 0.307 / 0.268
BiSeNetV20.539 / 0.791 / 0.8430.526 / 0.704 / 0.6660.470 / 0.695 / 0.6510.248 / 0.445 / 0.455
STDC10.541 / 0.789 / 0.8400.496 / 0.657 / 0.5970.283 / 0.518 / 0.5380.197 / 0.310 / 0.286
STDC20.577 / 0.823 / 0.8820.530 / 0.705 / 0.6550.269 / 0.443 / 0.4370.182 / 0.288 / 0.277
NAFNet0.568 / 0.818 / 0.882N/A*0.478 / 0.730 / 0.7180.285 / 0.529 / 0.533
DDRNet-230.517 / 0.800 / 0.8770.489 / 0.673 / 0.6260.462 / 0.704 / 0.6650.292 / 0.548 / 0.490
Cellpose30.662 / 0.858 / 0.9100.628 / 0.825 / 0.8480.555 / 0.758 / 0.7440.160 / 0.324 / 0.306
KDSNet-R500.557 / 0.809 / 0.8710.528 / 0.717 / 0.6840.275 / 0.465 / 0.4650.225 / 0.424 / 0.369
KDSNet-R1010.557 / 0.794 / 0.8350.525 / 0.718 / 0.6960.274 / 0.433 / 0.4230.230 / 0.432 / 0.372
MSDSeg0.518 / 0.816 / 0.9010.547 / 0.788 / 0.8280.463 / 0.726 / 0.7110.141 / 0.128 / 0.121
BEVANet0.552 / 0.804 / 0.8610.535 / 0.746 / 0.7310.484 / 0.718 / 0.6920.271 / 0.493 / 0.434
ESMDL-UNet0.518 / 0.806 / 0.8730.458 / 0.698 / 0.6840.391 / 0.631 / 0.5930.231 / 0.397 / 0.352
MSDCT-UNet (Ours)0.610 / 0.848 / 0.9250.610 / 0.783 / 0.7810.563 / 0.815 / 0.8970.183 / 0.290 / 0.226
Average 0.563 / 0.814 / 0.875 0.541 / 0.735 / 0.717* 0.425 / 0.646 / 0.636 0.214 / 0.378 / 0.345

Per-image comparison

Each dot is one validation image (n = 63). x = mIoU of the strongest baseline on that image; y = mIoU of MSDCT-UNet. Above the diagonal we win, below the baseline wins. Highlighted dots are the 8 cases in the qualitative panel; click one to jump there.

Ours wins Baseline wins Shown above (click to jump)

Method

MSDCT-UNet

MSDCT-UNet architecture: DCT extraction plus a 4-stage encoder-decoder with NeXtBlock, FiLM, AFASPP and DCT Attention

Left: DCT feature extraction (grayscale + Sobel → multi-scale local DCTs at n ∈ {3, 7, 15, 31} → 57-channel FDCT). Middle: 4-stage encoder-decoder with NeXtBlock + FiLM fusion, an Attentive Frequency ASPP bottleneck, and stage-specific DCT Attention. Right: Deep supervision over the main and three auxiliary heads.

Key modules

DCT Attention

DCT Attention

Multi-head, position-dependent frequency weighting over 57 DCT channels with SE-gating and temperature-scaled softmax.

FiLM Fusion

FiLM Fusion

Frequency-conditioned affine modulation prevents strong spatial edges from overriding weaker frequency cues.

AFASPP

AFASPP

Attentive Frequency ASPP fuses multi-scale spatial context with global frequency evidence at the bottleneck.

Full method details, including DCT extraction pseudocode and ablations, are available in the paper.

BOCCHI  ·  MSDCT-UNet  ·  National Taiwan University