Skip to yearly menu bar Skip to main content


Oral Session

Oral Session 4D: Visual Segmentation

Sat 6 Jun 1 p.m. PDT — 2:15 p.m. PDT
Abstract:
Chat is not available.

Sat 6 June 13:00 - 13:12 PDT

INSID3: Training-Free In-Context Segmentation with DINOv3

Claudia Cuttano ⋅ Gabriele Trivigno ⋅ Christoph Reich ⋅ Daniel Cremers ⋅ Carlo Masone ⋅ Stefan Roth

In-context segmentation (ICS) aims to segment arbitrary concepts, objects, parts, or personalized instances given a few annotated visual examples. Existing work relies on (i) fine-tuning vision foundation models (VFMs), which improves in-domain results but limits generalization, or (ii) combines multiple frozen VFMs, which preserves generalization but yields architectural complexity and fixed segmentation granularities. We revisit ICS from a minimalist perspective and ask: Can a single self-supervised backbone support both semantic matching and segmentation, without any supervision or auxiliary models? We show that scaled-up dense self-supervised features from DINOv3 exhibit strong spatial structure and semantic correspondence. We introduce INSID3, a training-free approach that segments concept at varying granularities only from frozen DINOv3 features, given an in-context example. INSID3 achieves state-of-the-art results across one-shot semantic, part, and personalized segmentation, outperforming previous work by +6.1 % mIoU, while using 3x fewer parameters and without any mask or category-level supervision.

Sat 6 June 13:12 - 13:25 PDT

MARCO: Navigating the Unseen Space of Semantic Correspondence

Claudia Cuttano ⋅ Gabriele Trivigno ⋅ Carlo Masone ⋅ Stefan Roth

Recent advances in semantic correspondence rely on dual-encoder architectures, combining DINOv2 with diffusion backbones. While accurate, these billion-parameter models generalize poorly beyond training keypoints, revealing a gap between benchmark performance and real-world usability, where queried points rarely match those seen during training.Building upon DINOv2, we introduce MARCO, a unified model for generalizable correspondence driven by a novel training framework that enhances both fine-grained localization and semantic generalization. By coupling a coarse-to-fine objective that refines spatial precision with a self-distillation framework, which extends sparse supervision beyond annotated regions, our approach transforms a handful of keypoints into dense, semantically coherent correspondences.MARCO sets a new state of the art on SPair-71k, AP-10K and PF-PASCAL, with gains that amplify at fine-grained localization thresholds (+10.3 PCK@0.01), strongest generalization to unseen keypoints (+3.8, SPair-U) and categories (+5.6, MP-100), while remaining 3× smaller and 10× faster than diffusion-based approaches.

Sat 6 June 13:25 - 13:37 PDT

PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation

Minjae Lee ⋅ Sungwoo Hur ⋅ Soojin Hwang ⋅ Won Hwa Kim

Visual Foundation Models (VFMs) such as the Segment Anything Model (SAM) have significantly advanced broad use of image segmentation. However, SAM and its variants necessitate substantial manual effort for prompt generation and additional training for specific applications. Recent approaches address these limitations by integrating SAM into in-context (one/few shot) segmentation, enabling auto-prompting through semantic alignment between query and support images. Despite these efforts, they still generate sub-optimal prompts that degrade segmentation quality due to visual inconsistencies between support and query images. To tackle this limitation, we introduce PR-MaGIC (Prompt Refinement via Mask Decoder Gradient Flow for In-Context Segmentation), a training-free test-time framework that refines prompts via gradient flow derived from SAM’s mask decoder. PR-MaGIC seamlessly integrates into in-context segmentation frameworks, being theoretically grounded yet practically stabilized through a simple top-1 selection strategy that ensures robust performance across samples.Extensive evaluations demonstrate that PR-MaGIC consistently improves segmentation quality across various benchmarks, effectively mitigating inadequate prompts without requiring additional training or architectural modifications.

Sat 6 June 13:37 - 13:50 PDT

R^2-Seg: Training-Free OOD Medical Tumor Segmentation via Anatomical Reasoning and Statistical Rejection

Shuaike Shen ⋅ Ke Liu ⋅ Jiaqing Xie ⋅ Shangde Gao ⋅ Chunhua Shen ⋅ Ge Liu ⋅ Mireia Crispin-Ortuzar ⋅ Shangqi Gao

Foundation models for medical image segmentation struggle under out-of-distribution (OOD) shifts, often producing fragmented false positives on OOD tumors. We introduce **R$^2$-Seg**, a **training-free** framework for robust OOD tumor segmentation that operates via a two-stage **Reason-and-Reject** process. First, the **Reason** step employs an LLM-guided anatomical reasoning planner to localize organ anchors and generate multi-scale ROIs. Second, the **Reject** step applies two-sample statistical testing to candidates generated by a frozen foundation model (BiomedParse) within these ROIs. This statistical rejection filter retains only candidates significantly different from normal tissue, effectively suppressing false positives. Our framework requires no parameter updates, making it compatible with zero-update test-time augmentation and avoiding catastrophic forgetting. On multi-center and multi-modal tumor segmentation benchmarks, **R$^2$-Seg** substantially improves Dice, specificity, and sensitivity over strong baselines and the original foundation models.

Sat 6 June 13:50 - 14:02 PDT

The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identification

Dante Wasmuht ⋅ Otto Brookes ⋅ Maximilian Schall ⋅ Pablo Palencia ⋅ Christopher Beirne ⋅ Tilo Burghardt ⋅ Majid Mirmehdi ⋅ Hjalmar Kühl ⋅ Mimi Arandjelovic ⋅ Sam Pottie ⋅ Peter Bermant ⋅ Brandon Asheim ⋅ Yi Jin Toh ⋅ Adam Elzinga ⋅ Jason Allan Holmberg ⋅ Andrew Whitworth ⋅ Eleanor Flatt ⋅ Laura Gustafson ⋅ Chaitanya Ryali ⋅ Yuan-Ting Hu ⋅ Baishan Guo ⋅ Andrew Westbury ⋅ Kate Saenko ⋅ Dídac Surís

Automated video analysis is critical for wildlife conservation. A foundational task in this domain is multi-animal tracking (MAT), which underpins applications such as individual re-identification and behavior recognition. However, existing datasets are limited in scale, constrained to a few species, or lack sufficient temporal and geographical diversity -- leaving no suitable benchmark for training general-purpose MAT models applicable across wild animal populations. To address this, we introduce SA-FARI, the largest open-source MAT dataset for wild animals. It comprises 11,609 camera trap videos collected over approximately 10 years (2014-2024) from 741 locations across 4 continents, spanning 99 species categories. Each video is exhaustively annotated culminating in $\sim$46 hours of densely annotated footage containing 16,224 masklet identities and 942,702 individual bounding boxes, segmentation masks, and species labels. Alongside the task-specific annotations, we publish anonymized camera trap locations for each video. Finally, we present comprehensive benchmarks on SA-FARI using state-of-the-art vision-language models for detection and tracking, including SAM 3, evaluated with both species-specific and generic animal prompts. We also compare against vision-only methods developed specifically for wildlife analysis. SA-FARI is the first large-scale dataset to combine high species diversity, multi-region coverage, and high-quality spatio-temporal annotations, offering a new foundation for advancing generalizable multi-animal tracking in the wild. The dataset is available at [ANONYMIZED]

Sat 6 June 14:02 - 14:15 PDT

VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation

Yulu Gao ⋅ Bohao Zhang ⋅ Zongheng Tang ⋅ Jitong Liao ⋅ wenjun wu ⋅ Si Liu

Instance-level object segmentation across disparate egocentric and exocentric views is a fundamental challenge in visual understanding, critical for applications in embodied AI and remote collaboration. This task is exceptionally difficult due to severe changes in scale, perspective, and occlusion, which destabilize direct pixel-level matching. While recent geometry-aware models like VGGT provide a strong foundation for feature alignment, we find they often fail at dense prediction tasks due to significant pixel-level projection drift, even when their internal object-level attention remains consistent. To bridge this gap, we introduce VGGT-Segmentor (VGGT-S), a framework that unifies robust geometric modeling with pixel-accurate semantic segmentation. VGGT-S leverages VGGT's powerful cross-view feature representation and introduces a novel Union Segmentation Head. This head operates in three stages: mask prompt fusion, coarse point-guided prediction, and iterative mask refinement, effectively translating high-level feature alignment into a precise segmentation mask. Furthermore, we propose a single-image self-supervised training strategy that eliminates the need for paired annotations and enables strong zero-shot generalization. On the challenging Ego–Exo4D benchmark, VGGT-S sets a new state-of-the-art, achieving 67.7% and 68.0% average IoU for Ego→Exo and Exo→Ego tasks, respectively, significantly outperforming prior methods. Notably, our zero-shot model surpasses most fully-supervised baselines, demonstrating the effectiveness and scalability of our approach.