TrackFish3D: Self-Supervised 3D Tracking of Schooling Fish from Multi-view Videos

Patt Phurtivilai1 Zhiyang Dou1,† Yifan Wu1 Kinfung Chu1 Yuan Liu2 Lei Yang1 Wenping Wang3 Taku Komura1,†

1 The University of Hong Kong, Hong Kong, China

2 Hong Kong University of Science and Technology, Hong Kong, China

3 Texas A&M University, College Station, TX, USA

† Corresponding authors

NeurIPS 2026

Overview

Abstract

TrackFish3D teaser showing multi-view fish tracking output.

Quantifying collective fish behavior requires accurate trajectories, yet multi-view 3D tracking remains challenging due to frequent occlusions, visually similar individuals, and the long-standing scarcity of identity annotations. We present TrackFish3D, a geometry-supervised framework for dense multi-camera 3D tracking of schooling fish. Instead of relying on appearance-based re-identification or manually annotated identities in the 3D tracking stage, TrackFish3D turns calibrated multi-view geometry into supervision: triangulation and reprojection consistency provide pseudo-associations, while a geometric encoder and global association transformer learn all-to-all cross-view correspondence within each frame.

To make these associations identity-aware, TrackFish3D introduces a self-supervised contrastive objective that separates co-visible individuals in the embedding space, together with a temporal predictor that preserves identities and bridges short occlusions across frames. The 3D association and tracking stage is trained once on unlabeled footage and applied directly to unseen test videos, requiring no cross-view identity labels, temporal annotations, 3D ground truth, appearance-based re-identification, or test-time optimization; upstream 2D perception models are pretrained or finetuned separately. On our benchmark, TrackFish3D (SAM3) improves 3D Multi-Object Tracking Accuracy from 87.7% for the strongest kinematic baseline to 95.8%. On the 3D-ZeF zebrafish benchmark, it achieves 81.1% MOTA, compared with 77.4% for the best geometric baseline. It also generalizes beyond fish, achieving strong results on real-world bird tracking.

Presentation

Video

Method

Framework

Overview diagram of the TrackFish3D framework.

Overview of the TrackFish3D framework. Synchronised multi-view video is processed by a Geometric Encoder and an Association Transformer that predict cross-view pairwise scores, supervised by geometric pseudo-supervision from triangulation consistency through a self-supervised association loss. Detections are fused via multi-view grouping and validation to produce 3D hypotheses, whose embeddings are separated by a self-supervised contrastive loss. Across frames, temporal linking and a self-supervised temporal loss enforce cross-frame consistency, enabling stable 3D trajectory recovery at inference.

Representation

Contrastive and Temporal Effects

Visualization of embedding separation with and without contrastive learning.

Contrastive Objective

Without contrastive training, same-fish detections overlap with other identities in the embedding space. With the contrastive objective, co-visible detections form compact, well-separated clusters, which makes cross-view grouping substantially more reliable in dense scenes.

Visualization of temporal consistency with and without temporal prediction training.

Temporal Prediction Objective

Without temporal training, same-fish embeddings drift across frames and cause identity switches. With the temporal prediction objective, embeddings remain close through short occlusions, enabling stable linking and improved identity consistency over time.

Evaluation

Results

SynFish Test Set

Across six held-out SynFish scenes, TrackFish3D achieves the best overall performance, reaching 96.6% MOTA with YOLO26+SAM2 inputs and 95.8% with SAM3, while also obtaining the highest MT% and strongest identity stability. The strongest baseline, 3D-SORT, reaches 87.7% MOTA, and MCTR reaches 83.7%, while TrackFish3D remains above 86% even on dense 16-fish scenes.

SynFish test-set averages

Method MOTA ↑ MT% ↑ ML% ↓ IDSw ↓ Frag ↓ MTBF ↑
SAM3 + Hungarian83.5%71.9%0.0%7.512.8183.8
SAM3 + Greedy83.5%71.9%0.0%7.512.8183.8
3D-SORT87.7%77.1%0.0%4.09.0209.4
YOLO26+SAM2 + Hungarian82.2%69.8%0.0%8.712.2183.1
YOLO26+SAM2 + Greedy82.2%69.8%0.0%8.712.2183.1
Self-MVA-64.6%0.0%96.9%1.72.215.9
ASNet33.3%49.6%11.5%20.829.060.7
SambaMOTR-10.5%11.7%73.8%0.07.851.8
MOTIP14.7%21.4%42.9%0.025.741.2
ReST61.3%74.8%0.0%173.778.517.5
MCTR83.7%82.7%0.0%1.714.2160.6
TrackFish3D (SAM3)95.8%95.8%0.0%0.06.5266.3
TrackFish3D (YOLO26+SAM2)96.6%95.8%0.0%0.23.7296.5
Qualitative trajectory comparison on SynFish test scenes.

SynFish Qualitative Comparison

On held-out SynFish scenes, TrackFish3D preserves identity continuity and produces cleaner 3D trajectories than representative baselines. Trajectory breaks indicate fragmentations or identity switches.

3D-ZeF Test Set

On the real 3D-ZeF test set, TrackFish3D reaches 81.1% MOTA and outperforms all baselines. The MOTA gap over simple geometric methods is modest because these scenes contain only five fish, but the advantage becomes clearer in identity stability, with fewer switches, lower fragmentation, and substantially higher mean time between failures.

3D-ZeF test-set averages

Method MOTA ↑ MT% ↑ ML% ↓ IDSw ↓ Frag ↓ MTBF ↑
Hungarian + nearest77.4%90.0%0.0%6.073.042.5
Greedy + nearest76.9%90.0%0.0%10.076.038.7
3D-SORT75.6%90.0%0.0%13.565.543.4
Self-MVA50.8%60.0%0.0%21.573.548.5
ASNet0.3%0.0%0.0%72.5175.010.6
SambaMOTR-26.4%0.0%100.0%0.51.06.5
MOTIP-45.9%0.0%100.0%1.56.52.7
ReST0.0%0.0%100.0%0.00.00.0
MCTR0.1%0.0%40.0%21.057.58.8
TrackFish3D81.1%90.0%0.0%4.030.593.6
Qualitative trajectory comparison on 3D-ZeF test sequences.

3D-ZeF Qualitative Comparison

TrackFish3D maintains identity consistency through real-world occlusions that fragment or swap baseline trajectories. ReST produces no trajectories because its cross-view association fails to match detections correctly.