Quantifying collective fish behavior requires accurate trajectories, yet multi-view 3D tracking
remains challenging due to frequent occlusions, visually similar individuals, and the long-standing
scarcity of identity annotations. We present TrackFish3D, a geometry-supervised framework for dense
multi-camera 3D tracking of schooling fish. Instead of relying on appearance-based re-identification
or manually annotated identities in the 3D tracking stage, TrackFish3D turns calibrated multi-view
geometry into supervision: triangulation and reprojection consistency provide pseudo-associations,
while a geometric encoder and global association transformer learn all-to-all cross-view
correspondence within each frame.
To make these associations identity-aware, TrackFish3D introduces a self-supervised contrastive
objective that separates co-visible individuals in the embedding space, together with a temporal
predictor that preserves identities and bridges short occlusions across frames. The 3D association
and tracking stage is trained once on unlabeled footage and applied directly to unseen test videos,
requiring no cross-view identity labels, temporal annotations, 3D ground truth, appearance-based
re-identification, or test-time optimization; upstream 2D perception models are pretrained or
finetuned separately. On our benchmark, TrackFish3D (SAM3) improves 3D Multi-Object Tracking
Accuracy from 87.7% for the strongest kinematic baseline to 95.8%. On the 3D-ZeF zebrafish
benchmark, it achieves 81.1% MOTA, compared with 77.4% for the best geometric baseline. It also
generalizes beyond fish, achieving strong results on real-world bird tracking.