DirtyMoCap: Robust Motion Capture from Unconstrained Markers
Authors and affiliations
1Zhejiang University2Westlake University3Fudan University4National University of Defense Technology5Nanjing University6National University of Singapore
SIGGRAPH Asia 2026
Abstract
Optical motion capture delivers high-fidelity human motion, but its reliance on strict marker layouts and clean trajectories severely limits its real-world applicability. In practice, tracking systems frequently output unconstrained markers - sparse, noisy, and unordered point clouds with unknown or varying configurations. To bridge the gap between corrupted raw markers and parametric human models, we introduce DirtyMoCap, a robust, marker-layout-free framework. Our core insight is to map unordered marker observations to a fixed set of "proxy anchors" - comprising skeletal joints and body surface points - acting as a stable intermediate representation. We first initialize and track these anchors over long sequences using a recurrent sliding-window architecture. Then, a custom differentiable Gauss-Newton solver fits the SMPL-H model to the tracked anchors to recover full-body pose, translation, and shape. By explicitly deriving geometric residuals, our solver learns adaptive observation confidence, smoothness, and prior weights end-to-end, adapting dynamically to the reliability of the input data. Extensive experiments on diverse, noisy marker configurations demonstrate that DirtyMoCap successfully generalizes across arbitrary layouts using only a single trained model. It consistently outperforms state-of-the-art configuration-specific baselines in both joint and vertex reconstruction accuracy, while our custom CUDA solver achieves up to a 100× speedup over standard PyTorch implementations. We further apply DirtyMoCap to heterogeneous raw optical MoCap recordings of traditional Chinese martial arts, yielding a Kung Fu motion dataset of temporally coherent SMPL-H reconstructions.
TL;DRRecovers 4D SMPL-H motion from dirty markers — sparse, noisy, and unordered — via proxy anchors and a differentiable and learnable Gauss–Newton solver, generalizing across marker configurations with a single model.
Method Overview

DirtyMoCap on CMU+GRAB Dataset
HKMALA-Motion Dataset Reconstructed by DirtyMoCap
We further apply DirtyMoCap to a heterogeneous collection of raw optical MoCap recordings of traditional Chinese martial arts, for which marker configurations and marker identities are unavailable. The collection contains real motion sequences captured between September 2013 and March 2022, averaging approximately 8,500 frames per sequence and totaling about 180 minutes across 22 normalized martial-arts style labels. Processing these noisy and unordered observations with DirtyMoCap yields temporally coherent SMPL-H motion and enables the construction of the Kung Fu motion dataset.
Acknowledgements
We thank International Guoshu Association Limited and the Institute of Chinese Martial Studies Limited for providing access to the raw martial-arts motion-capture recordings used to construct the HKMALA-Motion dataset. This work is funded by the Research Center for Industries of the Future (RCIF) at Westlake University, the Westlake Education Foundation.
