DirtyMoCap: Robust Motion Capture from Unconstrained Markers

Loading video · 0%

Authors and affiliations

1Zhejiang University2Westlake University3Fudan University4National University of Defense Technology5Nanjing University6National University of Singapore

SIGGRAPH Asia 2026

Abstract

Optical motion capture delivers high-fidelity human motion, but its reliance on strict marker layouts and clean trajectories severely limits its real-world applicability. In practice, tracking systems frequently output unconstrained markers - sparse, noisy, and unordered point clouds with unknown or varying configurations. To bridge the gap between corrupted raw markers and parametric human models, we introduce DirtyMoCap, a robust, marker-layout-free framework. Our core insight is to map unordered marker observations to a fixed set of "proxy anchors" - comprising skeletal joints and body surface points - acting as a stable intermediate representation. We first initialize and track these anchors over long sequences using a recurrent sliding-window architecture. Then, a custom differentiable Gauss-Newton solver fits the SMPL-H model to the tracked anchors to recover full-body pose, translation, and shape. By explicitly deriving geometric residuals, our solver learns adaptive observation confidence, smoothness, and prior weights end-to-end, adapting dynamically to the reliability of the input data. Extensive experiments on diverse, noisy marker configurations demonstrate that DirtyMoCap successfully generalizes across arbitrary layouts using only a single trained model. It consistently outperforms state-of-the-art configuration-specific baselines in both joint and vertex reconstruction accuracy, while our custom CUDA solver achieves up to a 100× speedup over standard PyTorch implementations. We further apply DirtyMoCap to heterogeneous raw optical MoCap recordings of traditional Chinese martial arts, yielding a Kung Fu motion dataset of temporally coherent SMPL-H reconstructions.

TL;DRRecovers 4D SMPL-H motion from dirty markers — sparse, noisy, and unordered — via proxy anchors and a differentiable and learnable Gauss–Newton solver, generalizing across marker configurations with a single model.

Method Overview

DirtyMoCap pipeline showing anchor initialization, anchor trajectory prediction, and the differentiable learnable Gauss-Newton solver
Overview of DirtyMoCap. The Anchor Initialization step (Sec. 3.1) uses the first-frame markers M1 to estimate initial anchors A1 and replicates them across the remaining frames in the current window to form Ainitt:t+W−1. The resulting anchor sequence, together with Mt:t+W−1, is then passed to the Anchor Trajectory Prediction module (Sec. 3.2) to refine At:t+W−1. Finally, a differentiable learnable Gauss–Newton solver (Sec. 3.3) optimizes the SMPL-H parameters, including pose, global translation, and shape. The estimates from the current window initialize the next one.

DirtyMoCap on CMU+GRAB Dataset

HKMALA-Motion Dataset Reconstructed by DirtyMoCap

We further apply DirtyMoCap to a heterogeneous collection of raw optical MoCap recordings of traditional Chinese martial arts, for which marker configurations and marker identities are unavailable. The collection contains real motion sequences captured between September 2013 and March 2022, averaging approximately 8,500 frames per sequence and totaling about 180 minutes across 22 normalized martial-arts style labels. Processing these noisy and unordered observations with DirtyMoCap yields temporally coherent SMPL-H motion and enables the construction of the Kung Fu motion dataset.

Wing Chung永春
Eagle Claw Fan Tsi Moon鷹爪翻子門
Huen Kuen洪拳
Eagle Claw Fan Tsi Moon鷹爪翻子門
Hung Sing Choy lee Fat鴻勝蔡李佛
Tai Sing Pap Kar Moon大聖劈掛門
Yang's Taiji楊式太極
Hung Sing Choy lee Fat鴻勝蔡李佛
Gao style Bagua Zhang高式八卦掌
Jin Wu Koon Shaolin Double Dragon少林雙龍派振武館
Ma's Tong Bei馬氏通備
Eagle Claw Fan Tsi Moon鷹爪翻子門

Acknowledgements

We thank International Guoshu Association Limited and the Institute of Chinese Martial Studies Limited for providing access to the raw martial-arts motion-capture recordings used to construct the HKMALA-Motion dataset. This work is funded by the Research Center for Industries of the Future (RCIF) at Westlake University, the Westlake Education Foundation.