Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation

Published in arXiv preprint arXiv:2609.01596, 2026

Assembly at sub-millimeter tolerances needs spatial precision, compliant interaction, and robustness to contact failures. Facet-0 unifies multimodal representation learning and RL post-training around a joint action-wrench proposal: a causal wrench history is aligned with vision-language semantics and kinematic state, and flow matching generates each action chunk together with the wrist-wrench profile it should induce. A distributional Action-Wrench Critic learned from deployment rollouts separates motions with similar task progress but different contact outcomes, while phase-aware rewards concentrate policy improvement on decisive interactions. Trained on ManuFacet-1K, a 1,000-hour force-synchronized corpus across three embodiments, it reaches 82% mean success on five sub-millimeter computer-assembly tasks versus 15% for the strongest baseline, at 0.5 mm placement accuracy and 50 ms command latency.

Preprint: arXiv:2609.01596

Recommended citation: Haoyuan Deng, Haichao Liu, Wenkai Guo, Yuan Ling, Zaijia Yang, Yuanjiang Xue, Haosheng Sun, Liangzi Wang, Ziwei Wang. (2026). "Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation." arXiv preprint arXiv:2609.01596.
Download Paper