UniManip: General-Purpose Zero-Shot Robotic Manipulation with Agentic Operational Graph
Published in arXiv preprint arXiv:2602.13086, 2026
End-to-end VLA models often lack the precision for long-horizon tasks, while classical hierarchical planners are semantically rigid in open-world settings. UniManip bridges the two with a Bi-level Agentic Operational Graph: an Agentic Layer orchestrates the task while a Scene Layer maintains a dynamic, object-centric state representation. The system instantiates scene graphs from unstructured perception, turns them into collision-free trajectories through a safety-aware local planner, and uses structured memory to diagnose and recover from failures. It outperforms state-of-the-art VLA and hierarchical baselines by 22.5% and 25.0% success rate respectively, and transfers zero-shot from fixed-base to mobile manipulation without fine-tuning.
Preprint: arXiv:2602.13086 ยท Project page: henryhcliu.github.io/unimanip
Recommended citation: Haichao Liu, Yuanjiang Xue, Yuheng Zhou, Haoyuan Deng, Yinan Liang, Lihua Xie, Ziwei Wang. (2026). "UniManip: General-Purpose Zero-Shot Robotic Manipulation with Agentic Operational Graph." arXiv preprint arXiv:2602.13086.
Download Paper
