CV

Yuanjiang Xue

E250086@e.ntu.edu.sg
(+86) 18744915886
Singapore, , SG

Summary

M.Sc. student in Computer Control & Automation at the School of EEE, Nanyang Technological University, working on robot learning, dexterous manipulation and embodied AI, with published work on robotic foundation models, human-in-the-loop reinforcement learning, zero-shot manipulation and force-aware teleoperation. Previously an Optoelectronic Information Science and Engineering graduate from Wuhan University with industry experience in image processing and embedded hardware development.

Education

  • Computer Control & Automation
    Present
    Nanyang Technological University
  • Optoelectronic Information Science and Engineering
    2023
    Wuhan University

Work Experience

  • Image Processing Engineer
    2023-07 - 2023-11
    TP-Link Technologies Co., Ltd.
    Shenzhen, China
    • Optimized algorithms to reduce mosquito trail effects in infrared night-vision scenarios for certain high-speed PTZ camera models, enhancing image processing performance for security cameras.
    • Adjusted algorithms to reduce noise in dark areas of indoor PTZ camera models in wide dynamic range daytime scenes, mitigating noise fluctuations and improving image processing performance for security cameras.
    • Investigated clarity quantification metrics suitable for multiple scenarios.
  • Hardware Development Engineer
    2023-12 - 2024-06
    Suzhou Sifei Technology Co., Ltd.
    Suzhou, China
    • Participated in the design and development of the EVM_MSPM0L1306 learning board for the MSPM0 series M0 chip.
    • Contributed to the design and development of the EVM_MSPM0G3507 learning board for the MSPM0 series M0 chip.

Skills

Programming Languages

  • C/C++
  • Python
  • Java

Platforms

  • Linux
  • Windows

Software

  • Matlab
  • Solidworks

Robot Learning

  • PyTorch
  • Reinforcement Learning
  • Imitation Learning
  • Real-Robot Deployment

Publications

  • Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation
    2026
    arXiv preprint arXiv:2609.01596
    H. Deng, H. Liu, W. Guo, Y. Ling, Z. Yang, Y. Xue, H. Sun, L. Wang, Z. Wang. A robotic foundation model that predicts and values the contact consequences of its actions, reaching 82% mean success on sub-millimeter computer-assembly tasks.
  • DexTeleop-0: Force-Aware Bimanual Dexterous Teleoperation with Ego-Centric Perception towards Shared Autonomy
    2026
    arXiv preprint arXiv:2606.23431
    H. Liu, Y. Jiang, H. Park, Y. Xue, Z. Wang. A tactile-driven adaptation strategy that translates coarse human tracking intent into precise, force-compliant robot commands for bimanual dexterous manipulation.
  • UniManip: General-Purpose Zero-Shot Robotic Manipulation with Agentic Operational Graph
    2026
    arXiv preprint arXiv:2602.13086
    H. Liu, Y. Xue, Y. Zhou, H. Deng, Y. Liang, L. Xie, Z. Wang. A bi-level Agentic Operational Graph unifying semantic reasoning and physical grounding, outperforming VLA and hierarchical baselines by 22.5% and 25.0% success rate.
  • E2HiL: Entropy-Guided Sample Selection for Efficient Real-World Human-in-the-Loop Reinforcement Learning
    2026
    IEEE Robotics and Automation Letters (RA-L)
    H. Deng, Y. Lin, Y. Xue, H. Du, Q. Wang, B. Zhou, Z. Wu, Z. Wang. Entropy-guided active sample selection that reduces the human interventions needed for real-world human-in-the-loop reinforcement learning.

Portfolio

  • Facet-0 — Robotic Foundation Model for Contact-Rich Precise Manipulation
    2026
    Ntu research — arxiv:2609.01596, 2026
    Robotic foundation model that predicts and values the contact consequences of its own actions. Flow matching generates each action chunk jointly with the future wrist-wrench profile it is expected to induce, over a representation aligning causal wrench history with vision-language semantics and kinematic state. A distributional Action-Wrench Critic learned from deployment rollouts separates motions with similar task progress but different contact outcomes, while phase-aware rewards and contact-selective credit concentrate policy improvement on decisive interactions; a lightweight bounded actor reuses the frozen representation for on-robot adaptation. Trained on ManuFacet-1K (1,000-hour force-synchronized corpus, three embodiments): 82% mean success on five sub-millimeter assembly tasks versus 15% for the strongest baseline, at 0.5 mm placement accuracy and 50 ms latency.
  • UniManip — General-Purpose Zero-Shot Manipulation with Agentic Operational Graph
    2026
    Ntu research — arxiv:2602.13086, 2026
    Bridges end-to-end VLA models, which lack long-horizon precision, and hierarchical planners, which are semantically rigid in open-world settings. A bi-level Agentic Operational Graph couples a high-level Agentic Layer for task orchestration to a low-level Scene Layer for dynamic state representation, continuously aligning abstract planning with live geometric constraints. It runs as a dynamic agentic loop: object-centric scene graphs instantiated from unstructured perception, parameterized into collision-free trajectories by a safety-aware local planner, with structured memory diagnosing and recovering from execution failures. Achieves 22.5% and 25.0% higher success than state-of-the-art VLA and hierarchical baselines, and transfers zero-shot from fixed-base to mobile manipulation without fine-tuning.
  • DexTeleop-0 — Force-Aware Bimanual Dexterous Teleoperation towards Shared Autonomy
    2026
    Ntu research — arxiv:2606.23431, 2026
    Attacks the contact-rich data bottleneck, where embodiment gaps break kinematic mapping and tactile and force feedback are typically absent. A tactile-driven adaptation strategy layers onto existing teleoperation pipelines: a real-time optimization loop estimates accurate contact points, reads a tactile-enabled fingertip force-sensing profile, and computes localized corrections through the operational space Jacobian with respect to joint-angle updates, turning coarse human tracking intent into precise force-compliant commands. Validated in simulation and on real hardware with higher success rates and execution efficiency on robust grasping, disturbance-resilient manipulation and complex dexterous tasks.
  • E2HiL — Entropy-Guided Sample Selection for Human-in-the-Loop Reinforcement Learning
    2026
    Ntu research — ieee ra-l, 2026
    Sample-efficient human-in-the-loop RL framework targeting the labor cost that limits real-world online learning. Built on the insight that a stable, rather than merely rapid, reduction of policy entropy yields a better exploration-exploitation trade-off. Influence functions model each sample's effect on policy entropy, efficiently estimated from the covariance of action probabilities and soft advantages; samples of moderate influence are retained while shortcut samples inducing sharp entropy drops and noisy samples with negligible effect are pruned. Across 10 real-world manipulation tasks spanning multiple embodiments and learning frameworks: 24.9% higher success rate with 9.3% fewer human interventions than state-of-the-art HiL-RL baselines, as a policy- and embodiment-agnostic plug-and-play module.
  • Self-Developed UAV at Wuhan University
    2022
    Bachelor thesis
    Completed the structural design, hardware assembly, flight control algorithm development, and flight debugging of the self-developed UAV at Wuhan University. Participated in the drone gimbal debugging project with Xiangtuo Technology. Assisted in the UAV video stabilization optimization work with Zhuomu Technology.
  • Event Camera-Based Visual Microphone
    2021
    Invention patent, national innovation and entrepreneurship project
    Proposed an event camera-based vibration sensing method and implemented an amplitude extraction algorithm based on event point statistics. This method uses laser illumination on the object to capture laser speckle event streams, and through denoising, time segmentation, and event point counting, it efficiently reconstructs vibration signals.
  • Adaptive Time-Slotted Clutter Map Constant False Alarm Rate Detection Method for External Radiation Source Radar
    2020
    Invention patent
    Developed an adaptive time division clutter map CFAR detection method for passive radar, improving false alarm performance in complex environments. Divided clutter intervals based on distance-Doppler spectra, utilizing the characteristics of separated transmission and reception to optimize clutter detection in the spatial domain. Established and initialized the clutter matrix, with real-time optimization through adaptive updates and forgetting factors to adapt to complex target signals and noise.
  • Emotion Recognition for the Elderly Using Wearable Devices
    2019
    Provincial-level innovation and entrepreneurship project
    Designed and developed a portable emotion monitoring device, including wearable device creation, physiological data collection, and emotion recognition model training and prediction.
  • 5th National College Student Integrated Circuit Innovation and Entrepreneurship Competition
    2020
    Competition - 3rd prize in central china region
    Modified chip IP core design, circuit design and soldering, code writing and debugging. Designed wireless intelligent development board, signal and information processing, IC design, and remote wireless real-time monitoring.
  • 6th National College Student Integrated Circuit Innovation and Entrepreneurship Competition
    2021
    Competition - 3rd prize in central china region, 1st prize in preliminary round
    Ported and accelerated target recognition based on neural networks to hardware development boards using FPGA-based CNN accelerators for SSD target detection.
  • Wuhan University "Zhixing Cup" Electronic Design Competition
    2020
    Competition - 1st prize
    Designed an intelligent tracking car with Bluetooth control, infrared tracking, ultrasonic obstacle avoidance, OLED display, and light-following docking at a designated position.

Interests

  • Embodied Foundation Models & VLA
    multimodal manipulation policies, RL post-training, real interaction grounding
  • Contact-Rich & Force-Aware Manipulation
    wrench-conditioned policies, sub-millimeter assembly, compliant interaction, tactile sensing
  • Sample-Efficient Real-World RL
    human-in-the-loop RL, active sample selection, policy entropy
  • Agentic Manipulation & Embodied Reasoning
    object-centric scene graphs, task orchestration, failure diagnosis and recovery
  • Dexterous Teleoperation & Shared Autonomy
    bimanual dexterous control, ego-centric perception, force feedback