E2HiL: Entropy-Guided Sample Selection for Efficient Real-World Human-in-the-Loop Reinforcement Learning

Published in IEEE Robotics and Automation Letters (RA-L), 2026

Human-in-the-loop reinforcement learning makes real-world robot training feasible, but the human corrections it depends on are expensive to collect. E2HiL uses an entropy-guided criterion to decide which samples are worth a human’s attention, so the policy reaches the same task performance with substantially fewer interventions on real robotic manipulation tasks.

Preprint: arXiv:2601.19969

Recommended citation: Haoyuan Deng, Yudong Lin, Yuanjiang Xue, Haoyang Du, Qianzhun Wang, Boyang Zhou, Zhenyu Wu, Ziwei Wang. (2026). "E2HiL: Entropy-Guided Sample Selection for Efficient Real-World Human-in-the-Loop Reinforcement Learning." IEEE Robotics and Automation Letters.
Download Paper