Jiahe Xu (Jay)

I am the CTO of Pinocchio AI, where we build realistic robots that interact with people through lifelike, animal-inspired behaviors and fit naturally into daily life. My research focuses on object-centric manipulation, bimanual manipulation, learning from human video, and 3D robot learning.

Previously, I was a systems engineer for sensors, systems, and software at CMU. I worked on sensor calibration for pinhole, fisheye, stereo, thermal cameras, IMU, and LiDAR; teleoperation for real-to-sim and real-to-real systems; real-time robotics systems with ROS1 and ROS2; and 3D scene reconstruction. I have worked with Jetson Orin/Xavier, NUC, iOS, iPhone, and Vision Pro platforms; Boston Dynamics Spot, custom RC cars, custom quadrotors, Franka Panda, Mobile Aloha, and UR5 robots; and simulation environments including PyBullet, SAPIEN, Isaac Gym/Isaac Lab, MuJoCo, and Genesis.

My goal is to build intelligent systems that can adapt to almost any robot from only a handful of demonstrations.

Email: xjh4438318846@gmail.com

Google Scholar     LinkedIn     Robot Videos     CV (Updated 05/21/2025)    

OC3D: Object-Centric 3D Diffusion Planner

Jiahe Xu*, Gang Tan
IROS2026
  |   Webpage

3D FlowMatch Actor: Unified 3D Policy for Single- and Dual-Arm Manipulation

Nikolaos Gkanatsios*, Jiahe Xu*, Matthew Bronars, Arsalan Mousavian, Tsung-Wei Ke, Katerina Fragkiadak
IROS 2026.
Paper   |   Webpage

Flying Hand: End-Effector-Centric Framework for Versatile Aerial Manipulation Teleoperation and Policy Learning

Guanqi He*, Xiaofeng Guo*, Luyi Tang, Yuanhang Zhang, Mohammadreza Mousaei, Jiahe Xu, Junyi Geng, Sebastian Scherer, Guanya Shi
IEEE Robotics and Automation Letters (RA-L), to be presented at ICRA 2025.
Paper   |   Webpage

Flying Calligrapher: Contact-Aware Motion and Force Planning and Control for Aerial Manipulation

Xiaofeng Guo*, Guanqi He*, Jiahe Xu, Mohammadreza Mousaei, Junyi Geng, Sebastian Scherer, Guanya Shi
IEEE Robotics and Automation Letters (RA-L), to be presented at ICRA 2025.
Paper   |   Webpage

PyPose: A Library for Robot Learning with Physics-based Optimization

CVPR 2023
Paper   |   Webpage

My skills tables

Robot Learning

Skills 🤖 Teleoperation 🧠 3D robot learning 👁️ VLA 📁 Data collection 🔁 Sim2real & real2sim 🗣️ Robot language control 🔄 Auto-eval
Proficiency Advanced Advanced Advanced Advanced Advanced Advanced Experienced
Details Hands-on teleoperation for dexterous hands and bimanual tasks, from 3D hand-point detection to retargeting with under 60ms delay Object-centric robot policies using point clouds and visual features for single-arm and bimanual manipulation tasks Training and evaluation for VLA models including SPOT, AnyPlace, iDP3, 3DDA, 3DFA, Pi0, and Pi0.5, with practical understanding of strengths and limits Real-time multi-camera recording and depth estimation for reliable robot data collection, with synchronized MP4 video segments and MCAP logs Real-to-sim tools for converting real objects into simulation assets, rescaling meshes, tracking rigid objects, and replaying calibrated scenes Voice-command pipeline connecting wake-word detection, speech recognition, an agent proxy, and TTS for natural robot interaction Auto-evaluation loop that detects rollout completion, records success or failure, and resets the robot and objects for the next trial

System & Sensor

Skills 🚀 System deployment 🎮 Simulation 🎯 Calibration 💻 Edge computing ⏱️ Real-time control 🎥 Video streaming 🗃️ Large dataset management
Proficiency Advanced Advanced Advanced Advanced Advanced Experienced Experienced
Details Linux deployment workflows with Docker, SSH host adoption, javis/deployer tooling, remote code sync, service setup, and Jetson/x86 payload management Policy testing, rollout replay, and scene reset workflows in PyBullet, SAPIEN, IsaacGym/IsaacLab, MuJoCo, and Genesis Calibration for pinhole, fisheye, stereo, and thermal cameras with IMU, LiDAR, and robot frames, reducing reprojection and FK errors Edge deployment optimization with TensorRT, VPI, GStreamer, Nvidia NVMEM encoders, camera capture, and device I/O on Jetson and NUC Real-time Aloha end-effector controller built from scratch, running up to 50Hz with feedback, limits, and watchdogs Low-latency robot video streaming over radio using DDS, GStreamer, and Nvidia NVMEM encoders Large robot dataset management across RH20T, Open X-Embodiment, and BridgeData, including rosbags, metadata, and train-test splits

Perception

Skills 🌐 3D scene graph 📡 Sensor fusion 🤝 Multi-robot system 📏 Depth estimation 🗺️ SLAM
Proficiency Advanced Advanced Advanced Experienced Experienced
Details Developed a bot-level 3D scene graph, combining object, place, and spatial relation layers into an LLM-ready world representation Sensor fusion across stereo, depth, thermal, IMU, LiDAR, radar, and gimbal data with careful time sync and calibration Merged perception outputs across heterogeneous robots, aligning LiDAR maps, point clouds, camera frames, TF trees, and shared object states into consistent robot-team context Benchmarked stereo and monocular depth pipelines including FoundationStereo, ZED X depth, Depth Anything V2, UniDepthV2, and Metric3D for robot data collection 3D mapping with RGB-D, LiDAR, and calibrated camera rigs, producing maps for planning and downstream robot use