Jiahe Xu (Jay)
I am the CTO of Pinocchio AI, where we build realistic robots that interact with people through lifelike, animal-inspired behaviors and fit naturally into daily life.
My research focuses on object-centric manipulation, bimanual manipulation, learning from human video, and 3D robot learning.
Previously, I was a systems engineer for sensors, systems, and software at CMU. I worked on sensor calibration for pinhole, fisheye, stereo, thermal cameras, IMU, and LiDAR; teleoperation for real-to-sim and real-to-real systems; real-time robotics systems with ROS1 and ROS2; and 3D scene reconstruction. I have worked with Jetson Orin/Xavier, NUC, iOS, iPhone, and Vision Pro platforms; Boston Dynamics Spot, custom RC cars, custom quadrotors, Franka Panda, Mobile Aloha, and UR5 robots; and simulation environments including PyBullet, SAPIEN, Isaac Gym/Isaac Lab, MuJoCo, and Genesis.
My goal is to build intelligent systems that can adapt to almost any robot from only a handful of demonstrations.
Email: xjh4438318846@gmail.com
Google Scholar
LinkedIn
Robot Videos
CV (Updated 05/21/2025)
OC3D: Object-Centric 3D Diffusion Planner
Jiahe Xu*, Gang Tan
IROS2026
|
Webpage
Flying Hand: End-Effector-Centric Framework for Versatile Aerial Manipulation Teleoperation and Policy Learning
Guanqi He*, Xiaofeng Guo*, Luyi Tang, Yuanhang Zhang, Mohammadreza Mousaei, Jiahe Xu, Junyi Geng, Sebastian Scherer, Guanya Shi
IEEE Robotics and Automation Letters (RA-L), to be presented at ICRA 2025.
Paper
|
Webpage
My skills tables
Robot Learning
| Skills | 🤖 Teleoperation | 🧠 3D robot learning | 👁️ VLA | 📁 Data collection | 🔁 Sim2real & real2sim | 🗣️ Robot language control | 🔄 Auto-eval |
|---|---|---|---|---|---|---|---|
| Proficiency | Advanced | Advanced | Advanced | Advanced | Advanced | Advanced | Experienced |
| Details | Hands-on teleoperation for dexterous hands and bimanual tasks, from 3D hand-point detection to retargeting with under 60ms delay | Object-centric robot policies using point clouds and visual features for single-arm and bimanual manipulation tasks | Training and evaluation for VLA models including SPOT, AnyPlace, iDP3, 3DDA, 3DFA, Pi0, and Pi0.5, with practical understanding of strengths and limits | Real-time multi-camera recording and depth estimation for reliable robot data collection, with synchronized MP4 video segments and MCAP logs | Real-to-sim tools for converting real objects into simulation assets, rescaling meshes, tracking rigid objects, and replaying calibrated scenes | Voice-command pipeline connecting wake-word detection, speech recognition, an agent proxy, and TTS for natural robot interaction | Auto-evaluation loop that detects rollout completion, records success or failure, and resets the robot and objects for the next trial |
System & Sensor
| Skills | 🚀 System deployment | 🎮 Simulation | 🎯 Calibration | 💻 Edge computing | ⏱️ Real-time control | 🎥 Video streaming | 🗃️ Large dataset management |
|---|---|---|---|---|---|---|---|
| Proficiency | Advanced | Advanced | Advanced | Advanced | Advanced | Experienced | Experienced |
| Details | Linux deployment workflows with Docker, SSH host adoption, javis/deployer tooling, remote code sync, service setup, and Jetson/x86 payload management | Policy testing, rollout replay, and scene reset workflows in PyBullet, SAPIEN, IsaacGym/IsaacLab, MuJoCo, and Genesis | Calibration for pinhole, fisheye, stereo, and thermal cameras with IMU, LiDAR, and robot frames, reducing reprojection and FK errors | Edge deployment optimization with TensorRT, VPI, GStreamer, Nvidia NVMEM encoders, camera capture, and device I/O on Jetson and NUC | Real-time Aloha end-effector controller built from scratch, running up to 50Hz with feedback, limits, and watchdogs | Low-latency robot video streaming over radio using DDS, GStreamer, and Nvidia NVMEM encoders | Large robot dataset management across RH20T, Open X-Embodiment, and BridgeData, including rosbags, metadata, and train-test splits |
Perception
| Skills | 🌐 3D scene graph | 📡 Sensor fusion | 🤝 Multi-robot system | 📏 Depth estimation | 🗺️ SLAM |
|---|---|---|---|---|---|
| Proficiency | Advanced | Advanced | Advanced | Experienced | Experienced |
| Details | Developed a bot-level 3D scene graph, combining object, place, and spatial relation layers into an LLM-ready world representation | Sensor fusion across stereo, depth, thermal, IMU, LiDAR, radar, and gimbal data with careful time sync and calibration | Merged perception outputs across heterogeneous robots, aligning LiDAR maps, point clouds, camera frames, TF trees, and shared object states into consistent robot-team context | Benchmarked stereo and monocular depth pipelines including FoundationStereo, ZED X depth, Depth Anything V2, UniDepthV2, and Metric3D for robot data collection | 3D mapping with RGB-D, LiDAR, and calibrated camera rigs, producing maps for planning and downstream robot use |


