Record interaction
Collect the robot’s applied forces and torques together with the object’s observed pose trajectory.
Recovering physical parameters from real robot interactions through vision-language priors and differentiable physics.
A robot can see an object move and feel the forces it applies. RigPI turns those observations into a physically grounded model.
Identifying mass, center of mass, inertia, and friction from real-world data is difficult when measurements are noisy and the starting guess is poor. RigPI combines vision-language model priors with a differentiable simulator to initialize, constrain, and refine these parameters.
RigPI uses recorded interaction data to refine a simulator until its predicted trajectory agrees with the observed motion.

Semantic priors help the optimization start in a plausible region; physics and observed motion guide the final estimate.
Collect the robot’s applied forces and torques together with the object’s observed pose trajectory.
Use visual semantic cues to initialize physical properties and define a feasible parameter range.
Differentiate through the simulated trajectory and update parameters to reduce the observed pose error.
Watch the project video for the real-world setup and trajectory reproduction.
Open video in DriveRigPI is an IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026 paper.
@article{he2026rigpi,
title={RigPI: Dynamic Parameter Identification of Rigid Body via VLM-Seeded Differentiable Simulation},
author={He, Xincheng and Zhang, Rongrong and Jiang, Wei and Xu, Wenqiang},
journal={arXiv preprint arXiv:2606.25212},
year={2026}
}