Learning to Explore Hidden Kinematicsfor Articulated Object Manipulation

*Equal contributionUniversity of Pennsylvania

Visual appearance can be misleading.

PartManip
Original · prismatic
✓
Modified · revolute
✕
Ours
Original · prismatic
✓
Modified · revolute
✓
Overview

We keep a belief over how a part is jointed, show it to the policy as a per-point articulation flow, and reward the policy for the uncertainty each interaction removes.

Abstract

The kinematics of an articulated object is often ambiguous from vision alone. Interaction resolves the ambiguity, and active perception methods exploit this by searching for the single action that most sharpens a belief over the kinematic parameters at each step. Such greedy search cannot be extended over a horizon without forward models of the contact and inertial dynamics, which are themselves unknown. We instead amortize action selection into training.

We maintain a belief distribution over joint type and parameters, initialized from a generative prior and updated by Bayesian filtering on the observed part motion. To condition the policy on this belief, we render it as a per-point articulation flow field, the motion that the current posterior predicts for every point on the object. Carrying the inductive bias of articulated motion, this representation generalizes better than a latent encoding of the belief or flow tracked from observation. We train the policy with reinforcement learning, rewarding the entropy that each interaction removes from the posterior, so that informative exploration becomes learned behavior rather than a search at every step.

Our method outperforms previous approaches across door and drawer manipulation on the PartManip benchmark, and on ArticuRiddle, a new dataset of objects whose appearance implies the wrong articulation.

Method

Belief, flow, and information gain

A generative prior seeds a set of articulation hypotheses. Each step, the belief becomes a flow field the policy observes, the policy acts, the observed motion updates the belief, and the entropy it removes is paid back as reward.

Pipeline: generative prior initializes particles; each hypothesis becomes a flow field; weighted flow and point cloud go to the actor; observed motion updates weights through a Bayes filter; entropy drop is the information reward for PPO.

Pipeline overview. Step through the four parts below.

Articulation belief

A belief over joint type and parameters

Each particle \(\theta^{(k)}=(c^{(k)},d^{(k)})\) is a joint type revolute prismatic fixed with its axis and pivot, drawn from the NAP generative prior and weighted by how well it explains the observed motion:

$$w_k \propto \exp\!\Big(-\tfrac{1}{2\sigma^2}\sum_j\big\|(\mathbf{x}'_j-\mathbf{x}_j)-\mathbf{f}_j(\theta^{(k)},\delta q_t)\big\|^2\Big)$$

Resampling keeps the hypotheses that keep explaining the motion.

Belief → flow

The belief, rendered as flow

The policy acts on point clouds, but the belief lives in joint space. Every hypothesis predicts a motion for every point: tangent to the swept circle for a revolute joint, uniform along the axis for a prismatic one. Fields are normalized per type and summed with the posterior weights, lifting each point from \(\mathbb{R}^6\) to \(\mathbb{R}^9\).

Recomputed every step, the flow is an implicit memory of past interactions. As the part moves, the hypothesis weights collapse onto the true axis.

Belief entropy

How much is still unknown

Particle weights alone mislead: once particles gather around the true axis their weights become similar, so weight entropy is high exactly when the belief has converged. We keep the type term exact and estimate each branch with a weighted kernel density over particle positions:

$$\mathbf{H}[b_t]=\mathbf{H}[P_t(c)]+\sum_{c\in\{r,p\}}P_t(c)\,\widehat{\mathbf{H}}[b_t(d\mid c)]$$
Information reward

Pay for the entropy removed

$$R_{\mathrm{info},t}=\mathbf{H}[b_t]-\mathbf{H}[b_{t+1}]$$

A difference of entropies, not a negated entropy. It vanishes once the belief stops moving, so each nat can be earned only once and standing still pays nothing.

A privileged state expert is distilled into the vision policy with DAgger, then fine-tuned with PPO on \(R=R_{\mathrm{task}}+\lambda_i R_{\mathrm{info}}\).

Benchmark

ArticuRiddle

Objects derived from PartManip where geometric cues point the wrong way: either the joint is swapped while the look stays the same, or the handle is moved or rotated so it suggests the wrong motion. Left: the PartManip original. Right: its ArticuRiddle edit.

PartManip
ArticuRiddle
Drag to orbit
Loading 3D
Results

The policy on each pair

What the policy sees

PartManip objects in IsaacGym, each paired with the point cloud the policy observes. Red arrows are the articulation flow from the current belief; many drawers start with a revolute-looking prior and switch to translation once the part moves.

Citation

BibTeX

@misc{liu2026learningexplorehiddenkinematics,
      title={Learning to Explore Hidden Kinematics for Articulated Object Manipulation},
      author={Ruiyao Liu and Boshu Lei and Zhuoyang Pan and Kostas Daniilidis},
      year={2026},
      eprint={2609.36553},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2609.36553},
}