Training signal from human video
Learning 3D articulation remains limited by the availability of real-world 3D annotations. Existing approaches rely on point clouds, RGB-D scans, CAD models or explicitly annotated articulation parameters, which are expensive to acquire and cover a narrow range of scenes and object categories. Models trained on such data generalize poorly to the diversity of objects encountered in daily life.
Egocentric video offers broader supervision. Datasets such as EPIC-KITCHENS, HOI4D and ARCTIC capture people interacting with diverse objects at a scale that 3D annotation cannot match, but they provide only RGB observations and 2D motion cues. EgoArt connects the two through an analytic trajectory generator: the predicted mask, interaction point, motion type, axis, hinge and extent define a 3D interaction trajectory, so a 2D hand trajectory and a 3D articulation label supervise the same parameters.
Records mined from human video
Each record is one frame with the mask of the actionable part (red), the track of the acting hand's knuckle over the interaction (green; ring at the interaction point, arrowhead at the end), a motion-type label and an instruction. These are mined automatically from three sources.
HOI4D
Table-top interactions with household objects









EPIC-KITCHENS
Kitchen activity recorded from a head-mounted camera









ARCTIC
Two-hand manipulation of small articulated objects









Abstract
Manipulating articulated objects requires identifying the relevant part and predicting how it can move. We present EgoArt, an instruction-conditioned model that grounds the task-relevant part from a single RGB image and predicts its 3D articulation affordance, including the mask, interaction point, motion type, axis, hinge, and motion extent. These predictions directly define a 3D interaction trajectory.
To reduce reliance on limited 3D annotations, we bridge 2D human hand trajectories and 3D articulation through an analytic trajectory generator, allowing egocentric videos without 3D labels to supervise explicit articulation parameters. The same formulation also supports joint training with 3D-labelled data, for which we derive a closed-form trajectory loss that requires neither explicit trajectory sampling nor annotated motion extent.
Trained jointly on SceneFun3D, HOI4D, EPIC-KITCHENS, and ARCTIC, EgoArt improves axis and hinge estimation over prior image-based articulation methods while generalizing to object categories seen only in 2D human video.
Method
One image and one instruction go in. A language-conditioned encoder produces a feature map; prediction heads estimate the part mask, the interaction point, the motion type, the axis, the hinge and the motion extent. An analytic decoder turns these parameters into the interaction trajectory, so a 2D hand track and a 3D annotation supervise the same quantities.

Real-robot demonstrations
EgoArt runs on a single frame from the gripper camera of a Boston Dynamics Spot, about a metre from the object, with the instruction shown. The robot walks to the predicted part, re-localizes the handle up close, and executes the predicted trajectory with its arm. Left: external view with the predicted trajectory (cyan) and axis (pink) drawn in. Right: the gripper camera. Clips play at 2x from the start of the reach to the end of the articulation.
Results
Each card shows the ground truth next to one method; pick the method below the images. Red: predicted part. Yellow: axis, drawn as the hinge line with its foot for revolute parts and as the direction of travel for prismatic ones. Green: motion of the interaction point (white) over the extent. Trajectories appear only for the ground truth, EgoArt and the SceneFun3D-only control, since the other baselines do not predict an interaction point; none of them predicts motion extent.
SceneFun3D test frames
The 3D-labelled benchmark. Baselines retrained on the same split; detectors show their oracle-matched instance.








Hand-video test frames
Held-out records from the human-video sources; the hand track is the ground-truth motion. "SceneFun3D-only" is the same model trained without human video. Categories absent from SceneFun3D are marked.










Phone photos outside every training set
Rooms that appear in neither SceneFun3D nor the hand-video sources. The photo is on the left; choose the model on the right. On the laptop, a category absent from SceneFun3D, only EgoArt finds the screen and calls it revolute; the SceneFun3D-only model predicts a slide on the keyboard.








Citation
@inproceedings{egoart2027,
title = {EgoArt: Learning 3D Articulation Affordances from Egocentric Human Interaction Videos},
author = {Andrew Sangwoo Ye and Jiaqi Chen and Marc Pollefeys and Taein Kwon},
booktitle = {Under review},
year = {2027}
}