Under review · ICRA 2027

EgoArt

Learning 3D Articulation Affordances from Egocentric Human Interaction Videos

Andrew Sangwoo Ye1, Jiaqi Chen1, Marc Pollefeys1, Taein Kwon1*

1ETH Zürich*Corresponding author

EgoArt learns 3D articulation from 2D human video
EgoArt learns 3D articulation from 2D human video. Trained on hand trajectories (top left) and 3D articulation labels (top right), it recovers the hinge of a laptop lid, a category seen only in human video, while other 3D-supervised methods fail.

Training signal from human video

Learning 3D articulation remains limited by the availability of real-world 3D annotations. Existing approaches rely on point clouds, RGB-D scans, CAD models or explicitly annotated articulation parameters, which are expensive to acquire and cover a narrow range of scenes and object categories. Models trained on such data generalize poorly to the diversity of objects encountered in daily life.

Egocentric video offers broader supervision. Datasets such as EPIC-KITCHENS, HOI4D and ARCTIC capture people interacting with diverse objects at a scale that 3D annotation cannot match, but they provide only RGB observations and 2D motion cues. EgoArt connects the two through an analytic trajectory generator: the predicted mask, interaction point, motion type, axis, hinge and extent define a 3D interaction trajectory, so a 2D hand trajectory and a 3D articulation label supervise the same parameters.

Records mined from human video

Each record is one frame with the mask of the actionable part (red), the track of the acting hand's knuckle over the interaction (green; ring at the interaction point, arrowhead at the end), a motion-type label and an instruction. These are mined automatically from three sources.

HOI4D

Table-top interactions with household objects
Open the laptop screen
“Open the laptop screen”revolute
Turn on the lamp
“Turn on the lamp”prismatic
Open the right cabinet door
“Open the right cabinet door”revolute
Close the laptop screen
“Close the laptop screen”revolute
Close the upper drawer
“Close the upper drawer”prismatic
Close the lamp arm
“Close the lamp arm”revolute
Push the pink toy car
“Push the pink toy car”prismatic
Close the lower fabric drawer
“Close the lower fabric drawer”prismatic
Close the lower cabinet door
“Close the lower cabinet door”revolute
Pull the small green toy car
“Pull the small green toy car”prismatic

EPIC-KITCHENS

Kitchen activity recorded from a head-mounted camera
Open drawer
“Open drawer”prismatic
Open fridge
“Open fridge”revolute
Close drawer
“Close drawer”prismatic
Close cupboard
“Close cupboard”revolute
Open fridge
“Open fridge”revolute
Open drawer
“Open drawer”prismatic
Open drawer
“Open drawer”prismatic
Open cupboard
“Open cupboard”revolute
Open fridge
“Open fridge”revolute
Open drawer
“Open drawer”prismatic

ARCTIC

Two-hand manipulation of small articulated objects
Close the waffle iron lid
“Close the waffle iron lid”revolute
Close the phone phone
“Close the phone phone”revolute
Close the laptop screen
“Close the laptop screen”revolute
Close the notebook cover
“Close the notebook cover”revolute
Close the capsule machine lever
“Close the capsule machine lever”revolute
Close the laptop screen
“Close the laptop screen”revolute
Open the phone phone
“Open the phone phone”revolute
Open the box lid
“Open the box lid”revolute
Open the scissors blades
“Open the scissors blades”revolute
Close the espresso machine lever
“Close the espresso machine lever”revolute

Abstract

Manipulating articulated objects requires identifying the relevant part and predicting how it can move. We present EgoArt, an instruction-conditioned model that grounds the task-relevant part from a single RGB image and predicts its 3D articulation affordance, including the mask, interaction point, motion type, axis, hinge, and motion extent. These predictions directly define a 3D interaction trajectory.

To reduce reliance on limited 3D annotations, we bridge 2D human hand trajectories and 3D articulation through an analytic trajectory generator, allowing egocentric videos without 3D labels to supervise explicit articulation parameters. The same formulation also supports joint training with 3D-labelled data, for which we derive a closed-form trajectory loss that requires neither explicit trajectory sampling nor annotated motion extent.

Trained jointly on SceneFun3D, HOI4D, EPIC-KITCHENS, and ARCTIC, EgoArt improves axis and hinge estimation over prior image-based articulation methods while generalizing to object categories seen only in 2D human video.

Method

One image and one instruction go in. A language-conditioned encoder produces a feature map; prediction heads estimate the part mask, the interaction point, the motion type, the axis, the hinge and the motion extent. An analytic decoder turns these parameters into the interaction trajectory, so a 2D hand track and a 3D annotation supervise the same quantities.

Overview of EgoArt
A language-conditioned encoder produces a feature map, from which articulation prediction heads estimate the part mask, interaction point, motion type, axis, hinge, and motion extent. An analytic decoder renders the corresponding interaction trajectory, enabling supervision from both 2D human hand tracks and 3D articulation annotations.

Real-robot demonstrations

EgoArt runs on a single frame from the gripper camera of a Boston Dynamics Spot, about a metre from the object, with the instruction shown. The robot walks to the predicted part, re-localizes the handle up close, and executes the predicted trajectory with its arm. Left: external view with the predicted trajectory (cyan) and axis (pink) drawn in. Right: the gripper camera. Clips play at 2x from the start of the reach to the end of the articulation.

“Open the cabinet door with the vertical handle”revolute
“Pull out the top drawer with the horizontal handle”prismatic
“Open the left cabinet door with the vertical handle”revolute
“Pull out the drawer in the middle with the horizontal handle”prismatic
“Open the right cabinet door”revolute
“Pull out the drawer in the bottom with the horizontal handle”prismatic

Results

Each card shows the ground truth next to one method; pick the method below the images. Red: predicted part. Yellow: axis, drawn as the hinge line with its foot for revolute parts and as the direction of travel for prismatic ones. Green: motion of the interaction point (white) over the extent. Trajectories appear only for the ground truth, EgoArt and the SceneFun3D-only control, since the other baselines do not predict an interaction point; none of them predicts motion extent.

SceneFun3D test frames

The 3D-labelled benchmark. Baselines retrained on the same split; detectors show their oracle-matched instance.

“Close the door next to the nightstand”revolute
Ground truth
Ground truth
EgoArt
EgoArt
“Open the top oven door”revolute
Ground truth
Ground truth
EgoArt
EgoArt
“Open the top left drawer of the wooden cabinet to the left of the bed”prismatic
Ground truth
Ground truth
EgoArt
EgoArt
“Open the right drawer of the wooden cabinet located next to the radiator”prismatic
Ground truth
Ground truth
EgoArt
EgoArt

Hand-video test frames

Held-out records from the human-video sources; the hand track is the ground-truth motion. "SceneFun3D-only" is the same model trained without human video. Categories absent from SceneFun3D are marked.

“Open the laptop screen”ARCTIC · category not in SceneFun3Drevolute
Ground truth
Ground truth
EgoArt
EgoArt
“Open the microwave door”ARCTIC · category in SceneFun3Drevolute
Ground truth
Ground truth
EgoArt
EgoArt
“Open drawer”EPIC-KITCHENS · category in SceneFun3Dprismatic
Ground truth
Ground truth
EgoArt
EgoArt
“Pull the toy car”HOI4D · category not in SceneFun3Dprismatic
Ground truth
Ground truth
EgoArt
EgoArt
“Close the upper drawer”HOI4D · category in SceneFun3Dprismatic
Ground truth
Ground truth
EgoArt
EgoArt

Phone photos outside every training set

Rooms that appear in neither SceneFun3D nor the hand-video sources. The photo is on the left; choose the model on the right. On the laptop, a category absent from SceneFun3D, only EgoArt finds the screen and calls it revolute; the SceneFun3D-only model predicts a slide on the keyboard.

“Move the mouse forward”prismatic
Photo
Photo
EgoArt
EgoArt
“Close the laptop”revolute
Photo
Photo
EgoArt
EgoArt
“Open the right closet”revolute
Photo
Photo
EgoArt
EgoArt
“Close the window”revolute
Photo
Photo
EgoArt
EgoArt

Citation

@inproceedings{egoart2027,
  title     = {EgoArt: Learning 3D Articulation Affordances from Egocentric Human Interaction Videos},
  author    = {Andrew Sangwoo Ye and Jiaqi Chen and Marc Pollefeys and Taein Kwon},
  booktitle = {Under review},
  year      = {2027}
}