ECHO: Ego-Centric modeling of
Human-Object interactions

ECCV 2026, Malmö

Paper arXiv Video Code (coming soon) BibTeX

TL;DR

ECHO reconstructs full-body human-object interactions from sparse egocentric tracking of the head and wrists.

👓🙌

Sparse egocentric input

First method to recover full-body HOI from head and wrist tracking alone.

🎲

Tri-variate diffusion

Human, object, and contact modeled jointly with independent noise schedules, allowing flexible input configurations.

🏃📦

Mixed-data training

Large-scale motion capture (AMASS) provides a strong motion prior; dedicated HOI datasets add interaction knowledge.

Abstract

Modeling human-object interactions (HOI) from an egocentric perspective is a critical yet challenging task, particularly when relying on sparse signals from wearable devices like smart glasses and watches. We present ECHO, the first unified framework to jointly recover human pose, object motion, and contact dynamics solely from head and wrist tracking. To tackle the underconstrained nature of this problem, we introduce a novel tri-variate diffusion process with independent noise schedules that models the mutual dependencies between the human, object, and interaction modalities. This formulation allows ECHO to operate with flexible input configurations, making it robust to intermittent tracking and capable of leveraging partial observations. Crucially, it enables training on a combination of large-scale human motion datasets and smaller HOI collections, learning strong priors while capturing interaction nuances. Furthermore, we employ a smooth inpainting inference mechanism that enables the generation of temporally consistent interactions for arbitrarily long sequences. Extensive evaluations demonstrate that ECHO achieves state-of-the-art performance, significantly outperforming existing methods lacking such flexibility.

Goal

Method

ECHO is a transformer-based diffusion model that predicts Human motion, Object pose, and Contacts given three-point tracking and the object conditioning. At its core is a tri-variate diffusion process with an independent denoising schedule per modality: setting the noise level to zero for an observed modality turns it into a condition, so a single network supports flexible input configurations: from pure head-and-wrist tracking to additional partial observations of the human or the object. The model operates in a per-frame head-centric coordinate system and generates arbitrarily long sequences via smooth inpainting, which blends past and current window predictions at every diffusion step, enabling online inference.

Results

ECHO is trained and evaluated on the union of BEHAVE, OMOMO, and AMASS, and achieves state-of-the-art performance in both human and object reconstruction. Where baselines often break contact with objects penetrating the body or floating mid-air, while ECHO produces physically plausible interactions.

Generalization to unseen motion and object from Aria Digital Twin:

Citation

@inproceedings{petrov2026echo,
  title={ECHO: Ego-Centric modeling of Human-Object interactions},
  author={Petrov, Ilya A and Guzov, Vladimir and Marin, Riccardo and Aksan, Emre and Chen, Xu and Cremers, Daniel and Beeler, Thabo and Pons-Moll, Gerard},
  booktitle={European Conference on Computer Vision},
  year={2026},
  organization={Springer}
}

Acknowledgments

Special thanks to Nikita Kister and Berna Kabadayi for the helpful discussions. This work is funded by the Deutsche Forschungsgemeinschaft - 409792180 (EmmyNoether Programme, project: Real Virtual Humans). G. Pons-Moll is a member of the Machine Learning Cluster of Excellence, EXC number 2064/1 - Project number 390727645. The authors thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting I. A. Petrov. R. Marin has been supported by the European Union's Horizon 2020 research and innovation program under the Marie Skłodowska-Curie grant agreement No 101109330. The project was made possible by funding from the Carl Zeiss Foundation. The computational resources for this project were provided by the Google Cloud grant. This work was supported by the European Research Council (ERC) Advanced Grant SIMULACRON and by the GNI Project “AI4Twinning”. The views and opinions expressed in this work are those of the authors and do not necessarily reflect the official policy or position of Google. The Google logo is a trademark of Google LLC. Its use does not imply endorsement of this work. Website is based on StyleGAN3 and Nerfies websites.

Carl-Zeiss-Stiftung
Tübingen AI Center
IMPRS-IS
University of Tübingen
MPII Saarbrücken
EU
Technical University of Munich
Munich Center for Machine Learning
Google Research