ECCV 2026 (Oral)

Interaction-Aware 4D Gaussian Splatting for
Dynamic Hand-Object Interaction Reconstruction

Hao Tian1,2,5   Chenyangguang Zhang3   Rui Liu4   Wen Shen1,5   Xiaolin Qin1,5 ✉

1 Chengdu Institute of Computer Applications, CAS
2 PetroChina (Beijing) Digital Intelligence Research Institute
3 Tsinghua University  |  4 Minzu University of China  |  5 University of Chinese Academy of Sciences

Paper arXiv Code Video
Teaser: previous methods vs. ours on HOI reconstruction

Figure 1. Previous methods rely on a single implicit deformation field with only 2D rendering supervision, leading to blurry and inconsistent reconstructions. Our Interaction-Aware 4D-GS introduces separate Hand / Object deformation fields coupled by explicit interaction parameters and 2D&3D joint supervision, yielding sharper, physically plausible results.

Abstract

This paper focuses on a challenging setting of simultaneously modeling geometry and appearance of hand-object interaction scenes without any object priors. We follow the trend of dynamic 3D Gaussian Splatting based methods, and address several significant challenges. To model complex hand-object interaction with mutual occlusion and edge blur, we present interaction-aware hand-object Gaussians with newly introduced optimizable parameters aiming to adopt a piecewise linear hypothesis for clearer structural representation. Moreover, considering the complementarity and tightness of hand shape and object shape during interaction dynamics, we incorporate hand information into the object deformation field, constructing interaction-aware dynamic fields to model flexible motions. To further address difficulties in the optimization process, we propose a progressive strategy that handles dynamic regions and static background step by step. Correspondingly, explicit regularizations are designed to stabilize the hand-object representations for smooth motion transition, physical interaction reality, and coherent lighting. Experiments show that our approach surpasses existing dynamic 3D-GS-based methods and achieves state-of-the-art performance in reconstructing dynamic hand-object interaction.

Presentation Video

Method

Pipeline of Interaction-Aware 4D Gaussian Splatting

Figure 2. Pipeline overview. Hand Gaussians and Object Gaussians are separately fed into their neural deformation fields (Hand Field / Object Field) parameterized by position encoding γ(ρ), rotation R, and timestamp t. Interaction-Aware Parameters couple the two streams. Deformed Gaussians are composited with Background Gaussians and rasterized. Supervision comes from five losses: ℒrender, ℒα, ℒinteraction, ℒHtrans, ℒOrot.

👐 Interaction-Aware Parameters

Each Gaussian carries a learnable weight w (interaction strength) and radius o (influence extent). These parameters let the hand and object fields communicate without explicit contact labeling, naturally capturing near-contact and in-contact dynamics.

🏐 Dual Neural Deformation Fields

Separate Hand Field and Object Field each take position-encoded Gaussian centers, rotations, and timestamps, and output per-Gaussian deformation offsets. Decoupled fields prevent entanglement while the interaction loss enforces consistency at boundaries.

🎯 2D & 3D Joint Supervision

Beyond standard photometric and alpha losses, we add 3D geometric constraints: a hand translation loss guides absolute hand position, and an object rotation loss leverages MANO average-rotation priors, substantially reducing degenerate solutions.

🌟 Background Disentanglement

A dedicated Background Gaussian set is maintained separately and composited during rasterization. This enables clean decomposed rendering 鈥?hand, object, and scene can be visualized or rendered independently without re-training.

Decomposed Rendering

Decomposed rendering: canonical, target, background

Figure 3. Our Gaussian representation supports explicit decomposition. From top to bottom: Ground-truth frames, canonical-space rendering, target-space rendering (after deformation), and background-only rendering. The clean separation enables downstream analysis and virtual object insertion.

Qualitative Comparison

Qualitative comparison on HOI4D, HO3D and HOLD vs 4DGS, Deform3DGS, SC-GS

Figure 4. Qualitative comparison on HOI4D (top), HO3D (middle), and HOLD vs. Ours on HO3D (bottom) against GT, 4DGS, Deform3DGS, and SC-GS. Our method consistently produces sharper textures, more accurate hand-object boundaries, and significantly fewer ghosting or blurring artifacts across all three benchmarks.

Novel View Synthesis

Real-time rendering from our trained 4D Gaussian representation across multiple sequences.

HOI4D RGB novel view

HOI4D — RGB Rendering

HOI4D Depth novel view

HOI4D — Depth Rendering

3D Wave RGB novel view

3D Wave — RGB Rendering

3D Wave Depth novel view

3D Wave — Depth Rendering

Translation sequence

Free-Viewpoint Translation

BibTeX

@inproceedings{tian2026interaction, title={Interaction-Aware 4D Gaussian Splatting for Dynamic Hand-Object Interaction Reconstruction}, author={Tian, Hao and Zhang, Chenyangguang and Liu, Rui and Shen, Wen and Qin, Xiaolin}, booktitle={European Conference on Computer Vision}, pages={405--424}, year={2026}, organization={Springer} }