← Back to journal
Journal of Virtual Convergence Research · Vol. 2, No. 3, July 2026

Robust 3D Human Pose Estimation against Occlusion using Cross-View Geometric Consistency and Temporal Interpolation

Subin Park

Published: 2026-07

Pages
1~22

Abstract

Accurate 3D human pose estimation in the Metaverse is often limited by occlusion and the difficulty of collecting 3D ground-truth (GT) data. We propose a markerless, annotation-free framework that needs no 3D GT or model training, using cross-view geometric consistency between calibrated stereo cameras as its only supervisory cue. The pipeline extracts 2D poses per view, matches people across views, and triangulates 3D skeletons. When occlusion makes triangulation fail, a temporal interpolation module recovers the missing 3D joints from neighboring frames within a 7-frame window. The system reached a 99.2% overall reconstruction success rate; under severe occlusion, the complete-skeleton ratio rose from 72% to 99%, recovering 87–92% of joints triangulation alone could not. As a feasibility study on a deliberately challenging two-person crossing-occlusion scenario, the results show that reliable 3D motion capture is achievable from only 2D detections and temporal constraints—without 3D supervision or training—making it suitable for interactive avatars and performances in the Metaverse.

Keywords

3D Human Pose EstimationMarkerlessOcclusion RobustnessAnnotation-Free ReconstructionGeometric ConsistencyMetaverse

Full text

Full text is not available yet.

Cite as

Subin Park. (2026). Robust 3D Human Pose Estimation against Occlusion using Cross-View Geometric Consistency and Temporal Interpolation. Journal of Virtual Convergence Research, 2(3), 1~22.