A monocular image sequence-based 3D hand pose estimation method and system
By combining a monocular camera with LSNetPose2D and OverLoCKGraphMLP architecture, the problems of high-cost hardware dependence, lack of depth information and dynamic motion blur are solved, and high-precision 3D hand pose estimation is achieved, which is suitable for human-computer interaction scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JILIN INST OF CHEM TECH
- Filing Date
- 2026-02-25
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, 3D hand pose estimation relies on high-cost hardware, monocular vision lacks depth information, dynamic motion blur and occlusion affect robustness, and the model's generalization ability is insufficient, leading to inaccurate estimation.
Image sequences are acquired using a monocular camera. The LSNetPose2D lightweight network and the OverLoCKGraphMLP multi-stage hybrid architecture are combined to achieve end-to-end 3D hand coordinate regression through a 3D pose regression module. The composite loss function is used to constrain anatomical rationality, and the AdamW optimizer and learning rate scheduling are combined to improve training stability.
It improves the accuracy and robustness of 3D hand pose estimation, reduces distortion, and enhances the generalization performance of the model, making it suitable for human-computer interaction scenarios such as robotic arm control.
Smart Images

Figure CN122116474A_ABST