A monocular image sequence-based 3D hand pose estimation method and system

By combining a monocular camera with LSNetPose2D and OverLoCKGraphMLP architecture, the problems of high-cost hardware dependence, lack of depth information and dynamic motion blur are solved, and high-precision 3D hand pose estimation is achieved, which is suitable for human-computer interaction scenarios.

CN122116474APending Publication Date: 2026-05-29JILIN INST OF CHEM TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JILIN INST OF CHEM TECH
Filing Date
2026-02-25
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing technologies, 3D hand pose estimation relies on high-cost hardware, monocular vision lacks depth information, dynamic motion blur and occlusion affect robustness, and the model's generalization ability is insufficient, leading to inaccurate estimation.

Method used

Image sequences are acquired using a monocular camera. The LSNetPose2D lightweight network and the OverLoCKGraphMLP multi-stage hybrid architecture are combined to achieve end-to-end 3D hand coordinate regression through a 3D pose regression module. The composite loss function is used to constrain anatomical rationality, and the AdamW optimizer and learning rate scheduling are combined to improve training stability.

Benefits of technology

It improves the accuracy and robustness of 3D hand pose estimation, reduces distortion, and enhances the generalization performance of the model, making it suitable for human-computer interaction scenarios such as robotic arm control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116474A_ABST
    Figure CN122116474A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of computer vision and image processing, and particularly relates to a 3D hand posture estimation method and system based on monocular image sequences, which comprises the following steps: acquiring an original image sequence of hand postures by using a monocular camera and performing preprocessing to obtain a monocular image sequence; inputting the monocular image sequence into a 2D posture estimation module constructed by using a LSNetPose2D lightweight network to obtain a joint heat map sequence; inputting the joint heat map sequence into a 3D posture regression module constructed by using an OverLoCKGraphMLP multi-stage hybrid architecture to obtain 3D hand coordinates, and realizing end-to-end 3D hand coordinate regression. The application effectively improves the estimation accuracy and robustness of hand three-dimensional postures under monocular vision conditions by combining deep learning and spatial mapping calibration technology, and solves the inaccurate estimation problem caused by missing parallax information, motion blur and occlusion in the prior art.
Need to check novelty before this filing date? Find Prior Art