Method for acquiring 2D hand image joint angle in real time in combination with MANO parameterized model
Through the dual-branch model and efficient feature aggregator design based on the MANO parameterized model, combined with multiple loss functions, the problem of insufficient real-time and accuracy in the existing technology is solved, and efficient real-time hand joint angle recognition is achieved on low-computing equipment.
Patent Information
- Application Number
- CN202510084828.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has shortcomings in terms of real-time and accuracy in processing, and it is difficult to meet the needs of real-time hand joint angle recognition, especially in the case of occlusion, depth blur and noise interference.
The dual-branch model and efficient feature aggregator design based on MANO parameterized model are adopted, combined with multiple loss functions for joint optimization, which significantly improves the accuracy of joint angle prediction, and supports deployment on low-computing equipment through lightweight design.
The method of obtaining joint angles of 2D hand images in real time is realized, which significantly improves the accuracy of joint angle prediction and operates efficiently on low-computing equipment, supporting integration into existing gesture recognition systems.
Smart Images

Figure CN120071391A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method for real-time obtaining joint angles of 2D hand images based on the MANO parametric model. Background Technique
[0002] In the fields of virtual reality, human-computer interaction, teleoperation, and rehabilitation medicine, gesture recognition and three-dimensional reconstruction have become research hotspots. At present, hand joint angles are mainly recognized through images or pose electrical signals. The specific implementation processes of the above two methods are quite different. However, with the development of deep learning technology and the improvement of computing power, it has been possible to obtain accurate hand joint angles in real time through images.
[0003] Compared with the method based on pose electrical signals, image recognition technology does not require any sensors or devices to be installed on the human body, and it is a non-invasive measurement method. It can be applied to a variety of scenarios, including virtual reality, augmented reality, gesture control, rehabilitation medicine, etc. Because it does not require specific hardware support, it has better versatility and flexibility.
[0004] However, the existing technologies still have deficiencies in terms of real-time performance and accuracy. Traditional methods rely on complex model structures and high-cost computing resources, and it is difficult to meet the real-time requirements. In addition, problems such as occlusion, depth blur, and noise interference further increase the difficulty of gesture reconstruction.
[0005] Based on the existing problems, the present invention proposes a method for real-time obtaining joint angles of 2D hand images based on the MANO parametric model. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technology, the present invention proposes a method for real-time obtaining joint angles of 2D hand images based on the MANO parametric model. First, through the design of a double-branch model and an efficient feature aggregator, the present invention enhances the real-time performance of obtaining 2D hand image joints; secondly, introducing a combination of multiple loss functions for joint optimization significantly improves the accuracy of joint angle prediction; finally, the present invention adopts a lightweight design, greatly reducing the number of network parameters, supporting deployment on low-computing-power devices, and the optimized network structure ensures efficient operation while maintaining the prediction accuracy. It can be integrated into existing gesture recognition systems or other human-computer interaction-based applications, saving development costs.
[0007] The present invention mainly includes the following steps:
[0008] S1: Data collection and processing.
[0009] S2: Building a training model based on the MANO parametric model.
[0010] S3: Transform the hand joint angles in combination with the MANO parametric model.
[0011] In step S1, specifically:
[0012] S101: The dataset of the present invention uses the existing dataset FreiHAND_pub_v2 and the actually collected data to enhance the generalization ability of the model. For the actually collected data, the following operations are performed on the data.
[0013] S102: Data cleaning: Identify and delete duplicate images by comparing the hash values or feature vectors of the images, delete low-quality, non-standard, and unreasonable images, and repair problems such as missing, noise, and artifacts in the data.
[0014] S103: Format alignment: Ensure that all images have the same size and number of channels. If the sizes and numbers of channels of the collected images are inconsistent, adjustments are required to convert the images into a format suitable for model input and align them with the FreiHAND_pub_v2 format.
[0015] S104: Data annotation: Annotate the objects or regions in the images to provide necessary information for subsequent training and evaluation.
[0016] S105: Data augmentation: Increase the diversity of the collected data through operations such as rotation, translation, scaling, and flipping to improve the generalization ability of the model.
[0017] S106: Combine the processed data and the FreiHAND_pub_v2 dataset to form the dataset used in the present invention, and divide the training set and the test set.
[0018] In step S2, specifically:
[0019] : S201: For the left and right hands, implement a dual-branch network model architecture, and use the HRNet-W32 network as the backbone network for feature extraction; HRNet-W32 mainly includes three modules: a multi-resolution parallel module, a feature fusion module, and a residual module.
[0020] S202: In the multi-resolution parallel module, initially extract features from the input image through a standard convolutional network to generate a preliminary high-resolution feature map; then, maintain the high-resolution feature map through multiple convolutional layers. These high-resolution feature maps maintain a large spatial resolution and can capture the detailed information of the image. Finally, obtain and process the low-resolution feature map through pooling or convolutional downsampling. The low-resolution branch helps capture more context information.
[0021] S203: In the feature fusion module, feature maps of different resolutions are fused. The low-resolution features are merged into the high-resolution branch through upsampling, and the high-resolution features are propagated to the low-resolution branch through downsampling, enhancing the complementarity between features and enabling the network to capture details and context information at various scales.
[0022] S204: In the residual module, the input information is added to the information after convolutional processing to form the final output. At each stage, HRNet-W32 allows information to propagate efficiently in the network through residual connections, avoiding the vanishing gradient problem and helping the network extract effective features at different levels, improving the training efficiency and convergence speed.
[0023] S205: Generate parameter feature maps, hand center feature maps, and partial segmentation feature maps based on the HRNet-W32 network processing.
[0024] S206: The above feature maps are respectively pre-processed through different MLPs (Multi-Layer Perceptrons). Each MLP network transforms the input image or parameters into a low-dimensional vector, which represents the high-level semantic information of the image or parameters. After being processed by the MLP, the obtained features can describe various geometric or pose information of the hand and are transformed into a parameterized vector suitable for the MANO model to process.
[0025] S207: Input the above parameterized vector into the MANO model to obtain the loss function of the hand mesh quality.
[0026] S208: For the hand center feature map and partial segmentation feature map generated in step S205, the center attention loss and partial attention loss are respectively introduced to optimize the positioning of the hand center map and improve the regional accuracy of the segmentation map. Finally, the loss function of the hand mesh quality of the MANO model and the above center attention loss and partial attention loss functions are processed by the feature aggregator for the calculation of PA-MPJPE and PA-MPVPE metrics.
[0027] S209: Rotate and translate the estimated joint positions and the true joint positions through Procrustes alignment, then calculate the error of each joint and take the average value to obtain PA-MPJPE. PA-MPJPE represents the joint position error, and its calculation formula is:
[0028] , where, represents the estimated joint position, represents the estimated true joint position.
[0029] The estimated mesh vertices and the true mesh vertices are matched through Procrustes alignment. Then, the error of each vertex is calculated and averaged to obtain PA-MPVPE. PA-MPVPE is an extension of PA-MPJPE and represents the joint position error in the predicted mesh. Its calculation formula is:
[0030] , where represents the predicted mesh vertex coordinates, represents the true mesh vertex coordinates.
[0031] In step S3, it can be specifically described as follows:
[0032] S301: According to the MANO model, the 3D coordinates of 21 joint points are output. According to the given three points , and calculate the angle between the two vectors: .
[0033] S302: Calculate the dot product of the vectors: .
[0034] S303: Calculate the 2-norm of each vector: .
[0035] S304: Calculate the cosine value of the angle according to the cosine value and the 2-norm: .
[0036] S305: Convert the cosine value to an angle according to the inverse cosine function: , which is the desired joint angle.
[0037] The present invention has the following advantages:
[0038] Improved real-time performance: Through the design of a dual-branch model and an efficient feature aggregator, real-time processing is achieved.
[0039] Improved accuracy: By jointly optimizing multiple loss functions, the accuracy of joint angle prediction is significantly improved.
[0040] Lightweight Design: By using the LocallyConnected2d layer in the contact_layers module, the number of parameters is reduced compared to the fully connected layer while maintaining the spatial hierarchy, thus reducing the computational overhead. Without sacrificing performance, the model becomes more lightweight. The segmentation-based attention mechanism utilizes SegmNet to generate hand segmentation masks and creates attention maps. By focusing on the relevant parts of the image, the model avoids unnecessary computations, thereby improving efficiency. The head layer of the model adopts BasicBlock, which has a simpler structure and less computational complexity, further reducing the computational load. The number of network parameters is significantly reduced, enabling deployment on low-computing-power devices. The optimized network structure ensures efficient operation while maintaining prediction accuracy.
[0041] Good Scalability: It can be integrated into existing gesture recognition systems or other human-computer interaction-based applications, saving development costs. Brief Description of the Drawings
[0042] Figure 1 This is the flowchart of the present invention.
[0043] Figure 2 This is the architecture diagram of the parametric hand feature extraction model combined with MANO.
[0044] Figure 3 This is the comparison diagram of the training results of the parametric hand feature extraction model combined with MANO. Detailed Embodiments
[0045] To enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] Referring to the attached Figure 1 , the specific embodiments of the present invention are as follows:.
[0047] S1: Data collection and processing.
[0048] S2: Build and train a model based on the MANO parametric model.
[0049] S3: Perform hand joint angle conversion in combination with the MANO parametric model.
[0050] In step S1, specifically:
[0051] S101: The dataset of the present invention uses the existing dataset FreiHAND_pub_v2 and the actually collected data to enhance the generalization ability of the model. For the actually collected data, the following operations are performed on the data.
[0052] S102: Data cleaning: Identify and delete duplicate images by comparing the hash values or feature vectors of the images, delete low-quality, non-standard, and unreasonable images, and repair problems such as missing, noise, and artifacts in the data.
[0053] S103: Format alignment: Ensure that all images have the same size and number of channels. If the sizes and numbers of channels of the collected images are inconsistent, adjustments are required to convert the images into a format suitable for model input and align them with the FreiHAND_pub_v2 format.
[0054] S104: Data annotation: Annotate the objects or regions in the images to provide necessary information for subsequent training and evaluation.
[0055] S105: Data augmentation: Increase the diversity of the collected data through operations such as rotation, translation, scaling, and flipping to improve the generalization ability of the model.
[0056] S106: Combine the processed data and the FreiHAND_pub_v2 dataset to form the dataset used in the present invention, and divide the training set and the test set.
[0057] Refer to the appendix Figure 2 , in step S2, specifically:
[0058] S201: For the left and right hands, implement a dual-branch network model architecture, and use the HRNet-W32 network as the backbone network for feature extraction; HRNet-W32 mainly includes three modules: a multi-resolution parallel module, a feature fusion module, and a residual module.
[0059] S202: In the multi-resolution parallel module, initially extract features from the input image through a standard convolutional network to generate a preliminary high-resolution feature map; then, maintain the high-resolution feature map through multiple convolutional layers. These high-resolution feature maps maintain a large spatial resolution and can capture the detailed information of the image. Finally, obtain and process the low-resolution feature map through pooling or convolutional downsampling. The low-resolution branch helps capture more context information.
[0060] S203: In the feature fusion module, feature maps of different resolutions are fused. The low-resolution features are merged into the high-resolution branch through upsampling, and the high-resolution features are propagated to the low-resolution branch through downsampling, enhancing the complementarity between features and enabling the network to capture details and context information at all scales.
[0061] S204: In the residual module, the input information is added to the information after convolutional processing to form the final output. At each stage, HRNet-W32 allows information to propagate efficiently in the network through residual connections, avoiding the vanishing gradient problem and helping the network extract effective features at different levels, improving the training efficiency and convergence speed.
[0062] S205: Generate parameter feature maps, hand center feature maps, and partial segmentation feature maps according to the HRNet-W32 network processing.
[0063] S206: The above feature maps are respectively preprocessed through different MLPs (Multi-Layer Perceptrons). Each MLP network converts the input image or parameters into a low-dimensional vector, which represents the high-level semantic information of the image or parameters. After being processed by the MLP, the obtained features can describe various geometric or pose information of the hand and be converted into a parameterized vector suitable for processing by the MANO model.
[0064] S207: Input the above parameterized vector into the MANO model to obtain the loss function of the hand mesh quality.
[0065] S208: For the hand center feature map and partial segmentation feature map generated in step S205, introduce the center attention loss and partial attention loss respectively to optimize the positioning of the hand center map and improve the regional accuracy of the segmentation map. Finally, process the loss function of the hand mesh quality of the MANO model and the above center attention loss and partial attention loss functions through a feature aggregator for the calculation of PA-MPJPE and PA-MPVPE metrics.
[0066] S209: Rotate and translate the estimated joint positions and the true joint positions for alignment through Procrustes alignment, then calculate the error of each joint and take the average value to obtain PA-MPJPE. PA-MPJPE represents the joint position error, and its calculation formula is:
[0067] , where represents the estimated joint position, represents the estimated true joint position.
[0068] The estimated mesh vertices and the true mesh vertices are matched through Procrustes alignment. Then, the error of each vertex is calculated and averaged to obtain PA-MPVPE. PA-MPVPE is an extension of PA-MPJPE and represents the joint position error in the predicted mesh. Its calculation formula is:
[0069] , where represents the predicted mesh vertex coordinates, represents the true mesh vertex coordinates.
[0070] After multiple rounds of training, the model converges, and its training results are compared with those of related models as shown in the appendix Figure 3 as shown.
[0071] In step S3, it can be specifically described as:
[0072] S301: According to the MANO model, the 3D coordinates of 21 joint points are output. According to the given three points , and calculate the angle between the two vectors: .
[0073] S302: Calculate the dot product of the vectors: .
[0074] S303: Calculate the 2-norm of each vector: .
[0075] S304: Calculate the cosine value of the angle according to the cosine value and the 2-norm: .
[0076] S305: Convert the cosine value to an angle according to the inverse cosine function: , which is the required joint angle.
[0077] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for real-time acquisition of 2D hand image joint angles combined with MANO parameterized model, characterized in that: The method for real-time acquisition of 2D hand image joint angles in combination with the MANO parameterized model comprises the following steps: S1: Data collection and processing; S2: Build a training model based on the MANO parameterized model; S3: Combined with MANO parameterized model to transform hand joint angles; In step S1, a method of mixing existing data sets and actual collected data is adopted, and data preprocessing is performed on the actual collected data to increase the generalization ability of the model; in step S2, the HRNet-W32 lightweight model is used for feature extraction, and the central attention loss and partial attention loss are introduced for joint optimization to significantly improve the accuracy of joint angle prediction; in step S3, the 3D coordinates of 21 joint points are output according to the MANO model. Given three points ,and Compute the angle between two vectors: ; Compute the dot product of vectors: ; Calculate the 2-norm of each vector: ; Calculate the cosine of the angle based on the cosine value and the 2-norm: ; Convert the cosine value to an angle using the arccosine function: .