Hand posture reconstruction method based on multi-stage joint point enhancement
Through multi-stage node enhancement network, the hand posture is reconstructed by feature extraction and Transformer module, the accuracy and calculation cost of gesture recognition under complex background and occlusion conditions in the prior art is solved, and the efficient posture reconstruction effect is achieved.
Patent Information
- Application Number
- CN202510822793.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
Existing gesture recognition and reconstruction technologies are difficult to accurately segment the hand area and perform posture estimation under complex backgrounds and variable lighting conditions. The detection is incomplete or inaccurate when the finger or palm is blocked, and the calculation cost of high-resolution depth map or point cloud data processing is high.
The multi-stage joint node enhancement network is adopted. First, the feature learning and prediction of 12 fingertips and palm root joints are performed through the feature extractor and EABlock module, combined with the Transformer module to conduct joint inference on the remaining 30 joint nodes, and the final 42 joint node coordinates are output through weighted fusion.
It improves the accuracy and efficiency of hand posture reconstruction and reduces calculation costs. MPJPE performs better than the existing methods and achieves an accuracy of 7.48mm.
Smart Images

Figure CN120340139A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning, and particularly relates to a hand pose reconstruction method based on multi-stage joint point enhancement. Background Art
[0002] Gesture recognition and reconstruction are research hotspots in the fields of computer vision and human-computer interaction, aiming to understand and reconstruct hand poses and gestures through computer algorithms to achieve natural human-computer interaction. With the progress of deep learning, image processing, and 3D sensing technologies, gesture recognition and reconstruction technologies have been widely applied and developed in applications such as virtual reality, augmented reality, games, robot control, and sign language translation. Due to the highly articulated and self-occluding nature of the hand, achieving accurate estimation is extremely challenging. Early gesture recognition relied on manually extracted image features such as edges, contours, skin color, etc. These methods required manual design of feature extraction algorithms and were sensitive to factors such as lighting changes and occlusion, with poor robustness; another early method was to use gloves or devices with sensors to capture hand movements, and these sensors could detect finger bending and displacement. The disadvantage of this type of method is that the device is expensive and not natural enough, and it is not suitable for large-scale applications. Existing methods are based on 2D images or 3D depth maps or point clouds. However, for existing methods, there are still some problems: First, accurately segmenting the hand region and performing pose estimation under complex backgrounds and changing lighting conditions is very challenging; Second, when fingers or palms are occluded from each other during hand pose estimation, the detected key points may be incomplete or inaccurate, which has a greater impact on the reconstruction accuracy; Third, processing high-resolution depth maps or point cloud data and using complex deep learning models for pose estimation will bring higher computational costs. Summary of the Invention
[0003] Aiming at the defects existing in the prior art, the present invention provides a hand pose reconstruction method based on multi-stage joint point enhancement. The specific technical solutions are as follows:
[0004] An embodiment of the present application provides a hand pose reconstruction method based on multi-stage joint point enhancement, and the steps are as follows:
[0005] Dataset preprocessing. Feature learning and prediction are performed on 12 fingertip and palm root joints through the first-stage part of the multi-stage joint enhancement network.
[0006] Feature learning is performed on the remaining 30 hand joints through the second-stage part of the multi-stage joint enhancement network.
[0007] In the second stage of the multi-stage joint enhancement network, the 30 hand joints are predicted and inferred by the Transformer module.
[0008] Output the 12 fingertip and palm root joint features obtained in the first stage and the 30 joint points inferred by the Transformer in the second stage after weighted fusion.
[0009] Reconstruct the hand pose and conduct quantitative analysis.
[0010] In a possible implementation, for the original data that needs to be feature-extracted, flip it and randomly crop it to a size of 256×256. Use the preprocessed image as the input image for both the first stage part and the second stage part of the multi-stage joint enhancement network.
[0011] In a possible implementation, input the preprocessed image into the 12-joint point prediction training network structure of the first stage of the multi-stage joint enhancement network. The network structure consists of a feature extractor and an EABlock. First, the feature extractor extracts features from the input image. After completing the feature extraction, the features are passed into the EABlock module to obtain adaptive interaction features. Then, the adaptive interaction features are connected with the corresponding hand features to generate the final adaptive hand features.
[0012] Subsequently, the adaptive hand features are respectively passed into the joint feature extractor and the joint enhancer. First, rough joint features are extracted by the joint feature extractor and used as the input of the joint enhancer to obtain the final joint features and predicted joint positions. Thus, the joint features and predicted joint positions of 10 fingertip joints and 2 palm root joints are obtained.
[0013] In a possible implementation, the EABlock module includes two stages, namely the extraction stage and the adaptation stage. The FuseFormer module in the extraction stage uses the left hand features and the right hand features to extract interaction features. Then, the interaction features and the hand features are combined and used as two groups of inputs to be passed into the adaptation stage. The FuseFormer modules of the two branches in the adaptation stage respectively process the two groups of inputs to generate the adaptive interaction features of the left hand and the adaptive interaction features of the right hand.
[0014] In a possible implementation, in the first stage, receive the preprocessed training data during training and conduct training before the start of the second stage training. The loss function used during training is MSE.
[0015] The training batch size is set to 40 epochs. After outputting the prediction results through forward propagation, calculate the loss using the above loss function, and then perform backpropagation. Use the adam optimizer, and the initial learning rate is set to When the loss saturates, the learning rate drops by a factor of ten.
[0016] In a possible implementation, the network structure of the second stage consists of an EABlock module, a MANO module, and a Transformer module. In the second stage, using the hand features extracted from the image, they are concatenated with the joint features of the 12 joint points predicted in the first stage, and the concatenated feature vector is input into the MANO module. Using the 12 joint coordinates output in the first stage as constraint conditions, the remaining 30 joint positions are inferred through the MANO parametric model.
[0017] Similar to the first stage, the EABlock module is used to extract hand features from the input image (i.e., the preprocessed training data), and joint feature extraction and joint feature enhancement are performed on the 30 extracted hand joint points. Different from the first stage, the 30 joint features predicted in this stage do not need to output the corresponding joint point coordinates.
[0018] In a possible implementation, in the Transformer module of the second stage, through the multi-head self-attention mechanism, the interdependent relationships between different joint points, as well as the associations between joint point geometric constraints and image features, are captured, thereby improving the prediction accuracy. The interdependent relationships between the 12 joint points and the 30 joint points to be predicted are captured through the multi-head self-attention mechanism, and at the same time, the context information extracted from the image features is combined. Finally, the enhanced 30 joint features are output.
[0019] In a possible implementation, the input data used during the training of the second stage is the preprocessed training data and the output of the first stage of the multi-stage joint enhancement network. The loss function used for training is as follows:
[0020]
[0021] where is the number of joint points in the second stage, is the position of the other joint points except the fingertips and the wrist predicted in the second stage, is the corresponding true value.
[0022]
[0023] where and are hyperparameters, and C is a constant used to make the values of the two cases continuous at |x| = w.
[0024] The comprehensive loss function of the second stage is:
[0025]
[0026] where and is the weight for balancing the losses in two stages.
[0027] In a possible implementation, through the feature fusion module, the 12 joint features output by the first stage are weighted and spliced with the enhanced 30 joint features output by the second stage, and the final 42 hand joint coordinates are obtained through regression by the MLP module.
[0028] In a possible implementation, the hand pose regression module is used to regress the final 42 hand joint features to reconstruct the hand pose, and the experimental results are quantitatively analyzed in combination with the MPJPE evaluation index.
[0029] The beneficial effects of the present invention are as follows:
[0030] 1. The fingertip error obtained by separately training the fingertips in the present invention is lower than that obtained by overall training, and thus a new two-stage joint point enhancement network is proposed.
[0031] 2. The present invention proposes a bimodal enhancement structure, which combines geometric constraints and visual constraints in the second stage for joint feature reasoning, thereby enhancing the accuracy of joint point prediction.
[0032] 3. Specifically, the final effect of the present invention has better performance compared with the existing methods. In terms of MPJPE performance, 7.48 mm of the present invention is better than 8.79 mm of IntagHand, 13.48 mm of Zhang et al., and 16.00 mm of Moon et al. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0034] Figure 1 is the flowchart of the method in the embodiment of the present invention.
[0035] Figure 2 is a schematic diagram of the overall network structure adopted in the embodiment of the present invention.
[0036] Figure 3 is a schematic diagram of the EABlock module structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0038] The present invention proposes a unique multi-stage joint point enhancement technology for determining the complete hand shape and pose from monocular RGB images. In the present invention, it is proposed that when only training 10 fingertip joints, the task of the network focuses on a specific subset. Due to the reduction in the number of joints, the network can focus on learning the features of the fingertips without having to process the information of other joints such as the palm and wrist at the same time. Such a training task is relatively simpler. Therefore, the network is more likely to reach a lower loss value. When training all 42 joints at the same time, the network needs to learn the features of all joints and optimize the pose estimation of the entire hand. This task is more complex, and the loss function needs to consider the error distribution of all joints, which may cause the learning effect of some joints to be interfered by the losses of other joints. Therefore, the individual loss of the obtained fingertip joints may be higher than the loss during focused training. In summary, the present invention proposes a multi-stage joint enhancement network, namely MSJENet. The first stage consists of a network that only trains the positions of the fingertip joints and the position of the palm root joint; the second stage combines the input image features obtained from the feature extractor and the 12 joint points output by the first stage, and performs joint reasoning through the Transformer module and the features extracted from the image to enhance the accuracy of the prediction of the remaining 30 joint positions. The main idea of this design is to combine the visual modality (image features) and the geometric modality (12 joint points), and improve the model's prediction ability for the remaining 30 joint positions through multi-modal fusion.
[0039] To achieve the above objectives, we first need to preprocess the input data, and then input the preprocessed images into the fingertip joint predictor in the first stage. After extracting features through the ResNet-18 module, a hand feature representation of a single RGB image is generated. Then, the EABlock is used to process the hand features of one or both hands to obtain the predicted fingertip joint points. After obtaining the 10 fingertip joint points and 2 palm root joint points output in the first stage, these 12 predicted joint points will be used as one of the inputs in the second stage, and the other input is the features of the remaining 15 joint points or 30 joint points learned by the EABlock module embedded in the second stage. In subsequent processing, the geometric inputs of the 12 joint points output in the first stage will be embedded into a high-dimensional space, added with position encoding, and input into the Transformer encoder layer. Subsequently, the features of these 12 joint points can be feature-fused with the joint point features learned by the EABlock in the second stage. At the same time, the 12 joint points output in the first stage will be input into the subsequent Transformer for decoding to infer the positions of the remaining 30 joint points, while receiving the positive supervision of the 15 or 30 joint point features learned by the EABlock in the second stage. Finally, the 30 predicted joint points output in the second stage and the 12 predicted joint points output in the first stage are weighted and fused to obtain the final output of 42 joint points. The network structure adopted by this method is as Figure 2 shown.
[0040] In a possible implementation, as Figure 1 shown, a hand pose reconstruction method based on multi-stage joint point enhancement of the present invention application is as follows:
[0041] Step 1, Dataset preprocessing.
[0042] Download the dataset Interhand2.6M, which contains 2.6M images. For the original data that needs to extract features, first preprocess it, flip it and randomly crop it to a size of 256×256. The preprocessed images are used as the input images for both the first stage part and the second stage part of the multi-stage joint enhancement network.
[0043] Step 2, Feature learning and prediction of 12 fingertips and palm root joints through the first stage part of the multi-stage joint enhancement network.
[0044] Input the preprocessed images into the 12-joint point prediction training network structure in the first stage of the multi-stage joint enhancement network. The network structure consists of a feature extractor and an EABlock. First, the feature extractor extracts features from the input images to obtain the left hand features and the right hand features ; The feature extractor uses a standard ReaNet-18. This adaptive structure enables the network to obtain higher-resolution feature maps in earlier layers while also reducing the overall receptive field. After feature extraction, the features are fed into the EABlock module to obtain adaptive interaction features. The EABlock module also consists of two stages, namely the extraction stage and the adaptation stage. Through the FuseFormer module in the extraction stage, the left hand features and the right hand features are first used to extract interaction features . Then, the interaction features or are combined with the hand features and respectively fed as two sets of inputs into the adaptation stage. Through the FuseFormer modules of the two branches in the adaptation stage, the two sets of inputs are respectively processed to generate the adaptive interaction features of the left hand and the adaptive interaction features of the right hand . After that, the adaptive interaction features are connected to the corresponding hand features to generate the final adaptive hand features ( or ). The formula for extracting the interaction features is expressed as follows:
[0045] (1)
[0046] where represents the learnable weights of the FuseFormer in the extraction stage.
[0047] In the adaptation stage, the EABlock (as shown in Figure 3 ) uses two additional Fuseformers to adapt the extracted interaction features to each hand. Although the interaction features obtained from the extraction stage help understand how the two hands interact, directly using them for 3D hand mesh recovery of each hand may not be optimal. Therefore, the interaction features are fused with the features of each hand to achieve two goals: 1) retain the interaction information between the two hands; 2) obtain the information unique to the left and right hands. Taking the left hand adaptation as an example, the left hand features and the interaction features are passed to a FuseFormer module, and they are fused through the FuseFormer module to output the interaction features adapted to the left hand . The right hand adaptation is carried out in the same way. The right hand features and the interaction features are passed to another FuseFormer module, and after fusion, the interaction features adapted to the right hand are output .
[0048] (2)
[0049] (3)
[0050] Among them and respectively represent the learnable weights in the Fuseformer adapted to the left and right hands. Finally, for each hand, the adaptive interaction features ( or ) are concatenated with their corresponding hand features ( or ) along the channel dimension as the final adaptive hand features for each hand, denoted by and .
[0051] Subsequently, and are respectively input into the joint feature extractor Joint-feature extractor and the joint enhancer Self-joint Transformer. First, after extracting the rough joint features through the joint feature extractor, they are used as the input of the joint enhancer to obtain the final joint features and predicted joint positions. Finally, the joint features and predicted joint positions of 10 fingertip joints and 2 palm root joints are obtained.
[0052] In the first stage, during training, the preprocessed training data obtained in step 1 is received and trained before the start of the second stage of training. The loss function used during training is MSE:
[0053] (4)
[0054] Among them is the number of joint points in the first stage (including fingertips and wrists), is the joint point positions of the fingertips and wrists predicted in the first stage, is the corresponding ground truth.
[0055] In the first stage, the training data is the preprocessed dataset in step 1, the batch size is set to 40 epochs. After outputting the prediction results through forward propagation, the loss is calculated using the above loss function, and then backpropagation is performed. The adam optimizer is used, and the initial learning rate is set to . When the loss saturates, the learning rate is decreased by a factor of ten.
[0056] Step 3: Feature learning of the remaining 30 hand joint points is performed through the second stage part of the multi-stage joint enhancement network.
[0057] The network structure of the second stage consists of an EABlock module, a MANO module, and a Transformer module. In the second stage, the goal is to utilize the features extracted from the image , concatenate them with the 12 joint points predicted in the first stage, input the concatenated feature vector into the MANO module, and use the 12 joint coordinates output in the first stage as the constraint conditions to infer the positions of the remaining 30 joints through the MANO parametric model.
[0058] Similar to the first stage, the EABlock module is used to extract hand features from the input image, i.e., the preprocessed training data in step 1, and perform joint feature extraction and joint feature enhancement on the 30 extracted hand joint points. Different from the first stage, the 30 joint features predicted in this stage do not need to output the corresponding joint point coordinates.
[0059] Step 4: In the second stage of the multi-stage joint enhancement network, the Transformer module predicts and infers the 30 hand joint points.
[0060] Fuse the 12 joint features output in the first stage with the 30 joint features obtained in step three through the Transformer module. This step realizes the fusion of the geometric constraints obtained in the first stage, i.e., the 12 joint point features, and the image constraints obtained in the second stage, i.e., the 30 joint point features, to achieve the purpose of dual-modal enhancement. In the Transformer module of the second stage, through the multi-head self-attention mechanism, the model can capture the interdependent relationships between different joint points, as well as the associations between joint point geometric constraints and image features, thereby improving the prediction accuracy. For the multi-head self-attention mechanism, first, perform an input linear transformation: Given the input matrix , where n is the length of the sequence, is the dimension of the input features. The input obtained by feature fusion is linearly transformed into the query (Query), key (Key), and value (Value) respectively:
[0061] (5)
[0062] Among them, , , are trainable parameter matrices, and their dimensions are all , where is the feature dimension of each head. Then, calculate the self-attention: For each head, calculate the attention score:
[0063] (6)
[0064] Among them, is the scaling factor, which is used to prevent the dot product value from being too large, so that the gradient of the softmax function becomes too small. Secondly, multi-head calculation is performed: the multi-head self-attention mechanism divides the input into h different subspaces and calculates multiple attention heads in parallel. Each head uses different 、 、 parameter matrices. The outputs of each head obtained will be concatenated together:
[0065] (7)
[0066] (8)
[0067] Among them, is the concatenated linear transformation matrix, which is used to map the output back to the original feature dimension, are the submatrices respectively split from and are used to calculate the output of the i-th attention head.
[0068] Finally, there are residual connection and layer normalization: the output of the multi-head self-attention is residually connected to the input and layer normalization is performed:
[0069] (9)
[0070] The dependencies between the 12 joint points and the 30 joint points to be predicted are captured via the multi-head self-attention mechanism, and the context information extracted from the image features is combined at the same time. Finally, 30 enhanced joint features are output.
[0071] The input data used during the training of this stage is the preprocessed training data in step 1 and the output of the first stage of the multi-stage joint enhancement network. The loss function used for training is as follows:
[0072] (10)
[0073] Among them is the number of joint points in the second stage, is the position of other joint points except the fingertips and wrists predicted in the second stage, is the corresponding true value.
[0074] (11)
[0075] Among them and are hyperparameters, and C is a constant, which is used to make the values of the two cases continuous at |x| = w.
[0076] The comprehensive loss function in the second stage is as follows:
[0077] (12)
[0078] Where and are the weights used to balance the losses in the two stages.
[0079] Step 5: Weightedly fuse and output the 12 fingertip and palm root joint features obtained in the first stage and the 30 joint points inferred by the Transformer in the second stage.
[0080] Through the feature fusion module, the 12 joint features output in the first stage are weightedly concatenated and fused with the enhanced 30 joint features output in the second stage, and the final 42 hand joint point coordinates are obtained through regression by the MLP module. The feature fusion formula for this step is as follows:
[0081] (13)
[0082] Where is the learnable weight matrix, is the bias term, and [⋅, ⋅] represents the concatenation operation.
[0083] Step 6: Reconstruct the hand pose and perform quantitative analysis.
[0084] Through the hand pose regression module Regressor (such as a fully connected layer), the 42 hand joint features obtained in Step 5 are regressed to reconstruct the hand pose, and the experimental results are quantitatively analyzed in combination with the MPJPE evaluation index, and compared with existing excellent methods to measure the accuracy and effectiveness of this method.
[0085] This invention uses MPJPE as the evaluation index, and the formula is as follows:
[0086] (14)
[0087] Where: N is the total number of joint points, is the predicted position of the i-th joint point, is the true position of the i-th joint point, represents the Euclidean distance (usually three-dimensional coordinates) between the predicted position and the true position of the i-th joint point.
[0088] Regarding the details of the neural network: Under the PyTorch framework, the optimizer is selected as Adam, and the initial learning rate is set to , when the loss saturates, the learning rate is decreased by a factor of ten. The input image size is set to 256×256, and the batch size is set to 32. The experimental data is shown as follows:
[0089] Methond Interhand2.6M(MPJPE / mm) IntagHand 8.79 Zhang et al. 13.48 Moon et al. 16.00 The present invention 7.48
[0090] Compared with other existing mainstream methods, trained in the same experimental environment and tested on the public dataset InterHand2.6M dataset, the present invention has achieved better performance, superior to 8.79 of IntagHand.
[0091] The above content is a further detailed description of the present invention in combination with specific / preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several alternatives or modifications can be made to these described embodiments, and these alternative or modified forms should all be regarded as belonging to the protection scope of the present invention.
[0092] The parts not detailed in the present invention belong to the well-known technologies in the art.
Claims
1. A hand pose reconstruction method based on multi-stage joint point enhancement, characterized in that The steps are as follows: Dataset preprocessing; Feature learning and prediction of 12 fingertip and palm root joints through the first-stage part of the multi-stage joint enhancement network; Feature learning of the remaining 30 hand joint points through the second-stage part of the multi-stage joint enhancement network; Predictive inference of 30 hand joint points by the Transformer module in the second stage of the multi-stage joint enhancement network; Weighted fusion and output of the 12 fingertip and palm root joint features obtained in the first stage and the 30 joint points inferred by the second-stage Transformer; Reconstruct the hand pose and conduct quantitative analysis.
2. The hand pose reconstruction method based on multi-stage joint point enhancement according to claim 1, wherein, For the original data that needs to extract features, flip it and randomly crop it to a size of 256×256; Use the preprocessed image as the input image for the first-stage part and the second-stage part of the multi-stage joint enhancement network.
3. A hand pose reconstruction method based on multi-stage joint point enhancement according to claim 1 or 2, characterized in that Feature learning and prediction of 12 fingertip and palm root joints through the first-stage part of the multi-stage joint enhancement network, including: Input the preprocessed image into the 12-joint point prediction training network structure of the first stage of the multi-stage joint enhancement network. The network structure consists of a feature extractor and an EABlock. First, the feature extractor extracts features from the input image. After feature extraction, the features are passed into the EABlock module to obtain adaptive interaction features. Then, the adaptive interaction features are connected with the corresponding hand features to generate the final adaptive hand features; Subsequently, the adaptive hand features are respectively passed into the joint feature extractor and the joint enhancer. First, rough joint features are extracted by the joint feature extractor and then used as the input of the joint enhancer to obtain the final joint features and predicted joint positions. Thus, the joint features and predicted joint positions of 10 fingertip joints and 2 palm root joints are obtained.
4. A hand pose reconstruction method based on multi-stage joint point enhancement according to claim 3, characterized in that The EABlock module includes two stages, namely the extraction stage and the adaptation stage. Through the FuseFormer module in the extraction stage, interaction features are extracted using the left hand features and the right hand features. Then, the interaction features and the hand features are combined and used as two groups of inputs to be passed into the adaptation stage. The FuseFormer modules of the two branches in the adaptation stage respectively process the two groups of inputs to generate the adaptive interaction features of the left hand and the adaptive interaction features of the right hand.
5. A hand pose reconstruction method based on multi-stage joint point enhancement according to claim 3, characterized in that In the first stage, receive the preprocessed training data during training and conduct training before the start of the second-stage training. The loss function used during training is MSE.
6. The hand pose reconstruction method based on multi-stage joint point enhancement according to claim 1, characterized in that, Feature learning of the remaining 30 hand joint points through the second-stage part of the multi-stage joint enhancement network, including: The network structure of the second stage consists of an EABlock module, a MANO module, and a Transformer module; in the second stage, using the hand features extracted from the image, they are concatenated with the joint features of the 12 joint points predicted in the first stage, and the concatenated feature vector is input into the MANO module. Taking the 12 joint coordinates output in the first stage as the constraint conditions, the positions of the remaining 30 joints are inferred through the MANO parametric model; Similar to the first stage, the EABlock module is used to extract hand features from the input image, and joint feature extraction and joint feature enhancement are performed on the 30 hand joint points extracted. Different from the first stage, the 30 joint features predicted in this stage do not need to output the corresponding joint point coordinates.
7. A hand pose reconstruction method based on multi-stage joint point enhancement according to claim 6, characterized in that, In the Transformer module of the second stage, through the multi-head self-attention mechanism, the interdependent relationships between different joint points, as well as the associations between joint point geometric constraints and image features, are captured, thereby improving the prediction accuracy; the interdependent relationships between the 12 joint points and the 30 joints to be predicted are captured through the multi-head self-attention mechanism, and at the same time, the context information extracted from the image features is combined; finally, the enhanced 30 joint features are output.
8. A hand pose reconstruction method based on multi-stage joint point enhancement according to claim 7, characterized in that, The input data used during the training of the second stage is the preprocessed training data in step one and the output of the first stage of the multi-stage joint enhancement network; the loss function used for training is as follows: ; wherein is the number of key points in the second stage, is the positions of other key points except the fingertips and the wrist predicted in the second stage, is the corresponding true value; ; where and are hyperparameters, and C is a constant used to make the values of the two cases continuous at |x| = w; The comprehensive loss function of the second stage is: ; wherein and are weights for balancing the losses in two stages.
9. A hand pose reconstruction method based on multi-stage joint point enhancement according to claim 7, characterized in that The specific method of weighted fusion is as follows: Through the feature fusion module, the 12 joint features output in the first stage and the enhanced 30 joint features output in the second stage are weighted and concatenated for fusion, and the final 42 hand joint point coordinates are obtained through regression by the MLP module.
10. A hand pose reconstruction method based on multi-stage joint point enhancement according to claim 9, characterized in that, Reconstruct the hand pose and conduct quantitative analysis, including: The hand pose is reconstructed by regressing the final 42 hand joint features through the hand pose regression module, and the experimental results are quantitatively analyzed in combination with the MPJPE evaluation index.
Citation Information
Patent Citations
Accurate three-dimensional hand posture estimation method
CN111401151A
Double-flow multi-scale hand posture estimation method based on single RGB image
CN113052030A
Three-dimensional hand posture estimation method based on monocular RGB image
CN115588237A
End-to-end hand object interaction attitude estimation method and system
CN118247851A
Two-stage hand key point identification method based on deep learning
CN119904908A