A hand pose estimation method based on dual aggregation network enhancement
By using a dual-aggregation network-enhanced hand posture estimation method in 3D gesture posture estimation, the problem of insufficient accuracy and robustness of hand posture estimation in the prior art is solved, and higher robustness and accuracy are achieved, especially in complex environments.
Patent Information
- Application Number
- CN202411607112.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-11-12
AI Technical Summary
The existing 3D gesture posture estimation method has problems of insufficient accuracy and robustness when dealing with high-dimensional data, significant changes in hand posture, finger appearance differences and self-occlusion.
Using the hand position estimation method based on dual aggregation network enhancement, the hand depth image and point cloud features are extracted and fused, and the point cloud image consistency aggregation module and dynamic map enhancement aggregation module are used to cycle three times to obtain accurate hand position pose.
Improve gesture estimation in occlusion situations, enhance hand joint output effect, and improve system robustness and accuracy, especially in complex environments.
Smart Images

Figure CN119418409B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a hand pose estimation method, in particular to a hand pose estimation method based on dual aggregation network enhancement, belonging to the technical field of visual processing and pose estimation. Background Art
[0002] 3D hand gesture estimation is a key technology in human-computer interaction applications, and has been widely used in virtual reality (VR), augmented reality (AR), and robotics. This research area has become a focus in computer vision, and the advancement of affordable depth sensors has rekindled interest in this field in recent years. However, 3D hand gesture estimation faces several challenges, including the high dimensionality of the data, significant variations in hand poses, subtle appearance differences between fingers, and severe self-occlusion, which hinder accurate and robust estimation of 3D hand gesture poses.
[0003] With the emergence of deep learning, deep learning-based methods have dominated the task of 3D hand gesture estimation. These deep learning-based methods can be divided into two categories according to the format of the input data: methods that utilize 2D images and methods that utilize 3D data. Image-based methods usually use 2D depth images as input and use 2D convolutional neural networks (CNNs) to extract local features from these images. The high parallelism of 2D convolution operations enables them to be efficiently computed on modern hardware, thereby improving the performance and speed of these methods in practical applications. In contrast, 3D data-based methods convert 2D depth images into 3D voxel representations or 3D point cloud representations, and then use 3D convolutional networks or point cloud networks for hand gesture estimation. These methods directly process 3D coordinate information instead of projecting it onto a 2D plane. Therefore, it avoids the information loss caused by projection and better preserves the spatial geometric structure of the original data.
[0004] Despite significant progress in image- and point-cloud-based 3D hand gesture estimation methods, each method still suffers from specific limitations. Image-based methods often fail to fully exploit the 3D features of depth data and have difficulty handling the complex nonlinear mapping required to convert 2D depth images into accurate 3D hand gesture poses. In addition, these methods are limited by the local receptive field of 2D convolutional neural networks (CNNs), limiting their ability to capture long-range dependencies and interactions. In contrast, point-cloud-based methods require the construction of dense and dynamic local neighborhoods and involve complex feature extraction processes, which increase computational costs and pose challenges for real-time performance and scalability. The unstructured nature of point clouds further complicates the extraction and registration of local features, which may affect the model's ability to generalize across a wide range of hand gesture poses. In addition, many existing methods rely on non-recursive architectures and lack the flexibility required to adapt to different resource constraints and accuracy requirements, a limitation that may hinder the effectiveness of these methods in a variety of applications where adaptability and accuracy are critical. Summary of the invention
[0005] The purpose of the present invention is to provide a hand pose estimation method based on dual aggregation network enhancement, which can improve hand gesture estimation under occlusion and enhance the hand joint output effect.
[0006] In order to achieve the above object, the present invention provides a hand pose estimation method based on dual aggregation network enhancement, comprising the following steps:
[0007] S1: First, from the hand depth image I in Generate hand point cloud position P in , and then the hand depth image I in and the hand point cloud position P in Input into the local coding fusion module to generate the fused image feature F f2D and point cloud features F f3D ;
[0008] S2: The fused 3D point cloud feature F f3D Input into the initial state generator to initialize the hidden state S 0 ;
[0009] S3: Initialize the hidden state S 0 Input to the regression module to obtain the initial estimate J of the joint point 0 ;
[0010] S4: Initial estimate J 0 And the fused image feature F f2D and point cloud features F f3D , input to the point cloud image consistency aggregation module (PICA) to generate enhanced point cloud features P JI ;
[0011] S5: Enhanced point cloud features P JI With the initial estimate J 0 The input resampling module outputs high-dimensional hand joint feature J P1 ;
[0012] S6: Enhanced point cloud features P JI and high-dimensional hand joint feature J P1 and the hidden state S of the previous stage 0 The joint features are input into the dynamic graph enhancement aggregation module (DGIA) to obtain the enhanced high-dimensional joint feature J P4 ;
[0013] S7: The enhanced high-dimensional joint feature J P4 Input to the regression module or the first iteration to estimate J 1 ;
[0014] S8: Repeat steps S4 to S7 twice to obtain the final hand joint point coordinate position J 3 .
[0015] The specific steps of step S2 of the present invention are as follows:
[0016] S21: First, 3D point cloud feature F f3D Generate the global vector J through MLP a1 ;
[0017] S22: Generate the initial hidden state S through three layers of bias induction layer (BIL) in sequence 0 , the global vector J a1 is copied J times and input into three bias induction layers (BIL) to generate the hidden state S of each joint in [J,512] dimensions. 0 ,The three bias inducing layers (BILs) provide joint-independent biases which can be viewed as a learnable embedding of position information, enabling different joints to be uniquely mapped from the same global features.
[0018] The specific steps of step S4 of the present invention are as follows:
[0019] S41: Initial estimate j 0 First project it into 2D plane space, then pass through similarity mapping fusion module and fused image feature F f2D Perform splicing and fusion to obtain the enhanced image feature F that integrates the prior information fs2D ;
[0020] S42: Enhanced image features F fs2D , hand point cloud position P in , 3D point cloud features F f3DInput the bilinear grid sampling module together to obtain the image feature F mapped to the 3D space fsb2D ;
[0021] S43: Enhanced image features F fs2D With 3D point cloud feature F f3D After splicing, the enhanced point cloud feature P that combines prior information and image features is obtained through 1*1 convolution fusion. JI .
[0022] The specific steps of step S6 of the present invention are as follows:
[0023] S61: High-dimensional hand joint features J P1 First, input the GCN module to model the graph relationship and obtain the enhanced joint feature J P2 , the joint graph G = (V, E) is composed of the joint set V = {v i |i=1,…,J} and the edge set E={e i |i=1,…,M}, where M represents the total number of limbs defined artificially;
[0024] set up is the joint v of the lth layer i The representation of the feature, A∈[0,1] J×J Defined as the adjacency matrix of graph G, if the i-th joint is connected to the j-th joint, then a ij Set to 1, so the graph convolution can be formulated as: S 1 =σ(WS 1-1 A), where σ represents a nonlinear function, such as ReLU, W∈R dhid×dhid is the trainable weight;
[0025] S62: Enhanced joint feature J P2 , Enhanced point cloud features P JI , hidden state S 0 The common input GRU module generates enhanced joint point features J P3 ;
[0026] S63: Enhanced joint feature J P3 Input the dynamic graph convolution module to obtain the enhanced high-dimensional joint feature J P4 , the attention matrix A of dynamic graph convolution d is obtained by calculation, for each node v in the hand skeleton i ∈R dq , firstly, by performing node feature S i ∈R Cin Apply a trainable linear transformation to compute the query vector q i ∈R dq , key vector ki ∈R dk and a numeric vector v i ∈R dv , using the shared parameter W q ∈R Cin×dq , W k ∈R Cin×dk and W v ∈R Cin×dv , then, for each hand joint, a query-key addition operation is performed to obtain a weighted score a ij , which quantifies the strength of the association between two joints, the attention score between the two joints It is given by the following formula: When obtaining the dynamic adjacency matrix A d Finally, the DGCN module:
[0027] Compared with the prior art, the present invention first extracts and fuses the hand depth image and point cloud features to obtain enhanced 2D depth map features and 3D point cloud features, then estimates the initial pose estimate from the point cloud features, and then inputs the initial pose estimate, enhanced image features and point cloud features into the point cloud image consistency aggregation module (PICA) and the dynamic graph aggregation module (DGIA) for feature aggregation and enhancement, and repeats three times to obtain accurate hand pose. The dual aggregation network proposed in the present invention combines two different aggregation forms to utilize their unique characteristics and improve hand gesture estimation under occlusion. The point cloud image consistency aggregation module (PICA) module uses learnable interpolation and sampling techniques to seamlessly integrate point cloud and image features from depth data. The dynamic graph enhancement aggregation module (DGIA) module uses a learnable graph convolutional network to enhance the hand joint output by combining the hidden state of GRU with local point features. The present invention can improve hand gesture estimation under occlusion and enhance the hand joint output effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 A flowchart of the invention;
[0029] Figure 2 Flowchart of the point cloud image consistency aggregation module;
[0030] Figure 3 is a schematic diagram of the similarity mapping fusion module;
[0031] Figure 4 Flowchart of the dynamic graph convolution gated recurrent module. DETAILED DESCRIPTION
[0032] The present invention will be further described below in conjunction with the accompanying drawings.
[0033] like Figure 1-Figure 4 As shown, a hand pose estimation method based on dual aggregation network enhancement includes the following steps:
[0034] S1: First, from the hand depth image I in Generate hand point cloud position P in , and then the hand depth image I in and the hand point cloud position P in Input into the local coding fusion module to generate the fused image feature F f2D and point cloud features F f3D ;
[0035] S2: The fused 3D point cloud feature F f3D Input into the initial state generator to initialize the hidden state S 0 ;
[0036] The specific steps of step S2 are as follows:
[0037] S21: First, 3D point cloud feature F f3D Generate the global vector J through MLP a1 ;
[0038] S22: Generate the initial hidden state S through three layers of bias induction layer (BIL) in sequence 0 , the global vector J a1 is copied J times and input into three bias induction layers (BIL) to generate the hidden state S of each joint in [J,512] dimensions. 0 ,The three bias inducing layers (BILs) provide joint-independent biases which can be viewed as a learnable embedding of position information, enabling different joints to be uniquely mapped from the same global features.
[0039] S3: Initialize the hidden state S 0 Input to the regression module to obtain the initial estimate J of the joint point 0 ;
[0040] S4: Initial estimate J 0 And the fused image feature F f2D and point cloud features F f3D , input to the point cloud image consistency aggregation module (PICA) to generate enhanced point cloud features P JI ;
[0041] The specific steps of step S4 are as follows:
[0042] S41: Initial estimate j 0 First project it into 2D plane space, then pass through similarity mapping fusion module and fused image feature F f2DPerform splicing and fusion to obtain the enhanced image feature F that integrates the prior information fs2D ;
[0043] Project the hand joints onto the image plane, represented by J = {j i ∈R I×2 Considering the sparsity of joints, this paper proposes an adaptive similarity projection fusion method to generate a dense feature map D∈R from sparse three-dimensional joint features. H×W×C3D , which involves identifying, for each specified pixel q in the dense map, the k nearest neighbors in the image plane projection, as Figure 3 As shown;
[0044] Next, a multi-layer perceptron (MLP) is used, and then the features are combined through the mean aggregation process;
[0045] Incorporating 2D similarity measures into the fusion module enhances its robustness, especially in complex scenes involving overlapping objects. This integration enables the module to utilize dense 2D features to guide the densification of sparse 3D features. After this enhancement, the joint features are concatenated with the input image features. Then, 1×1 convolutions are applied to these concatenated features to reduce the dimensionality. By reintegrating the initially estimated coarse 3D hand joint coordinates with the 2D image features, the system provides prior information that is critical to the subsequent learning stage. This approach helps to refine and evaluate the initial estimates, thereby improving the overall accuracy of gesture estimation.
[0046] S42: Enhanced image features F fs2D , hand point cloud position P in , 3D point cloud features F f3D Input the bilinear grid sampling module together to obtain the image feature F mapped to the 3D space fsb2D ;
[0047] Initially, the points are projected to the image plane to obtain the corresponding 2D image features. If the coordinates are non-integer, bilinear interpolation is used to retrieve the 2D image features of the points;
[0048] S43: Image features F mapped to 3D space fsb2D With 3D point cloud feature F f3D After splicing, the enhanced point cloud feature P that combines prior information and image features is obtained through 1*1 convolution fusion. JI .
[0049] S5: Enhanced point cloud features P JI With the initial estimate J 0 The input resampling module outputs high-dimensional hand joint feature J P1 ;
[0050] S6: Enhanced point cloud features P JI and high-dimensional hand joint feature J P1 and the hidden state S of the previous stage 0 The joint features are input into the dynamic graph enhancement aggregation module (DGIA) to obtain the enhanced high-dimensional joint feature J P4 ;
[0051] The specific steps of step S6 are as follows:
[0052] S61: High-dimensional hand joint features J P1 First, the graph convolutional neural network (GCN) module is input to model the graph relationship and obtain the enhanced joint point feature J P2 , the joint graph G = (V, E) is composed of the joint set V = {v i |i=1,…,J} and the edge set E={e i |i=1,…,M}, where M represents the total number of limbs defined artificially;
[0053] set up is the joint v of the lth layer i The representation of the feature, A∈[0,1] J×J Defined as the adjacency matrix of graph G, if the i-th joint is connected to the j-th joint, then a ij Set to 1, so the graph convolution can be formulated as: S 1 =σ(WS 1-1 A), where σ represents a nonlinear function, such as ReLU, W∈R dhid×dhid is the trainable weight;
[0054] S62: Enhanced joint feature J P2 , Enhanced point cloud features P JI , hidden state S 0 The common input gated recurrent unit GRU module generates enhanced joint point features J P3 ;
[0055] S63: Enhanced joint feature J P3 Input the dynamic graph convolution module to obtain the enhanced high-dimensional joint feature J P4 , the attention matrix A of dynamic graph convolution d is obtained by calculation, for each node v in the hand skeleton i ∈R dq First, we use the node feature S i ∈R Cin Apply a trainable linear transformation to compute the query vector q i ∈R dq , key vector k i ∈R dk and a numeric vector vi ∈R dv , using the shared parameter W q ∈R Cin×dq , W k ∈R Cin×dk and W v ∈R Cin×dv , then, for each hand joint, a query-key addition operation is performed to obtain a weighted score a ij , which quantifies the strength of the association between two joints, the attention score between the two joints It is given by the following formula: When obtaining the dynamic adjacency matrix A d Finally, the DGCN module:
[0056] S7: The enhanced high-dimensional joint feature J P4 Input to the regression module or the first iteration to estimate J 1 .
[0057] S8: Repeat steps S4 to S7 twice to obtain the final hand joint point coordinate position J 3 .
[0058] The present invention is applied to virtual office, and an embodiment of the present invention is given.
[0059] Assuming that the user is wearing an AR headset and is in a virtual office environment, the device needs to accurately identify the user's hand posture to support the user's precise virtual desktop operations. Users can use gesture control to adjust the layout of the virtual desktop, select applications, browse documents, etc. All operations rely on hand posture estimation technology. In order to provide a smoother interactive experience, it is necessary to ensure that the device can stably and efficiently capture hand movements in complex environments, while eliminating the sensitivity of traditional methods to lighting changes, occlusions, and depth perception errors.
[0060] Implementation steps
[0061] Environment setup and data collection:
[0062] The AR headset is equipped with multiple high-resolution depth cameras located on the front and sides of the device to fully capture the user's hand movements. The camera scans the user's hand to obtain high-precision depth information from the finger joints to the palm.
[0063] When the device is initialized, the system automatically scans and builds a 3D model of the user's hand. The camera records every subtle movement of the fingers and palm in space in real time, generating a dynamic 3D skeleton based on the depth information of the hand.
[0064] Dual Aggregation Network Enhanced Hand Pose Estimation:
[0065] Network architecture: The system uses a dual aggregation network (DANet) to enhance hand pose estimation. The system uses a dual aggregation network consisting of two main components: one is a point cloud image consistency aggregation module, which is responsible for capturing the overall position and movement trend of the hand; the other is a dynamic image enhancement aggregation module, which focuses on accurately modeling the detailed movement of each finger.
[0066] Point cloud image consistency aggregation module: This module identifies the relative position of the hand by analyzing the overall movement of the user's hand, such as the direction of hand movement in space, wrist rotation, etc. This module can process hand data collected by multiple depth sensors and generate a unified hand dynamic representation by fusing multimodal data.
[0067] Dynamic graph enhancement aggregation module: The dynamic graph enhancement aggregation module enhances the accuracy of hand movements by carefully modeling each finger joint. Through dynamic graph enhancement aggregation, the system can track the movement trajectory and bending angle of each finger, especially in complex environments, and can reduce recognition errors caused by different camera angles.
[0068] Hand motion recognition and virtual desktop interaction:
[0069] When the user performs a gesture, such as clicking a virtual button with a finger or scrolling a virtual document by swiping, the depth camera captures the user's hand movements in real time, and the system estimates the hand pose through a dual aggregation network.
[0070] When the user makes a "pointing" action, the system will determine the exact position of the finger and accurately map a cursor on the virtual desktop. The user only needs to point the finger at the application or file to be selected, and the system will sense and perform the corresponding operation. For example, the user can select different application windows by pointing, or use the "drag" gesture to move the window to a different position on the screen.
[0071] When the user performs a pinch gesture, the system will recognize the pinching degree of the fingers and respond to the zoom operation in the virtual interface. When the user pinches his fingers, the virtual document or image will be enlarged or reduced, thus achieving a more intuitive operation experience.
[0072] Hand pose estimation in complex environments:
[0073] In actual applications, users may be in complex backgrounds or environments, such as uneven lighting or obstructions. In this case, traditional hand pose estimation methods may be greatly disturbed, resulting in reduced recognition accuracy.
[0074] Through the dual aggregation network enhancement method of the present invention, the device can effectively suppress occlusion and environmental noise with the support of depth information. The point cloud image consistency aggregation module can correct local errors based on a large range of spatial information, while the dynamic image enhancement aggregation module can focus on repairing the estimation errors of the finger details, thereby improving the robustness of the system.
[0075] System feedback and user interaction:
[0076] The system provides real-time feedback based on the estimated hand posture and the user's gestures. When the user performs each gesture, the device will display the corresponding virtual interface response, such as highlighting the virtual button, moving or scaling the virtual window, etc.
[0077] Implementation results and effects
[0078] Through the dual-aggregation network enhancement method in this implementation case, the AR headset device can stably estimate hand poses in complex environments, greatly improving the fluency and accuracy of virtual desktop operations.
[0079] The system can maintain high-precision response when the user performs various gestures, especially in environments with changing light, occlusion or complex hand movements, the device can still accurately track the user's hand movements.
[0080] Users can easily interact with the virtual interface through natural gestures without relying on a traditional mouse, keyboard or external controller, which improves the immersion and convenience of the interaction.
[0081] Compared with traditional gesture recognition methods, the hand pose estimation method based on dual aggregation network enhancement can significantly reduce misrecognition and operation delays, ensuring that users can interact with the virtual environment in real time.
Claims
1. A hand pose estimation method based on dual aggregation network enhancement, characterized in that: The following steps are involved: S1: First, from the hand depth image Generate hand point cloud position , and then the hand depth image And the hand point cloud position Input into the local coding fusion module to generate fused image features and point cloud features ; S2: The fused 3D point cloud features Input into the initial state generator to initialize the hidden state ; S3: Initialize the hidden state Input to the regression module to obtain the initial estimate of the joint points ; S4: Initial Estimate And the fused image features and point cloud features , input to the point cloud image consistency aggregation module to generate enhanced point cloud features ; S5: Enhanced point cloud features With initial estimate The input resampling module outputs high-dimensional hand joint features ; S6: Enhanced point cloud features and high-dimensional hand joint features and the hidden state of the previous stage The joint features are input into the dynamic graph enhancement aggregation module to obtain enhanced high-dimensional joint features. ; S7: Enhanced high-dimensional joint features Input into the regression module to obtain the first iteration estimate ; S8: Repeat steps S4 to S7 twice to obtain the final hand joint coordinate position ; The specific steps of step S4 are as follows: S41: Initial Estimates First project it into 2D plane space, then pass through similarity mapping fusion module and fused image features Perform splicing and fusion to obtain enhanced image features that incorporate prior information ; S42: Enhanced image features , hand point cloud position , 3D point cloud features Input the bilinear grid sampling module together to get the image features mapped to 3D space ; S43: Obtain image features mapped to 3D space With 3D point cloud features After splicing, the enhanced point cloud features that combine prior information and image features are obtained through 1*1 convolution fusion. .
2. The hand pose estimation method based on dual aggregation network enhancement according to claim 1 is characterized in that: The specific steps of step S2 are as follows: S21: First 3D point cloud features Generate global vector through MLP ; S22: Generate the initial hidden state by executing three bias-inducing layers in sequence , the global vector Copied times and input into three bias induction layers to generate Hidden states of individual joints ,The three bias inducing layers provide joint-independent biases which are regarded as ,learnable position information embedding, enabling different joints to be uniquely mapped from the ,same global features.
3. The hand pose estimation method based on dual aggregation network enhancement according to claim 1 is characterized in that: The specific steps of step S6 are as follows: S61: High-dimensional hand joint features First, input the GCN module to model the graph relationship and obtain enhanced joint point features , joint diagram By joint collection and edge set Composition, among which, represents the total number of limbs defined artificially; S62: Enhanced joint features , Enhance point cloud features , hidden state Common input GRU module generates enhanced joint point features ; S63: Enhanced joint features Input dynamic graph convolution module to obtain enhanced high-dimensional joint features .
Citation Information
Patent Citations
Human hand three-dimensional posture estimation method and device based on three-dimensional point cloud
CN110222580A
Three-dimensional hand posture estimation and recognition method based on image sequence
CN114882493A