Gesture reconstruction method based on image point cloud multi-modal fusion and joint guidance

By combining multimodal fusion technology with depth images and point cloud data in three-dimensional gesture reconstruction, using joint guidance and attention mechanisms, the problems of low accuracy of three-dimensional gesture reconstruction and complex feature extraction in the existing technology are solved, achieving higher reconstruction accuracy and robustness.

CN119963740AActive Publication Date: 2025-05-09CHONGQING UNIV OF POSTS & TELECOMM
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510125414.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-09
Estimated Expiration
2045-01-27

Smart Images

  • Figure CN119963740A_ABST
    Figure CN119963740A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of human-computer interaction and computer vision, and particularly relates to a gesture reconstruction method based on image point cloud multi-modal fusion and joint guidance. Respectively extracting key point features of the depth image and the three-dimensional point cloud by using a feature extraction network for the two-dimensional depth image and the three-dimensional point cloud; a multi-modal feature and global feature fusion module is used to provide additional global information for the multi-modal features; thirdly, further optimizing and updating fusion features by using a feature fusion iteration module guided by joint coordinates so as to improve the reconstruction precision of the three-dimensional gesture, and meanwhile, realizing accurate estimation of each key point in the hand posture through multiple times of iteration updating; according to the method, depth image information and point cloud space geometric structure features are effectively combined, and feature information of respective modals is aggregated by adopting key point features, so that redundant interaction of invalid features is reduced, and the efficiency of multi-modal fusion is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of human-computer interaction and computer vision, and in particular relates to a gesture reconstruction method based on image point cloud multimodal fusion and joint guidance. Background Art

[0002] With the rapid development of computer vision and artificial intelligence technologies, gesture recognition and 3D gesture reconstruction have become important research directions in the fields of human-computer interaction, virtual reality (VR), augmented reality (AR), etc. Many of the more challenging application scenarios require higher accuracy in 3D hand gesture estimation. For example, augmented reality application scenarios involve frequent hand movements and occluded hand-object interactions. In the task of 3D gesture reconstruction, accurate prediction of the coordinates of the key points of the hand can effectively analyze the user's hand movements and behaviors, which is the basis for achieving efficient human-computer interaction.

[0003] With the advancement of deep learning technology and the popularity of low-cost depth cameras, 3D hand pose estimation based on depth images has made significant progress. However, there are still some scenarios where difficult and challenging problems exist, mainly including the following aspects: (1) Most of the early gesture reconstruction methods are based on depth images. For example, the method proposed in the patent application with publication number CN102262783A processes the depth data as a single-channel two-dimensional image, which inevitably ignores the three-dimensional nature of the depth data. (2) In the dynamic gesture recognition task, some methods directly use the hand point cloud data to extract local features of the hand, such as the method proposed in the patent application with publication number CN11 Patent applications No. 8887736A and CN118202987A require very complex feature extractors due to their disordered and unstructured characteristics. In addition, since their task goal is recognition and the focus is not on 3D gesture reconstruction, the accuracy of the obtained hand joint coordinates is low. (3) Currently, there are very few methods that use depth images and point cloud data at the same time. For example, the patent application with publication number CN117351573A only uses depth images to obtain 2D hand joint coordinates, ignoring the ability of depth images to describe local contours and reflect local depth changes, and does not fully explore and utilize the features of depth images and hand point clouds. Summary of the invention

[0004] In view of the problems existing in the prior art, the present invention proposes a gesture reconstruction method based on image point cloud multimodal fusion and joint guidance, constructs a gesture reconstruction model, and uses the model to reconstruct gestures from depth images. The reconstruction process includes the following steps:

[0005] 101. Obtain a hand depth image and its point cloud data, an image feature extraction network extracts local image features and global image features from the depth image, and a point cloud feature extraction network extracts global point cloud features from the point cloud data;

[0006] 102. The global image features and the global point cloud features are concatenated together, and a bias embedding vector is generated for each joint point through an induced bias layer. The vectors of all joint points constitute the global key point embedding.

[0007] 103. Global key point embedding projects the high-dimensional embedding into the three-dimensional space through linear transformation to obtain the initial three-dimensional joint coordinates, and obtains the local point cloud features according to the three-dimensional joint points;

[0008] 104. The three-dimensional joint point coordinates are updated through the graph convolution network to obtain updated global features, and the global features are respectively spliced ​​with the local image features and the local point cloud features to obtain image joint features and point cloud joint features;

[0009] 105. Based on the attention mechanism, the image joint features and point cloud joint features are weighted respectively, and then the weighted features are updated using the graph convolutional network;

[0010] 106. Use the updated image joint features and point cloud joint features of the graph convolutional network to predict the three-dimensional joint coordinates, and determine whether the iteration limit is reached. If not, update the point cloud local features according to the predicted three-dimensional joint points and return to step 104. Otherwise, the currently predicted three-dimensional joint coordinates are used as the reconstructed structure.

[0011] Furthermore, the process of extracting local image features and global image features from the depth image by the image feature extraction network includes:

[0012] The depth image is processed by the autoencoder network to obtain a two-dimensional local feature map F 2d ∈R C×H×W , where C, H, and W are the two-dimensional local feature maps F 2d Number of channels, height, and width;

[0013] Two-dimensional local feature map F 2d Input to the convolutional residual network to generate the heat map corresponding to each joint point The 2D UV coordinates of the joints are estimated based on the normalized heat map, and the image features of the joints are aggregated from the local feature map according to the heat map to obtain the local image features. Where J is the number of hand joints in the depth image, d 2d is the image feature dimension corresponding to each joint;

[0014] Two-dimensional local feature map F 2d A global image feature is output through global pooling.

[0015] Furthermore, the process of extracting local point cloud features and global point cloud features from point cloud data by the point cloud feature extraction network includes:

[0016] Extract 3D local geometric features F from point cloud data through point set convolutional layer 3d ∈R N×d , where N represents the number of points in the point cloud, and d represents the point cloud feature dimension corresponding to each point;

[0017] 3D local geometric features F 3d ∈R N×d After passing through the one-dimensional convolution layer and the maximum pooling layer, the global point cloud features will be obtained;

[0018] The two-dimensional local feature map F 2d After being projected to the feature space of the 3D local geometric features, it is spliced ​​with the 3D local geometric features to generate the fused point cloud feature F fuse ;

[0019] Using the initialized 3D joint coordinates as the starting position, further guide the point cloud feature F fuse Aggregation to generate the final local point cloud features where d 3d Indicates the point cloud feature dimension corresponding to each joint.

[0020] Furthermore, the process of weighting the image joint features and the point cloud joint features based on the attention mechanism includes:

[0021] The image joint features and the point cloud joint features are mapped respectively through two linear modules to obtain the query vector, key vector and value vector of the image joint features and the point cloud joint features;

[0022] Concatenate the query vectors, key vectors, and value vectors of the image joint features and the point cloud joint features;

[0023] The concatenated query vector, key vector and value vector are input into the self-attention module to obtain the attention weights, and the image joint features and point cloud joint features are weighted using the attention weights.

[0024] Compared with the existing technology, the present invention effectively combines the depth image information with the spatial geometric structure features of the point cloud, and uses key point features to aggregate the feature information of each modality, reducing the redundant interaction of invalid features and improving the efficiency of multimodal fusion; in addition, the present invention deeply mines the complementary information between multiple modalities through the joint coordinate guidance strategy, which helps to reduce the interference of ambiguous information in cross-modal feature fusion and can significantly improve the accuracy and robustness of gesture 3D reconstruction. In addition, the method proposed in the present invention has been experimentally verified on challenging public gesture datasets such as NYU and DexYCB, and the results show that the method is much better than previous methods in 3D gesture reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a flow chart of the three-dimensional hand gesture reconstruction method based on image point cloud multimodal fusion and joint guidance of the present invention;

[0026] Figure 2 Graphs of input data of two different modalities in the present invention, where (a) is a deep depth image and (b) is hand point cloud data;

[0027] Figure 3 Schematic diagram of the effect of 3D gesture reconstruction in the present invention, wherein Figure (a) is the reconstruction result on a single-hand dataset, and Figure (b) is the reconstruction result on a hand-object interaction dataset;

[0028] Figure 4 It is the network framework of the three-dimensional gesture reconstruction method based on image point cloud multimodal fusion and joint guidance of the present invention. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0030] The present invention proposes a gesture reconstruction method based on image point cloud multimodal fusion and joint guidance, constructs a gesture reconstruction model, and uses the model to reconstruct gestures from depth images. Figure 1 , the reconstruction process includes the following steps:

[0031] 101. Obtain a hand depth image and its point cloud data, an image feature extraction network extracts local image features and global image features from the depth image, and a point cloud feature extraction network extracts global point cloud features from the point cloud data;

[0032] 102. The global image features and the global point cloud features are concatenated together, and a bias embedding vector is generated for each joint point through an induced bias layer. The vectors of all joint points constitute the global key point embedding.

[0033] 103. Global key point embedding projects the high-dimensional embedding into the three-dimensional space through linear transformation to obtain the initial three-dimensional joint coordinates, and obtains the local point cloud features according to the three-dimensional joint points;

[0034] 104. The three-dimensional joint point coordinates are updated through the graph convolution network to obtain updated global features, and the global features are respectively spliced ​​with the local image features and the local point cloud features to obtain image joint features and point cloud joint features;

[0035] 105. Based on the attention mechanism, the image joint features and point cloud joint features are weighted respectively, and then the weighted features are updated using the graph convolutional network;

[0036] 106. Use the updated image joint features and point cloud joint features of the graph convolutional network to predict the three-dimensional joint coordinates, and determine whether the iteration limit is reached. If not, update the point cloud local features according to the predicted three-dimensional joint points and return to step 104. Otherwise, the currently predicted three-dimensional joint coordinates are used as the reconstructed structure.

[0037] In this embodiment, a specific implementation process of a gesture reconstruction method based on image point cloud multimodal fusion and joint guidance is proposed, which specifically includes the following steps:

[0038] Step 1: Use different feature extraction networks to extract corresponding multimodal features, global features, and initialized 3D key point coordinates from the input hand depth image and 3D point cloud. The depth image feature extraction network mainly captures local details and depth change information of the hand, and the point cloud feature extraction network mainly obtains the 3D structure information and spatial features of the hand. In this embodiment, the multimodal features include depth image features and 3D point cloud features.

[0039] Step 2: Use the feature fusion module to fuse the multimodal features with the global features to obtain the multimodal key point features that integrate the global information. By fusing the global information with the key point features of different modalities, the loss of global features is compensated and the error caused by occlusion on the joint point prediction is reduced.

[0040] Step 3: Use the initialized 3D key point coordinates to guide the depth map key point features and point cloud key point features to further interact and merge, and improve the semantic consistency and spatial consistency of the features. This step uses the initial key point coordinates as a guide to further couple and optimize the depth map and point cloud data through feature interaction.

[0041] Step 4: Update the multimodal features based on the idea of ​​iterative correction, use the two-dimensional visual features and three-dimensional geometric structure information to gradually correct the gesture, and finally obtain an accurate and robust three-dimensional gesture model. Through the iterative optimization mechanism, the gradual approximation of the joint position and the accurate reconstruction of the model are achieved.

[0042] like Figure 4The present invention extracts two-dimensional local feature maps and three-dimensional local features from depth maps and point cloud data respectively through image encoders and point cloud encoders, and obtains global features based on the two-dimensional local feature maps and three-dimensional local features; the two-dimensional local feature maps and three-dimensional local features are further processed to obtain point cloud joint features and depth joint features, and the point cloud joint features and the depth joint features are respectively spliced ​​with the global features and then feature fusion is performed, and the joint regressor predicts the three-dimensional joint point coordinates based on the input fused features.

[0043] In this embodiment, the hand depth image is a cropped image focusing on the hand area, which may be a single hand image or a hand-object interaction image. The hand depth image is an image captured by a depth camera. Figure 2 (a) An example of a hand depth image is given; the 3D point cloud is converted from the depth image using the camera parameters, that is, the point cloud data is generated by back-projecting the depth map. Figure 2 (a) Depth map projection generation Figure 2 (b) Corresponding point cloud data, the depth image and point cloud data are processed by the image processing branch and the point cloud processing branch respectively to extract the corresponding local features and global information, and initialize the three-dimensional key point coordinates.

[0044] In the image processing branch, the depth image is first processed by the autoencoder network to obtain a two-dimensional local feature map F 2d ∈R C×H×W ; Then, the local feature map F 2d Input to the convolutional residual network to generate the heat map corresponding to each joint point And further aggregate the image features of the joints from the local feature map according to the heat map In particular, the 2D UV coordinates of the joints can be estimated through the heat map, which is supervised by the real UV coordinates, ensuring the accuracy and robustness of image feature extraction.

[0045] In order to optimize the predicted joint point UV coordinates (U' j ,V j ') and the real joint point UV coordinates (U j ,V j ) (i.e., Ground Truth, GT), this embodiment introduces the Smooth L1 loss function to supervise it, and its loss formula is as follows:

[0046]

[0047] Where J is the number of hand joints in the depth image; L(x) is the standard Smooth L1 loss function, which is defined as:

[0048]

[0049] This loss function is robust. When the prediction error is small, the loss is in the form of a quadratic function, which can provide a strong gradient signal. When the error is large, the loss is in a linear form, which avoids the excessive influence of abnormal points on the training process, thereby improving the stability of network training. Through the optimization of this loss function, the deviation between the predicted UV coordinates and the real coordinates can be effectively reduced, providing higher accuracy guarantee for subsequent 3D gesture reconstruction.

[0050] In the point cloud processing branch, the 3D local geometric features F are extracted from the point cloud data through a point set convolutional layer (such as PointNet or PointNet++, etc.) 3d ∈R N×d In order to enhance the multimodal interaction between the depth map and the point cloud, the two-dimensional local feature map F 2d The projected feature F proj And the local feature F of the point cloud 3d Splice to generate the fused point cloud feature F fuse In addition, the initialized 3D key point coordinates are used as the starting position to further guide the aggregation of local features and generate the final point cloud joint features.

[0051] Specifically, the process of guiding local feature aggregation includes, for each joint point, taking the three-dimensional coordinates of the joint point as the center, using the ball query algorithm to find K neighboring points within a radius r around the position in the point cloud data, and using the geometric features of these points as the initial point cloud features of the joint point. This step ensures that the point cloud features can effectively capture the local geometric information near each joint point, laying the foundation for subsequent fusion processing.

[0052] At the same time, after the depth image and point cloud data pass through their respective feature extraction branches, not only the corresponding local features are obtained, but also the image global vector G containing global information is generated. img and the point cloud global vector G pc ; Next, the two global vectors are concatenated and replicated J times to provide them to the induced bias layer to generate an independent bias embedding vector for each joint point. This operation is equivalent to generating a learnable position embedding for each joint point, which helps to improve the positioning accuracy of the key points; after generating the global key point embedding, the high-dimensional embedding is projected into the three-dimensional space through a linear transformation to obtain the initial three-dimensional joint coordinate position.

[0053] In order to improve the accuracy of the initial 3D joint coordinate prediction and provide more accurate geometric reference and spatial guidance information for subsequent fusion interaction, the loss function of the initial 3D joint position is introduced here. This loss function adopts the form of Smooth L1 loss and is defined as follows:

[0054]

[0055] in, is the three-dimensional joint position coordinates of the j-th joint point predicted by the network, is the real 3D coordinate of the j-th joint point.

[0056] In the present invention, global features refer to the key point embedding generated in the first step, which is further processed by a graph convolutional network to capture the geometric relationship between joint points. One of the cores of a graph convolutional network is a learnable adjacency matrix. Traditional graph convolutional networks usually rely on a predefined adjacency matrix. However, in multimodal feature processing and global feature extraction tasks, predefined adjacency relationships may be difficult to fully adapt to the actual data distribution. In order to overcome this limitation, the present invention introduces the design of a dynamic adjacency matrix, which is a learnable kinematic matrix that can adaptively adjust the connection strength relationship between nodes according to the characteristics of the input data, thereby more accurately modeling the semantic and geometric associations between key points. By dynamically adjusting the adjacency relationship, GCN can not only capture local geometric relationships, but also mine global semantic associations, so that the model has a higher representation ability for the complex relationships between key points. It provides a more accurate semantic and geometric basis for the feature-level connection of subsequent image and point cloud features. At the same time, its adaptive characteristics ensure the generalization ability of the model to the input data and improve the overall performance of the multimodal feature fusion and interaction module.

[0057] After completing the update of the global features, they are concatenated with the features extracted from the image branch and the point cloud branch, thereby achieving effective fusion of global features and multimodal features. The specific steps include:

[0058] Joint image features extracted from image branches Combined with the updated keypoint embeddings in the following way:

[0059]

[0060] Among them, Concat represents the feature-level connection operation, updated_embedding is the key point embedding; the generated image joint feature F' j It not only preserves the characteristics of the image modality, but also integrates the global joint semantics and geometric relationships;

[0061] Joint point cloud features extracted from point cloud branches It is also combined with the updated keypoint embedding in a similar way:

[0062]

[0063] Among them, the generated point cloud joint feature F" j While retaining the modal characteristics of point clouds, the global information is introduced to supplement them.

[0064] Based on the respective modal characteristics, the image joint features and point cloud joint features of the present invention introduce key point embedding containing global information, so that the features of each modality are more consistent in spatial geometric structure and semantic information. The fused multimodal features have stronger expressive ability and provide rich feature representation for the subsequent multimodal feature interaction module.

[0065] The present invention uses the initialized three-dimensional key point coordinates to guide the further interaction and fusion between the depth map key point features and the point cloud key point features. After the above-mentioned feature fusion module, joint image features and joint point cloud features that have been fused with global information have been generated. In order to further enhance the semantic relevance and spatial consistency between these features, three-dimensional joint coordinates are introduced as additional guiding information to ensure the precise alignment of features in space and promote deep interaction between different modalities.

[0066] The fusion between joint image features and point cloud features is performed through the attention mechanism. Specifically, the image features and point cloud features are fed into a multi-head attention module, in which the relationship between each joint feature and other joint features is dynamically adjusted by calculating its attention weight. In this way, the correlation between features is significantly enhanced, while the influence of irrelevant information is suppressed.

[0067] In order to ensure that the image and point cloud features can be accurately aligned in space, the initialized 3D joint coordinates are introduced into the attention mechanism as guiding information. Specifically, the 3D coordinates affect the attention calculation through an adaptive adjustment mechanism, so that the interaction of features not only depends on the semantic information of the features themselves, but also takes into account their positional relationship in space. In the implementation process, the depth map joint features and point cloud joint features are passed to two different linear blocks respectively. In the calculation process of each linear block, the 3D joint coordinates are used as conditional input to generate a set of adjustment parameters. These adjustment parameters are used to control the behavior of the multi-head self-attention and multi-layer perceptron modules in the attention mechanism, thereby adjusting the interaction mode of the features.

[0068] During the calculation process, the depth map joint feature F' and the point cloud joint feature F" are processed separately and passed

[0069] jj

[0070] The attention mechanism is integrated, which includes the following steps:

[0071] First, two linear modules are used to j ' and F j "The features are linearly transformed to generate three vectors: query (Q), key (K), and value (V);

[0072] Next, the query, key, and value vectors of the image and point cloud features are concatenated to form a joint query, key, and value matrix - q, k, v;

[0073] By splicing the features of these two different modalities together, the model can simultaneously focus on the relationship between the two modalities, input the joint query, key, and value matrix into the self-attention module AttenBlock, and calculate the attention output;

[0074] The attention module is responsible for weighted summing of features according to the attention weights to obtain the weighted "'

[0075] Output F j and F j ;

[0076] Finally, the attention output is weighted updated according to the scaling, offset, and gating signal vectors obtained from the previous two linear blocks. The weighted update ensures that the features can adjust their weights according to the spatial relationship between joints during the interaction with the three-dimensional coordinate guidance.

[0077] After being processed by the previous attention module, the processed features are input into the graph convolutional network (GCN). At this time, GCN uses a predefined joint adjacency matrix, which is constructed based on the spatial relationship and physical connection between joints, and can capture the topological structure and geometric constraints between joints. The graph convolution operation will fuse the information of neighboring nodes through the adjacency matrix, so that the features of each node can contain the semantic and geometric information of other nodes in its neighborhood.

[0078] For the image joint features and point cloud joint features obtained after graph convolutional network processing, this embodiment adopts an iterative update method to further process the feature data. First, new three-dimensional joint coordinates are predicted based on the image joint features and point cloud joint features obtained after graph convolutional network processing; next, these updated features will redirect the entire process, continue feature fusion and interaction, and further optimize the spatial structure and geometric form of the gesture through gradual correction.

[0079] During the iterative correction process, the model uses the output features of the previous GCN to predict new 3D joint coordinates. The updated 3D coordinates not only take into account the 2D visual features, but also combine the 3D geometric information to ensure that the spatial structure of the gesture is more accurate; and the new 3D joint coordinates will replace the initialized 3D joint coordinates during the iteration process to extract more accurate joint point cloud features from the point cloud space. Based on the new 3D joint coordinates, the loss between the real 3D coordinates is calculated, and the model is further optimized. This process helps to gradually reduce the error between the predicted and real 3D coordinates. The loss function of the iterative process is expressed as:

[0080]

[0081] Among them, l s represents the loss function of one feature iteration; P j '(t) represents the predicted three-dimensional joint coordinates of the j-th joint point after the t-th iteration update.

[0082] In summary, the total loss of the gesture reconstruction model of the present invention is expressed as:

[0083]

[0084] Among them, L total is the total loss of the gesture reconstruction model; α, β, γ are the weight coefficients of the loss function; T is the number of feature iterations.

[0085] Through multiple rounds of iterations, the gesture reconstruction model of the present invention gradually corrects the predicted three-dimensional joint coordinates. Each round of iteration performs feature update and error feedback on the basis of the previous round. The optimization process provides the model with more spatial information and geometric constraints, ensuring that the final three-dimensional gesture prediction is more accurate and robust. After several rounds of iterative correction, the final three-dimensional joint coordinate prediction is obtained. At this time, the gesture reconstruction model not only accurately reflects the spatial layout of the gesture, but also has strong robustness and can adapt to different postures and external interference.

[0086] This embodiment also provides a performance comparison of a 3D gesture reconstruction method based on image point cloud multimodal fusion and joint guidance, as follows:

[0087] 1) Dataset

[0088] NYU dataset: Contains 8252 test frames and 72,757 training frames, including captured RGBD data and real annotations of hand poses. This dataset contains multiple viewpoints and annotations, suitable for training and evaluation of gesture estimation models. Each frame in the dataset contains RGBD data from three Kinect cameras, providing one front view and two side views. The data of the training set comes from a single user, while the test set includes samples from two users. In addition, the annotations of the dataset include the positions of 36 joints, and we selected 14 of them for training and evaluation.

[0089] 2) Performance comparison

[0090] Figure 3 (a) shows the effect of single-hand reconstruction, where the blue lines represent the real joints and the red lines represent the joints predicted by the present invention. Figure 3 (b) The hand-object reconstruction structure is also given. In order to verify the effectiveness of the proposed method, this embodiment makes a detailed comparison between the proposed method and the current reconstruction method based on depth image or point cloud, and also includes a method combining depth image and point cloud. The performance comparison between the present invention and the prior art is shown in Table 1.

[0091] Table 1 Comparison of the objective performance of the gesture reconstruction method of the present invention and other reconstruction methods

[0092] method Input Data MPJPE DepthPrior++ Depth 12.24 A2J Depth 8.61 AWR Depth 7.48 HandPointNet Point 10.54 HandFolding Point 8.58 HandR2N2 Point 7.27 IPNET Depth&Point 7.17 Hand DAGT Depth&Point 7.12 Method of the present invention Depth&Point 7.05

[0093] In Table 1, the input data Depth represents the depth image, Point represents the point cloud, and Depth&Point represents both modalities; this embodiment uses the mean joint error (MPJPE) to measure the error between the predicted value and the true value. In Table 1, this embodiment compares in detail the gap between the mainstream indicator of the mean joint error of the three-dimensional gesture reconstruction method based on image point cloud multimodal fusion and joint guidance of the present invention. Experiments show that the three-dimensional gesture reconstruction method based on image point cloud multimodal fusion and joint guidance of the present invention has a smaller experimental error than other reconstruction methods, and its objective performance exceeds the current mainstream gesture reconstruction method.

[0094] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A gesture reconstruction method based on image point cloud multimodal fusion and joint guidance, characterized in that: Construct a gesture reconstruction model and use it to reconstruct gestures from depth images. The reconstruction process includes the following steps:

101. Obtain a hand depth image and its point cloud data, an image feature extraction network extracts local image features and global image features from the depth image, and a point cloud feature extraction network extracts global point cloud features from the point cloud data; 102. The global image features and the global point cloud features are concatenated together, and a bias embedding vector is generated for each joint point through an induced bias layer. The vectors of all joint points constitute the global key point embedding; 103. Global key point embedding projects the high-dimensional embedding into the three-dimensional space through linear transformation to obtain the initial three-dimensional joint coordinates, and extracts local point cloud features from the point cloud data according to the three-dimensional joints; 104. Update the global key point embedding through the graph convolutional network to obtain updated global features, and concatenate the global features with the image local features and the point cloud local features to obtain image joint features and point cloud joint features; 105. Based on the attention mechanism, the image joint features and point cloud joint features are weighted respectively, and the weighted features are sent to the graph convolutional network to make full use of the topological information for further feature fusion and update; 106. Use the image joint features and point cloud joint features obtained after the graph convolution network is updated to predict the three-dimensional joint coordinates, and determine whether the iteration limit is reached. If not, update the point cloud local features according to the predicted three-dimensional joint points and return to step 104. Otherwise, the currently predicted three-dimensional joint coordinates are used as the reconstruction result.

2. The method for hand gesture reconstruction based on image point cloud multimodal fusion and joint guidance according to claim 1, characterized in that: The process of extracting local image features and global image features from the depth image by the image feature extraction network includes: The depth image is processed by the autoencoder network to obtain a two-dimensional local feature map F 2d ∈R C×H×W , where C, H, and W are the two-dimensional local feature maps F 2d Number of channels, height, and width; Two-dimensional local feature map F 2d Input to the convolutional residual network to generate the heat map corresponding to each joint point The 2D UV coordinates of the joints are estimated based on the normalized heat map, and the image features of the joints are aggregated from the local feature map according to the heat map to obtain the local image features. Where J is the number of hand joints in the depth image, d 2d is the image feature dimension corresponding to each joint; Two-dimensional local feature map F 2d A global image feature is output through global pooling.

3. The method for hand gesture reconstruction based on image point cloud multimodal fusion and joint guidance according to claim 2, characterized in that: The process of extracting local point cloud features and global point cloud features from point cloud data by the point cloud feature extraction network includes: Extract 3D local geometric features F from point cloud data through point set convolutional layer 3d ∈R N×d , where N represents the number of points in the point cloud, and d represents the point cloud feature dimension corresponding to each point; 3D local geometric features F 3d ∈R N×d After passing through the one-dimensional convolution layer and the maximum pooling layer, the global point cloud features will be obtained; The two-dimensional local feature map F 2d After being projected to the feature space of the 3D local geometric features, it is spliced ​​with the 3D local geometric features to generate the fused point cloud feature F fuse ; Using the initialized 3D joint coordinates as the starting position, further guide the point cloud feature F fuse Aggregation to generate the final local point cloud features where d 3d Indicates the point cloud feature dimension corresponding to each joint.

4. The method for hand gesture reconstruction based on image point cloud multimodal fusion and joint guidance according to claim 1 or 3, characterized in that: The process of updating the local features of the point cloud according to the three-dimensional joint points includes: for each joint point, taking the three-dimensional coordinates of the joint point as the center, using the ball query algorithm to find K neighboring points within a radius r around the position in the point cloud data, and using the geometric features of these points as the local features of the point cloud of the joint point.

5. The method for hand gesture reconstruction based on image point cloud multimodal fusion and joint guidance according to claim 1, characterized in that: A bias embedding vector is generated for each joint point through the induced bias layer. The process includes: after the global image features and the global point cloud features are concatenated, they are first copied J times to obtain J feature maps. Then each feature map is sent to the one-dimensional convolution layer, and a learnable position embedding is added to each feature map for each joint point. The learnable position embedding provides an independent bias for each joint point to generate a bias embedding vector.

6. The method for hand gesture reconstruction based on image point cloud multimodal fusion and joint guidance according to claim 1, characterized in that: The process of weighting the image joint features and point cloud joint features based on the attention mechanism includes: The image joint features and the point cloud joint features are mapped respectively through two linear modules to obtain the query vector, key vector and value vector of the image joint features and the point cloud joint features; Concatenate the query vectors, key vectors, and value vectors of the image joint features and the point cloud joint features; The concatenated query vector, key vector and value vector are input into the self-attention module to obtain the attention weights, and the image joint features and point cloud joint features are weighted using the attention weights.

7. The method for hand gesture reconstruction based on image point cloud multimodal fusion and joint guidance according to claim 1, characterized in that: The loss function of the gesture reconstruction model during training is expressed as: Among them, L total is the total loss of the gesture reconstruction model; α, β, γ are the weight coefficients of the loss function; l uv is the loss function of the image feature extraction network; l init The loss function of the point cloud feature extraction network; is the loss function at the tth feature iteration, and T is the number of feature iterations.

8. The method for hand gesture reconstruction based on image point cloud multimodal fusion and joint guidance according to claim 7, characterized in that: The loss function of the image feature extraction network is expressed as: Where J is the number of hand joints in the depth image; L(x) is the standard Smooth L1 loss function; (U' j ,V j ') represents the UV coordinate of the jth hand joint point calculated by the normalized heat map, (U j ,V j ) is the real UV coordinate of the j-th joint point.

9. The method for hand gesture reconstruction based on image point cloud multimodal fusion and joint guidance according to claim 7, characterized in that: The loss function of the point cloud feature extraction network is expressed as: Where J is the number of hand joints in the depth image; is the three-dimensional joint position coordinates of the j-th joint point predicted by the network, is the real 3D coordinate of the j-th joint point.

10. The method for hand gesture reconstruction based on image point cloud multimodal fusion and joint guidance according to claim 7, characterized in that: The loss function during feature iteration is expressed as: Among them, l s represents the loss function of one feature iteration; J is the number of hand joints in the depth image; P j '(t) represents the three-dimensional joint coordinates predicted by the j-th joint point after the t-th iteration update, is the real 3D coordinate of the j-th joint point.

Citation Information

Patent Citations

  • A method and system for three-dimensional gesture motion reconstruction

    CN102262783A

  • Pesticide application method of visual spraying robot based on augmented reality remote control

    CN118202987A

  • Method for identifying dynamic gesture under point cloud data

    CN118887736A

  • Depth image gesture estimation method based on semi-supervised learning

    CN111797692A

  • Cascaded three-dimensional hand posture estimation method

    CN117351573A