Hand posture estimation method based on RGB and Depth images
By combining feature fusion methods of RGB and Depth images, and utilizing PointNet++ encoders and cross-modal feature fusion networks, the problem of insufficient hand pose estimation accuracy in existing technologies is solved, achieving higher accuracy and robustness in hand pose estimation, and improving the dexterity of the hand.
Patent Information
- Application Number
- CN202511092064.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-28
AI Technical Summary
Existing hand pose estimation methods based on RGB and depth images are insufficient in terms of accuracy and robustness, especially in complex environments where it is difficult to accurately identify hand poses.
A fusion method based on RGB and Depth images is adopted. By extracting features from the RGB image and initial keypoint features from the Depth image, and combining the PointNet++ encoder and cross-modal feature fusion network, the feature fusion is performed using bidirectional attention and cross-attention mechanisms to improve the accuracy of hand pose estimation.
It improves the accuracy and robustness of hand pose estimation, especially the ability to identify key hand points in complex environments, thereby enhancing the operational precision of dexterous hands.
Smart Images

Figure CN121033931A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of hand pose estimation and dexterous hand teleoperation, and particularly relates to a hand pose estimation method based on RGB and depth images. BACKGROUND
[0002] Dexterous manipulation is a fundamental aspect of human-robot interaction, enabling the execution of complex tasks such as grasping and teleoperation. Currently, there are two main approaches to teleoperating dexterous robotic hands: wearable devices and vision-based methods. Wearable devices, including data gloves and electromyography (EMG) sensors, are commonly used to capture hand poses. These setups often require expensive hardware, specialized engineering, expert operators, or devices that limit the natural fluid motion of the demonstrator's hand. In contrast, vision-based methods typically utilize depth cameras to obtain three-dimensional information of the hand. Vision-based teleoperation effectively addresses many of the inherent limitations of wearable devices due to its non-contact and simplicity.
[0003] Teleoperating dexterous robotic hands faces significant challenges due to their high degree of freedom and complex structure. These complexities pose considerable demands on vision-based teleoperation methods. Specifically, accurate control of dexterous hands requires highly accurate three-dimensional hand pose estimation. Existing hand pose estimation techniques can generally be divided into two methods: RGB-based methods and depth-based methods. Although some progress has been made on single RGB or depth methods, they still face challenges due to the inherent limitations of their single-modal approach. RGB-based methods are susceptible to color ambiguity and lack of detailed local geometric information. In contrast, depth-based methods often encounter noise and depth inconsistency, especially at the edges of the hand and objects, which become more pronounced during motion.
[0004] Therefore, there is a need for a hand pose estimation method based on RGB and depth images. SUMMARY
[0005] Therefore, the present application provides a hand pose estimation method based on RGB and depth images, extracts RGB features based on hand RGB images, extracts initial hand key point features based on hand depth images, fuses RGB features and initial hand key point features, and performs hand pose estimation. The problem of insufficient hand pose estimation accuracy is solved.
[0006] To this end, the present application provides the following technical solutions: A hand pose estimation method based on RGB and depth images, comprising: extracting RGB features based on the hand RGB image, and extracting initial hand key point features based on the hand depth image; projecting the initial hand key point features into a three-dimensional space to obtain three-dimensional point cloud features; fusing the RGB features and the three-dimensional point cloud features through an encoder to obtain aggregated point cloud features; fusing the aggregated point cloud features and the RGB features to obtain fusion features; obtaining a hand pose by using the fusion features through a fully connected layer.
[0007] Further, the RGB features based on the hand RGB image are extracted, including: obtaining a feature map of the hand RGB image through a linear layer, and dividing the feature map into a plurality of tokens; obtaining key tokens, value tokens and query tokens corresponding to the feature map through linear projection; constructing an affinity matrix of the feature map based on the key tokens and the query tokens; calculating the local density of the visual token based on the affinity matrix; performing clustering based on the local density of the visual token to obtain a merged token; splicing the merged token and the visual token to obtain a fusion token; obtaining RGB encoded multi-scale features based on the fusion token; obtaining the RGB features by using the RGB encoded multi-scale features through a progressive decoder.
[0008] Further, the initial hand key point features based on the hand depth image are extracted, including: inputting the hand depth image into a depth key point extraction module to obtain the initial hand key point features; The depth key point extraction module comprises a resnet18 network and a fully connected layer.
[0009] Further, the aggregated point cloud features obtained by fusing the RGB features and the three-dimensional point cloud features through the encoder include:
[0010] wherein, the aggregated point cloud features, the RGB features, the three-dimensional point cloud features, the encoder is a point cloud feature encoder of PointNet++, the weight is a learnable weight. and the three-dimensional point cloud features, the encoder is a point cloud feature encoder of PointNet++, the weight is a learnable weight.
[0011] Further, the fusion features obtained by fusing the aggregated point cloud features and the RGB features include: Token generated by linear layer for RGB feature and three-dimensional point cloud feature; Pre-fusion feature is obtained based on the token through a bidirectional attention mechanism and a cross-attention mechanism; The pre-fusion feature is weighted and summed to obtain a fusion feature.
[0012] Further, the multi-scale feature encoded by RGB is obtained by a progressive decoder, including:
[0013]
[0014]
[0015] wherein, represents the processed RGB feature, represents an intermediate feature, as the RGB feature .
[0016] Further, the pre-fusion feature is obtained based on the token through a bidirectional attention mechanism and a cross-attention mechanism, including:
[0017] wherein, BA is Bi-attention, which is a bidirectional routing attention mechanism, and CA is a cross-attention mechanism; is a bidirectional RGB pre-fusion feature, is a bidirectional aggregated point cloud pre-fusion feature, is a cross-RGB pre-fusion feature, is a cross-aggregated point cloud pre-fusion feature; is an RGB hand query token, is an RGB hand key token, is a depth hand value token, is a depth hand query token, is a depth hand key token, is an RGB hand value token.
[0018] Advantages and positive effects of the present application: The method improves the accuracy of hand feature extraction by extracting and fusing hand features of RGB images and depth images respectively.
[0019] 1) The RGB feature is extracted by a region-aware token clustering transformer combined with a progressive decoder, and the region-aware token clustering transformer allocates more tokens to high-value image regions, so that the tokens can focus on key regions, learn more comprehensive image representations, and improve the accuracy of RGB feature extraction.
[0020] 2) The cross-modal feature fusion module obtains fused features through an attention mechanism on the RGB features and the three-dimensional point cloud features. By integrating the complementary advantages of the two modalities, the limitations of a single modality in recognizing hand key points can be alleviated, thereby improving the accuracy of recognizing hand key points. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0022] Figure 1 is a flowchart of the hand pose estimation method based on RGB and Depth images in embodiment 1; Figure 2 is a flowchart of the hand pose estimation method based on RGB and Depth images in embodiment 1; Figure 3 is an architecture diagram of the RGB feature extraction module in embodiment 1; Figure 4 is a key point visualization diagram of the hand pose estimation NYU dataset and DexYCB dataset in embodiment 1; Figure 5 is a comparative experiment diagram of the average error of the hand pose estimation in embodiment 1; Figure 6 is an experimental diagram of the hand pose estimation for dexterous hand remote control in embodiment 1; Figure 7 is a flowchart of the hand pose estimation method based on RGB and Depth images in embodiment 2. DETAILED DESCRIPTION
[0023] In order to make the person skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and in the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] The present application provides a hand pose estimation method based on RGB and Depth images, comprising: S1: using a zed2i depth camera to collect RGB images and Depth images; S2: pre-processing the collected images: using a YOLOv5 algorithm to detect the hand target, and cutting the RGB and Depth images according to the detected bounding box to obtain hand RGB images and hand Depth images; S3: training and constructing a hand pose estimation model using a public dataset NYU dataset, and then inputting the pre-processed hand images into the hand pose estimation model for real-time detection of hand key points to obtain the hand key points; S4: mapping the hand key points to the key points of the dexterous hand to control the dexterous hand.
[0026] Embodiment 1 A hand pose estimation method based on RGB and Depth images, comprising: S3 comprises: constructing a hand pose estimation model based on RGB and Depth images for hand key point detection: 1) inputting the hand RGB image into an RGB feature extraction module; extracting the RGB feature in the region-aware token cluster transformer (RTCT) and the progressive decoder; Wherein, the RTCT is a cluster-based transformer, including: region-aware token cluster attention (RTCA) and transformer block can extract multi-scale RGB features for subsequent processing.
[0027] 2) input the hand depth image into the depth feature extraction module to extract the initial hand key point feature.
[0028] The depth feature extraction module comprises a resnet18 and a fully connected layer.
[0029] 3) project the initial hand key point feature into a three-dimensional space to obtain a depth feature, and perform interpolation operation on the depth feature to obtain a three-dimensional point cloud feature.
[0030] 4) input the three-dimensional point cloud feature and the RGB feature into the Encoder of PointNet++ for local aggregation to obtain an aggregated point cloud feature. 3) input the aggregated point cloud feature and the RGB feature into the cross-modal feature fusion network to perform cross-modal feature interaction to obtain a fusion feature, and input the fusion feature into a linear layer to obtain a hand key point.
[0031] The Region-aware Token Cluster Transformer extracts multi-scale RGB features, and the process includes: An image I is received, and a set of feature maps is obtained through linear projection and a series of Region-aware Token Cluster Attention.
[0032] 1) the data processing process of the Region-aware Token Cluster Attention module includes: 1) divide each feature map into non-overlapping tokens , and perform linear projection to obtain key, value and query, denoted as: . Perform average pooling operation on to obtain , and generate affinity matrix A through The formula of the affinity matrix is:
[0033] 2) based on the affinity matrix, determine the local density of the visual token through (k-Nearest Neighbors-based Density Peak Clustering, DPC-kNN), and the calculation formula is:
[0034] wherein, denotes the visual token with index i, j in A; denotes the visual token local density of visual tokens, KNN represents the k nearest neighbors of visual tokens .
[0035] 3) Determine the cluster center of visual tokens based on the local density of visual tokens: For each visual token, define a distance indicator , representing the minimum distance between the measured visual token and other tokens with higher density. The formula for defining the distance of each visual token is:
[0036] wherein represents the distance indicator, represents the local density. The score of the cluster center visual token is the product of the local density and the minimum distance , denoted as . The higher the score, the higher the possibility of the visual token having greater density and distance as a cluster center. The visual tokens with the top k scores are taken as cluster centers, and the remaining tokens are assigned to the nearest cluster center with higher density.
[0037] 4) Cluster each visual token according to the cluster center to obtain the merged token, and the formula is:
[0038] wherein represents the set of the i-th cluster, represents the visual token, represents the importance score corresponding to each visual token, represents the merged token.
[0039] Connect the merged token with the visual token to obtain the fusion token, and the formula is:
[0040] wherein represents the connected key token, represents the connected value token.
[0041] 5) Input the token into the transformer block for processing, and integrate the importance score into the attention mechanism to obtain the RGB-encoded multi-scale feature:
[0042] wherein represents the channel dimension of the query, represents the RGB-encoded multi-scale feature, represents the importance score.
[0043] 2. Progressive decoder, including: point-wise convolution and depthwise separable convolution; RGB features are obtained by the progressive decoder based on multi-scale features of RGB coding, including:
[0044]
[0045]
[0046] wherein, represents the processed RGB feature, represents the intermediate feature, as the final RGB feature .
[0047] 3. Cross-modal feature fusion module, responsible for fusing features of two modalities: 1) linear layer is applied to RGB feature and aggregated point cloud feature to generate hand feature token . 2) based on the hand feature, pre-fusion features are obtained through bidirectional attention mechanism and cross attention mechanism.
[0048]
[0049] wherein, BA is Bi-attention, which is a bidirectional routing attention mechanism, and CA is a cross attention mechanism; is a bidirectional RGB pre-fusion feature, is a bidirectional aggregated point cloud pre-fusion feature, is a cross RGB pre-fusion feature, is a cross aggregated point cloud pre-fusion feature; is an RGB hand query token, is an RGB hand key token, is a depth hand value token, is a depth hand query token, is a depth hand key token, is an RGB hand value token.
[0050] The pre-fusion features are weighted and summed to obtain the fusion feature:
[0051] wherein, is a trainable weight, is the fusion feature.
[0052] Embodiment 2 A hand pose estimation method based on RGB and Depth images includes: S1: Use a zed2i depth camera to acquire RGB and depth images; the image size is [missing information]. The number of channels is 3 and 1 respectively.
[0053] S2: Preprocess the acquired images: Use the pre-trained YOLOv5 algorithm to detect targets on the hand, and crop the RGB and Depth images according to the detected bounding boxes to obtain the hand image; S3: Use the NYU dataset to train and build a hand pose estimation model, and then input the preprocessed hand image into the hand pose estimation model to perform real-time detection of hand key points to obtain the key points of the hand.
[0054] The structure of the hand pose estimation model is as follows: Figure 3 As shown, it includes: a feature extraction module, a key point feature aggregation module, and a cross-modal fusion module.
[0055] The feature extraction module includes an RGB feature extraction module and a depth keypoint extraction module.
[0056] The RGB feature extraction module includes: Region-aware Token Cluster Transformer and progressive decoder.
[0057] The deep key point extraction module includes ResNet18 and a fully connected layer.
[0058] 1. Input the RGB image into the RGB feature extraction module to obtain the RGB features. .
[0059] 1) The input image is processed through a linear layer, and a hierarchical feature map is generated through a series of Region-aware Token ClusterAttention (RTCA) blocks. .
[0060] Assume the feature size of the i-th RTFA block is Feature map Divided into Non-overlapping tokens , . respectively through matrix Linear projection yields:
[0061] in, To query the token, For key tokens, is a value token.
[0062] Construct a region-aware affinity matrix.
[0063] to compress the spatial redundancy of local regions. Then, the affinity matrix A is calculated by multiplying and transpose, the formula is expressed as:
[0064] Based on the affinity matrix, the local density of visual tokens is determined by (k-Nearest Neighbors-based Density Peak Clustering, DPC-kNN), and the formula is:
[0065] wherein, denotes the visual token with index i, j in A; denotes the local density of visual token , and KNN denotes the k nearest neighbors of visual token .
[0066] Based on the local density of visual tokens, the cluster centers of visual tokens are determined: For each visual token, a distance indicator is defined, which represents the minimum distance between the measured visual token and other tokens with higher density. The formula for defining for each visual token is:
[0067] wherein, denotes the distance indicator, denotes the local density. The score of the cluster center visual token is the product of the local density and the minimum distance , denoted as . The higher the score, the higher the possibility of the visual token as a cluster center. The visual tokens with the top k scores are selected as cluster centers, and the remaining tokens are assigned to the nearest cluster center with higher density.
[0068] According to the cluster centers, the visual tokens are clustered to obtain merged tokens, and the formula is:
[0069] wherein, Let i represent the set of the i-th cluster. Indicates a visual token. This represents the importance score for each visual token. This indicates a merge token.
[0070] The merge token is obtained by connecting the merge token and the visual token, as shown in the formula:
[0071] in, This indicates the key token after connection. This represents the value token after the connection.
[0072] By incorporating importance scores into the attention mechanism, multi-scale features encoded in RGB are obtained:
[0073] in, Indicates the channel dimension of the query. This represents the multi-scale features of RGB encoding. Indicates the importance score.
[0074] 2. Progressive decoder, including: multiple MCM modules (multi-level convolution modules); the MCM modules include: pointwise convolution and depthwise separable convolution; Multi-scale features based on RGB encoding obtain RGB features through a progressive decoder, including:
[0075]
[0076]
[0077] in, Indicates the processed RGB features. Indicates intermediate features, As the final RGB feature .
[0078] The deep key point extraction module includes ResNet18 and a fully connected layer.
[0079] Deep features extracted after ResNet18 .
[0080] Interpolation is performed on the depth features to obtain 3D point cloud features.
[0081] The key point feature aggregation module aggregates 3D point cloud features and RGB features to generate aggregated point cloud features.
[0082] The RGB features and 3D point cloud features are input into the PointNet++ Encoder network for feature aggregation near keypoint features, resulting in aggregated point cloud features, expressed by the formula:
[0083] in, For 3D point cloud features, For the generated RGB features, the Encoder is a point cloud feature encoder from PointNet++. and These are learnable weights; This represents the aggregated point cloud features.
[0084] 4. Cross-modal fusion module: fuses RGB features and aggregated point cloud features to obtain fused features.
[0085] 1) RGB features and aggregated point cloud features After passing through a linear layer, a hand feature token is generated. ; 2) Based on hand features, pre-fused features are obtained through bidirectional attention and cross-attention mechanisms.
[0086]
[0087] Among them, BA stands for Bi-attention, which is a bidirectional routing attention mechanism, and CA stands for cross-attention mechanism.
[0088] The pre-fused features are weighted and summed to obtain the fused features:
[0089] in, For trainable weights, Fusion characteristics.
[0090] 5. Input the fused features into the linear layer to obtain key points.
[0091] 6. Map the key points of the human hand to the key points of the dexterous hand to control the dexterous hand.
[0092] Example 3 The effectiveness of this method was verified through comparative experiments. The average error of the experimental results is shown in Tables 1 and 2: Table 1
[0093] Table 2
[0094] In summary, this method, which combines RGB and Depth image features for hand keypoint extraction, improves the accuracy of hand keypoint capture, thereby enhancing the operational precision of dexterous hands.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hand pose estimation method based on RGB and Depth images, characterized in that, include: RGB features are extracted from the RGB image of the hand, and initial key point features of the hand are extracted from the Depth image of the hand. The initial hand key point features are projected into three-dimensional space to obtain three-dimensional point cloud features; Aggregated point cloud features are obtained by fusing RGB features and 3D point cloud features using an encoder; The fused features are obtained by fusing and aggregating point cloud features and RGB features; Hand pose is obtained through a fully connected layer by utilizing fusion features.
2. The method according to claim 1, characterized in that, RGB features are extracted from hand RGB images, including: The feature map of the hand RGB image is obtained through a linear layer, and the feature map is divided into several tokens; The key token, value token, and query token corresponding to the feature map are obtained through linear projection. Construct an affinity matrix for the feature map based on key tokens and query tokens; The local density of the visual token is calculated based on the affinity matrix; Clustering based on the local density of visual tokens to obtain merged tokens; The merge token and the visual token are concatenated to obtain the fusion token; Multi-scale features of RGB encoding are obtained based on the fusion token; RGB features are obtained through a progressive decoder by utilizing the multi-scale features of RGB encoding.
3. The method according to claim 1, characterized in that, The extraction of initial hand key point features based on the hand depth image includes: The hand depth image is input into the depth key point extraction module to obtain the initial hand key point features; The deep key point extraction module includes: a ResNet18 network and a fully connected layer.
4. The method according to claim 1, characterized in that, The process of obtaining aggregated point cloud features by fusing RGB features and 3D point cloud features through an encoder includes: in, For aggregated point cloud features, It is an RGB feature. For 3D point cloud features, the Encoder is a PointNet++ point cloud feature encoder. and These are learnable weights.
5. The method according to claim 1, characterized in that, The fusion of point cloud features and RGB features to obtain fused features includes: Tokens that generate RGB features and 3D point cloud features through linear layers; Pre-fused features are obtained based on the token using a bidirectional attention mechanism and a cross-attention mechanism; The pre-fused features are weighted and summed to obtain the fused features.
6. The method according to claim 2, characterized in that, The method of obtaining RGB features through a progressive decoder using RGB-encoded multi-scale features includes: in, Indicates the processed RGB features. Indicates intermediate features. As an RGB feature .
7. The method according to claim 5, characterized in that, The method of obtaining pre-fused features based on the token through bidirectional attention and cross-attention mechanisms includes: Among them, BA stands for Bi-attention, which is a bidirectional routing attention mechanism, and CA stands for cross-attention mechanism; It is a bidirectional RGB pre-fusion feature. It is a bidirectional aggregated point cloud pre-fusion feature. Cross-RGB pre-fusion features Cross-aggregated point cloud pre-fusion features; It is an RGB hand query token. It's an RGB hand button token. It is a depth hand value token. It is a deep hand query token. It is a deep hand key token. It is an RGB hand value token.