A hand-object estimation method, system, electronic device and storage medium

Through three-dimensional interactive modeling and vertex comparison learning, the problem of difficult to capture hand-object interaction details and small object modeling in the existing technology is solved, and the performance and accuracy of hand-object posture estimation are improved.

CN119920014BActive Publication Date: 2025-06-10NANCHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510406323.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-10
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The prior art is difficult to capture hand-object interaction details when estimating interactive hand-object poses from monocular images, especially inadequate modeling and feature acquisition of small objects in complex contexts, resulting in limited performance of 3D interactive hand-object pose estimation methods.

Method used

Through the three-dimensional interactive modeling method, the mapping of multi-scale two-dimensional to three-dimensional spatial features is realized, the fine-grained interaction correlation between the hand and the object in the three-dimensional space is captured, and the model's ability to process hand-object pose estimation in complex scenarios is enhanced through the comparative learning of positive and negative samples between the hand vertices in the three-dimensional space.

Benefits of technology

The performance of the three-dimensional interactive hand-object pose estimation method is improved, the interactive detail representation and model processing capabilities of complex backgrounds are enhanced, and more accurate hand-object interaction pose estimation is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920014B_ABST
    Figure CN119920014B_ABST
Patent Text Reader

Abstract

This application belongs to the field of computer vision technology and discloses a hand-object estimation method, system, electronic device, and storage medium. The method includes: obtaining a hand-object interaction image and extracting deep feature maps of the hand and the object; predicting the 2D coordinates of the hand vertices and the 2D coordinates of the object; performing pose alignment on the 2D coordinates of the hand vertices and the deep feature map of the hand, and on the 2D coordinates of the object and the deep feature map of the object respectively, to obtain the pose-aligned features of the hand and the object, and then mapping them to the 3D space to obtain the 3D space features of the hand and the object; capturing the fine-grained interactivity between the 3D space features of the hand and the object through a cross-attention mechanism to generate enhanced features of the hand and the object; using contrastive learning to enhance the similarity between hand vertices and their difference from the environmental voxels to obtain optimized hand space features; and finally decoding the enhanced features of the object and the optimized hand space features to output the hand-object estimation result. This method can improve the performance of 3D interactive hand-object pose estimation methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a hand-object estimation method, system, electronic device and storage medium based on three-dimensional interactive modeling and vertex contrast learning. Background Art

[0002] 3D hand-object pose estimation is one of the popular research directions in the field of computer vision and plays a crucial role in fields such as virtual reality and human-computer interaction. This task generally estimates or reconstructs the hand and the object being manipulated in an image. In recent years, with the rapid development of deep neural network technology and the emergence of large-scale hand-object interaction datasets, many breakthrough results have been achieved in the research on the poses of hands and objects. Currently, relatively good results have been obtained in estimating the pose of a single hand and a single object from a monocular image. However, due to the high degree of freedom of the hand and the mutual occlusion caused by the frequent contact between the hand and the object, it is more challenging to estimate the pose of an interactive hand-object from a monocular image.

[0003] Currently, the interactive hand-object pose estimation method has obtained relatively satisfactory results, but there are still some problems: (1) The core of the 3D interactive hand-object pose estimation task lies in accurately modeling the interaction relationship between the hand and the object. In recent years, although the basic position and pose characteristics of the hand and the object can be captured, there are deficiencies in the details of the hand-object interaction, especially since the interaction between the hand and the object mainly occurs on the surface, and the surface contour information of the hand and the object is not well utilized. (2) Jointly estimating the hand-object pose in a complex background, especially modeling small objects, is challenging. Because small objects occupy a relatively small area in the image, resulting in less effective information extracted from the image. In addition, in a complex background, the image may contain multiple objects that may be similar in color or shape to the object being manipulated. However, the current mainstream methods have insufficient acquisition of the characteristics of these small objects or similar objects. Therefore, the above two problems limit the performance of the 3D interactive hand-object pose estimation method. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a hand-object estimation method, system, electronic device and storage medium, aiming to improve the performance of the three-dimensional interactive hand-object pose estimation method. On the one hand, through the three-dimensional interactive modeling method, the mapping of multi-scale two-dimensional to three-dimensional space features is realized, and at the same time, the fine-grained interaction correlation between the hand and the object in the three-dimensional space is captured, enhancing the representation of interaction details; on the other hand, through the contrast learning of positive and negative samples between the hand vertices in the three-dimensional space, the similarity and difference between the hand mesh vertices and the environmental voxels in a complex background are enhanced, thereby strengthening the model's ability to process hand-object pose estimation in complex scenarios.

[0005] To solve the above technical problems, the present application is implemented as follows:

[0006] In a first aspect, an embodiment of the present application provides a hand-object estimation method, and the method includes:

[0007] Obtain a hand-object interaction image;

[0008] Extract features from the hand-object interaction image to respectively obtain a hand deep feature map and an object deep feature map;

[0009] Predict hand joint points for the hand deep feature map to obtain hand vertex 2D coordinates; perform object heatmap encoding on the object deep feature map to obtain object 2D coordinates;

[0010] Align the hand vertex 2D coordinates and the hand deep feature map in terms of pose to obtain a hand pose alignment feature; align the object 2D coordinates and the object deep feature map in terms of pose to obtain an object pose alignment feature;

[0011] Perform 3D feature mapping on the hand pose alignment feature and the object pose alignment feature to respectively obtain a hand 3D spatial feature and an object 3D spatial feature;

[0012] According to the cross-attention mechanism, capture the fine-grained interaction relationship between the hand 3D spatial feature and the object 3D spatial feature to obtain a hand enhanced feature and an object enhanced feature;

[0013] Use contrastive learning of positive and negative samples to enhance the similarity between hand vertices in the hand enhanced feature and the difference between hand vertices in the hand enhanced feature and environmental voxels to obtain a hand spatial optimization feature;

[0014] Decode the object enhanced feature and the hand spatial optimization feature to obtain a hand-object estimation result.

[0015] As an optional implementation manner of the first aspect of the present application, the step of extracting features from the hand-object interaction image to respectively obtain a hand deep feature map and an object deep feature map includes: using a Restnet50-FPN single and double branch hybrid network to perform shallow feature extraction on the hand-object interaction image to obtain a hand shallow feature and an object shallow feature; using multi-scale backbones to respectively perform multi-scale feature extraction on the hand shallow feature and the object shallow feature to obtain a hand deep feature map and an object deep feature map.

[0016] As an alternative implementation of the first aspect of the present application, the steps of predicting hand joint points from the hand deep feature map to obtain the 2D coordinates of the hand vertices include: extracting hand features from the hand deep feature map through an hourglass module, introducing residual learning through residual connections, and processing the hand features through basic blocks to obtain three different feature vectors, namely, a two-dimensional heat map, 2D joint point coordinates, and a latent feature map.

[0017] As an alternative implementation of the first aspect of the present application, the steps of encoding the object deep feature map to obtain the 2D coordinates of the object include: passing the object deep feature map through a convolutional block and an activation function to obtain a two-dimensional heat map representation of the object, where the two-dimensional heat map representation of the object includes a height H and a width W; calculating the mean values on the height H and the width W to quantify the two-dimensional heat map coordinates and convert them into corresponding one-dimensional vectors respectively; assuming that the tensors of dimension W and H represent the heat map indices on the width W and the height H respectively; multiplying the normalized one-dimensional vectors by the corresponding heat map indices to obtain the 2D coordinates of the object.

[0018] As an alternative implementation of the first aspect of the present application, before the step of performing 3D feature mapping on the hand pose alignment feature and the object pose alignment feature, it further includes: preprocessing the hand pose alignment feature and the object pose alignment feature, where the preprocessing includes: performing four-fold downsampling on the hand pose alignment feature and simplifying the number of hand vertices; simplifying the number of object vertices of the object pose alignment feature.

[0019] As an alternative implementation of the first aspect of the present application, the steps of capturing the fine-grained interaction relationship between the hand 3D spatial feature and the object 3D spatial feature according to the cross-attention mechanism to obtain the hand enhanced feature and the object enhanced feature include: using a one-dimensional convolutional layer to extract the query values and feature values of the hand 3D spatial feature and the object 3D spatial feature respectively; obtaining the corresponding attention-weighted features according to the query values, the feature values, and the total feature dimension; using a one-dimensional convolutional layer to fuse the hand 3D spatial feature and the object 3D spatial feature with the corresponding attention-weighted features respectively to obtain the hand enhanced feature and the object enhanced feature.

[0020] As an alternative implementation of the first aspect of the present application, the steps of enhancing the similarity between hand vertices in the hand enhancement feature and enhancing the difference between hand vertices and environmental voxels in the hand enhancement feature by contrastive learning of positive and negative samples to obtain an optimized hand spatial feature include: using environmental voxels to form an embedding representation of vertex coordinates, where each environmental voxel represents the feature information of a local area in the hand enhancement feature to capture the similarity and difference between data. For each hand vertex, the positive sample is the predicted hand vertex coordinate value, and the negative sample is the environmental voxel; based on the hand vertex and the environmental voxel, using the hand InfoNCE loss function to obtain the optimized hand spatial feature, and the hand InfoNCE loss function includes a first contrastive loss function between hand vertices and a second contrastive loss function between hand vertices and environmental voxels.

[0021] In a second aspect, an embodiment of the present application provides a hand-object estimation system, which includes:

[0022] An image acquisition module for acquiring hand-object interaction images;

[0023] A three-dimensional dense modeling module for feature extraction of the hand-object interaction image to respectively obtain a hand deep feature map and an object deep feature map; predicting hand joint points for the hand deep feature map to obtain 2D coordinates of hand vertices; performing object heatmap encoding on the object deep feature map to obtain 2D coordinates of the object; aligning the 2D coordinates of the hand vertices and the hand deep feature map in terms of pose to obtain a hand pose alignment feature; aligning the 2D coordinates of the object and the object deep feature map in terms of pose to obtain an object pose alignment feature; performing 3D feature mapping on the hand pose alignment feature and the object pose alignment feature to respectively obtain a hand 3D spatial feature and an object 3D spatial feature; capturing the fine-grained interaction relationship between the hand 3D spatial feature and the object 3D spatial feature according to the cross-attention mechanism to obtain a hand enhancement feature and an object enhancement feature;

[0024] A vertex contrastive learning module for enhancing the similarity between hand vertices in the hand enhancement feature and enhancing the difference between hand vertices and environmental voxels in the hand enhancement feature by contrastive learning of positive and negative samples to obtain an optimized hand spatial feature;

[0025] An output image module for decoding the object enhancement feature and the optimized hand spatial feature to obtain a hand-object estimation result.

[0026] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0027] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0028] Compared with the prior art, the present invention proposes a hand-object estimation method: by acquiring hand-object interaction images and separately extracting the deep feature maps of the hand and the object, the key visual information of the hand and the object can be captured, providing a data basis for subsequent processing; predicting the hand joint points of the hand deep feature map to obtain the 2D coordinates of the hand vertices, and encoding the object deep feature map to obtain the 2D coordinates of the object, which can accurately represent the spatial position information of the hand and the object; aligning the 2D coordinates of the hand vertices with the hand deep feature map to obtain the hand pose alignment feature, and at the same time aligning the 2D coordinates of the object with the object deep feature map to obtain the object pose alignment feature, ensuring the consistency of the spatial positions of the hand and the object and reducing the estimation error; performing 3D feature mapping on the pose alignment features of the hand and the object to respectively obtain the 3D spatial features of the hand and the 3D spatial features of the object, providing three-dimensional spatial information for capturing the interaction relationship between the hand and the object; capturing the fine-grained interaction relationship between the 3D spatial features of the hand and the 3D spatial features of the object through the cross-attention mechanism to obtain the enhanced hand feature and the enhanced object feature, which can more accurately represent the interaction details between the hand and the object; using contrastive learning of positive and negative samples to enhance the similarity between the hand vertices in the enhanced hand feature and the difference between the hand vertices and the environmental voxels to obtain the optimized hand spatial feature, further improving the representation ability of the hand feature and enhancing the robustness of the estimation; finally, decoding the enhanced object feature and the optimized hand spatial feature to achieve accurate estimation of hand-object interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a flowchart of a hand-object estimation method provided by the first embodiment of the present invention;

[0030] Figure 2 is a process division diagram of a hand-object estimation method in the first embodiment of the present invention;

[0031] Figure 3 is a diagram showing the 2D-3D vertex mapping process in the first embodiment of the present invention;

[0032] Figure 4 is a diagram of the 3D hand-object cross-attention method in the first embodiment of the present invention;

[0033] Figure 5 It is a diagram showing the grid vertex contrast learning method in the first embodiment of the present invention;

[0034] Figure 6 It is a schematic structural diagram of a hand-object estimation system provided in the second embodiment of the present invention. Detailed implementation manners

[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0036] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.

[0037] In order to illustrate the technical solutions described in the present application, the following will be described through specific embodiments.

[0038] Embodiment 1

[0039] Please refer to Figure 1 , which is a flowchart of a hand-object estimation method proposed in the first embodiment of the present application. Please refer to Figure 2 , which is a process division diagram of a hand-object estimation method in the first embodiment of the present application. The steps of the proposed method are as follows.

[0040] Step S01: Obtain a hand-object interaction image.

[0041] Step S02: Extract features from the hand-object interaction image to obtain a hand deep feature map and an object deep feature map respectively.

[0042] Step S03: Predict hand joint points on the hand deep feature map to obtain the 2D coordinates of hand vertices; perform object heatmap encoding on the object deep feature map to obtain the 2D coordinates of the object.

[0043] Step S04: Align the 2D coordinates of the hand vertices and the hand deep feature map in terms of pose to obtain the hand pose-aligned feature; align the 2D coordinates of the object and the object deep feature map in terms of pose to obtain the object pose-aligned feature.

[0044] Step S05: Perform 3D feature mapping on the hand pose-aligned feature and the object pose-aligned feature to obtain the hand 3D spatial feature and the object 3D spatial feature respectively.

[0045] Step S06: According to the cross-attention mechanism, capture the fine-grained interaction relationship between the hand 3D spatial feature and the object 3D spatial feature to obtain the hand enhanced feature and the object enhanced feature.

[0046] It should be noted that in this embodiment, the processes of steps S02 to S06 are the three-dimensional interaction modeling (3D Interaction Modeling, DIM) process.

[0047] For step S02, in some embodiments, a Restnet50-FPN single and double branch hybrid network is used to perform shallow feature extraction on the hand-object interaction image to obtain the hand shallow feature and the object shallow feature; multi-scale backbones are used to perform multi-scale feature extraction on the hand shallow feature and the object shallow feature respectively to obtain the hand deep feature map and the object deep feature map.

[0048] Specifically, modifications are made to the residual network Restnet50, which is divided into 5 levels: level 0 to level 4. Among them, levels 0 and 1 are still single-branch structures, and the hand and the object are used as a whole foreground for feature learning in the early stage. The intermediate stages, levels 2 and 3, are double-branch structures, which respectively learn the hand and object features. Level 4 is a single-branch structure, unifying the feature spaces of the hand and the object for subsequent interaction operations. Then, two pyramid structures (FPN) are used to continue the double-branch output. The first branch performs feature fusion operations on the feature maps output from levels 1-4, and the second branch obtains the feature maps at different scales to strengthen the feature representations of the hand and the object. The entire multi-scale output backbone network outputs the deep feature map representations of the hand and the object 、 , as well as the multi-scale feature sets of the hand and the object respectively. Among them is the hand deep feature map representation, is the object deep feature map representation, and H and W are the height and width of the input image respectively.

[0049] For step S03, in some embodiments, the deep hand feature map is used to extract hand features through an hourglass module, introduce residual learning through residual connections, and process the hand features through basic blocks to obtain three different feature vectors, namely a two-dimensional heat map, 2D joint coordinates, and a latent feature map.

[0050] Specifically, for the hand 2D joint coordinates , the deep hand feature map output by the backbone network is passed through an hourglass module (Hourglass module), which focuses on the extraction of hand features. Subsequently, residual learning is introduced through residual connections to enhance the representation ability of the network. Finally, the features are further processed through basic blocks and three different feature vectors are output: a two-dimensional heat map , 2D joint coordinates , where and , and a latent feature map .

[0051] For step S03, in some embodiments, the deep object feature map is passed through a convolutional block and an activation function to obtain a two-dimensional heat map representation of the object, where the two-dimensional heat map representation of the object includes a height H and a width W; the mean value is calculated on the height H and the width W to quantify the two-dimensional heat map coordinates and convert them into corresponding one-dimensional vectors; let W and H dimensional tensors represent the heat map indices on the width W and the height H respectively; the normalized one-dimensional vector is multiplied by the corresponding heat map index to obtain the object 2D coordinates.

[0052] Specifically, for the deep feature map of the object , it is passed through a convolutional block stacked by multiple convolutions, and finally the corresponding activation function is used to obtain the two-dimensional heat map representation of the object . Among them , represents the minimum number of vertices of the object, B is the training batch, and H and W are the height and width of the input image respectively. Specifically, define the two-dimensional pixel point projection position coordinates of the object , and represent the horizontal value and the vertical value of the coordinates respectively. For , the mean value is calculated separately in its width and height dimensions to quantify the two-dimensional heat map coordinates and convert them into two independent one-dimensional vectors , . Let be a tensor of W and H dimensions, representing the heat map index. The normalized one-dimensional vector is multiplied by the heat map index to obtain the 2D coordinate representation of the object. As shown in the formula:

[0053]

[0054]

[0055]

[0056]

[0057]

[0058] Among them, represents element-wise multiplication, represents the 2D coordinates of an object.

[0059] For step S04, specifically, after obtaining the 2D coordinates of the hand and the object respectively, 2D pose alignment is applied to the hand-object latent feature maps at different scales as follows:

[0060]

[0061] Among them, represents the feature representation after alignment, represents the hand-object feature maps at different scales, represents the 2D coordinates of the hand or the object.

[0062] In step S05, preprocessing is performed on the hand pose alignment feature and the object pose alignment feature. The preprocessing includes: performing four-fold downsampling on the hand pose alignment feature and streamlining the number of hand vertices; streamlining the number of object vertices of the object pose alignment feature.

[0063] Specifically, as Figure 3 shown, a learnable parameter is introduced to achieve feature conversion from the 2D pose space to the 3D mesh vertex space. However, for the hand model, directly mapping to the complete set of mesh vertices (such as the 778 mesh vertices of the MANO hand model) often faces problems of large computational amount and high computational complexity. The present invention uses a downsampling strategy. Specifically, the hand model is downsampled four-fold, and the original 778 mesh vertices are streamlined to 49. For the object model, the minimum number of mesh vertices is represented as 21, and through experiments, the present invention sets the object mesh vertices to 63. This strategy not only simplifies the problem but also enables the design of a more compact and efficient mapping matrix M i , which maps the 2D features into the 3D space. It can be expressed as:

[0064]

[0065] Among them, representation j Grid vertex feature representation in three-dimensional space representation j Feature representation after alignment j representation of the hand h or object o , where is the number of key points, and the number of key points for both the hand and the object is 21.

[0066] For step S06, in some embodiments, a one-dimensional convolutional layer is used to extract the query values and feature values of the hand 3D space features and the object 3D space features respectively; the corresponding attention-weighted features are obtained according to the query values, feature values and the total feature dimension; a one-dimensional convolutional layer is adopted to fuse the hand 3D space features and the object 3D space features with the corresponding attention-weighted features respectively to obtain the hand enhanced features and the object enhanced features respectively.

[0067] Specifically, when the hand and the object interact, there is a correlation between them. For example, when picking up a pair of scissors or a plate, the hand postures are significantly different. However, previous work often models the interaction relationship between the hand and the object based on sparse predefined key points or joint points. The present invention uses a cross-attention mechanism to aggregate the hand-object fine-grained interaction features and enhance the feature expressions of the hand and the object. Specifically, as Figure 4 shown, the grid vertex features of the hand and the object in three-dimensional space and , hereinafter denoted by and , are used to extract their respective query values (query values), feature values (key values) and key-value pairs (value values) by using a one-dimensional convolutional layer. For the hand model, that is and , representing the query, key and value of specific grid vertices of the hand. At the same time, the query, key and value of the object model can also be calculated and represented as , and , where is the total feature dimension of the hand or the object. Aggregate all the object vertex features for each hand grid vertex, and the same applies to each object vertex. Calculate their respective attention weights using the formula and perform normalization processing, as shown in the following formula:

[0068]

[0069]

[0070] Here, the hand feature query and the object feature key The similarity between them, and the attention weight of each object feature to the hand feature is obtained through the normalized exponential function . Among them, represents the degree of attention of each element in the hand feature set to all elements in the object feature set. Similarly, the object feature query and the hand feature key The similarity between them is calculated, and the attention weight of the hand feature to the object feature is obtained . And represents the degree of attention of each element in the object feature set to all elements in the hand feature set.

[0071] Finally, the original features (i.e., hand 3D spatial features and object 3D spatial features) and the attention-weighted features are concatenated together, and a one-dimensional convolutional layer is used for the fusion operation, as shown in the following formula:

[0072]

[0073] Among them, represents the i feature in the final three-dimensional space, where i represents the hand feature or the object feature , represents the above-mentioned or , represents the above-mentioned or , represents the concatenation operation, represents the fusion operation. Using the three-dimensional interaction modeling DIM, the model's understanding and processing ability of the fine-grained interaction relationship between the two in the 3D space is enhanced.

[0074] Step S07: Using the contrastive learning of positive and negative samples, enhance the similarity between the hand vertices in the enhanced hand feature, and enhance the difference between the hand vertices and the environmental voxels in the enhanced hand feature, to obtain the optimized hand spatial feature.

[0075] It should be noted that in this embodiment, the process of step S07 is the process of the Vertex Contrastive Learning Network (VCL-Net).

[0076] In some embodiments, environmental voxels are used to form an embedded representation of vertex coordinates, where each environmental voxel represents the feature information of a local area in the hand enhancement features to capture the similarities and differences between data. For each hand vertex, the positive sample is the predicted hand vertex coordinate value, and the negative sample is the environmental voxel; based on the hand vertex and the environmental voxel, the hand InfoNCE loss function is used to obtain the optimized hand space features. The hand InfoNCE loss function includes a first contrast loss function between hand vertices and a second contrast loss function between hand vertices and environmental voxels.

[0077] Specifically, the goal of contrastive learning is usually to learn an effective data representation by comparing the similarities and differences between data samples, so that the representations of similar data in the feature space are closer, while the representations of different data in the feature space are more dispersed. The InfoNCE function is used to set up the contrastive learning process, as shown in the formula:

[0078]

[0079] where, represents the InfoNCE loss, represents the logarithmic function the input value of represents the exponential function, represents the anchor image, is the embedded space representation of represents the positive sample embedding, represents the negative embedding, is the negative sample set. All embeddings are normalized. represents the inner product, is the temperature hyperparameter used to control the smoothness of the distribution. The larger , the flatter the distribution output by the softmax function, and the lower the ability of the model to distinguish the similarities of different samples;

[0080] The numerator part queries the similarity between and the positive sample , and the temperature parameter is used to control the smoothness of the similarity distribution, where function converts the obtained similarity into a non-negative value and amplifies the similarity score of the positive sample. The denominator part represents the sum of the positive sample similarity score and all negative sample similarity scores. Overall, it makes the model pay more attention to the positive sample during training while minimizing the similarity of the negative sample.

[0081] The present invention believes that in a multi-object background, due to the severe occlusion between the hand and the object, using the 3D vertex annotation information of the hand may not provide sufficiently accurate supervision. Specifically, during feature learning, the area occupied by the hand mesh vertices is relatively small, while the vast majority are background elements, such as the objects being manipulated. These background elements occlude the hand, greatly affecting the learning and representation of hand features. To address this challenge, Figure 5 A grid vertex contrastive learning method is described, which provides global supervision for the 3D feature information of the hand. For the interaction process between the hand and the object, it makes the representations of the hand and the object in the feature space more dispersed, reducing the interference of background elements such as objects on the hand feature learning, so as to achieve accurate hand pose estimation.

[0082] First, effective key points are formed. Anchor point indices are randomly selected from a specific set based on a predefined number of anchor points, and then the consistency between the anchor point indices and the vertex mask is determined to obtain a label mask. According to the mask and the number of instances to be sampled, the index instances of the corresponding instances are obtained . Let the set represent all possible vertex indices, where is the batch size, is the number of grid vertices. A subset M is randomly selected from the set as the set of anchor point indices, where . For each anchor point index m ∈ M , there is a corresponding label . The instance index solution is shown in the equation:.

[0083]

[0084]

[0085] where, represents the set of positive instance indices, represents the set of negative instance indices. Finally, effective key points can be obtained through the instance indices in the 3D space.

[0086] Secondly, the optimization of hand gestures. There is a correlation between the same type of hand mesh vertices. Using mesh vertex contrast learning to strengthen the connection between them enables the model to learn better vertex feature representations, thereby improving the accuracy of vertex coordinate prediction. In addition, in the 3D feature space, the hand mesh vertices are also compared with other environmental voxels, and the feature representations are further strengthened by learning the differences. After obtaining the 3D hand feature space, environmental voxel embedding is used to form an embedding representation related to the vertex coordinates, where each environmental voxel embedding represents the feature information of a local area in the feature space, ensuring that the feature space can fully capture the similarities and differences between the data. For each hand mesh vertex, the positive sample is the predicted mesh vertex coordinate value, and the negative samples are some environmental voxels.

[0087] Based on the key points (i.e., the hand vertices of the present invention) and the 3D embedding representation (i.e., the environmental voxels of the present invention), the present invention performs optimization operations in the 3D hand space and uses the hand InfoNCE loss function to obtain more accurate hand features. Specifically, for the hand-environment contrast learning optimization, the optimal parameter settings are selected according to multiple experimental comparisons. The detailed definition of the second contrast loss function is shown in the formula:

[0088]

[0089] Where, represents the second contrast loss between the hand vertex and the environmental voxel, is the set of mesh vertex embeddings of the positive samples, represents the positive sample, represents the negative samples, i.e., the environmental voxels, C represents the input value of the logarithmic function of, represents the hand vertex coordinates.

[0090] For the hand-hand gesture contrast learning, the number of anchor points selected is the same as that of the hand-environment contrast learning. The specific definition of the first contrast loss function is shown in the formula:

[0091]

[0092] Where, represents the first contrast loss between the hand vertices, D represents the input value of the logarithmic function of.

[0093] Step S08: Decode the object enhanced features and the hand space optimized features to obtain the hand-object estimation result.

[0094] In this embodiment, it should be noted that the present invention conducts experiments on common benchmarks.

[0095] For the experimental dataset, HO3D v2 is a 3D hand and object dataset consisting of 77K images from 68 video sequences. It includes 10 objects and 10 subjects, with 66034 training samples and 11524 evaluation samples. This patent uses the official training set and test set splitting protocol and submits the test results to the official website for test evaluation. Dex-YCB is a recently introduced large-scale dataset in which 10 objects grasping 20 different objects are recorded from 8 different perspectives with 582K RGB-D frames over 1000 sequences. It can be applied to multiple tasks such as 2D hand and keypoint detection, object pose estimation, and 3D hand pose estimation. This patent mainly focuses on the 3D hand pose estimation problem and then gives the results according to the official s0 splitting protocol.

[0096] For the evaluation metrics of hand pose, PA-MPJPE / MPVPE measures the average of the Euclidean distances (mm) between the hand coordinates estimated by the model and the ground truth coordinates at each mesh vertex position, including both alignment and non-alignment analysis. F-Score is the harmonic mean of Precision and Recall. In 3D pose estimation tasks, F-Score can be used to evaluate the similarity between two meshes. At a given distance threshold , namely 5mm / 15mm, the distances between the vertices of the two meshes are compared. By calculating the F-Score at different thresholds, the performance of the algorithm under different precision requirements can be evaluated.

[0097] For the evaluation metrics of object pose, ADD-0.1D is a common evaluation metric in object pose estimation tasks, which is the average mesh vertex error within 10% of the object diameter. Here, ADD-0.1D combines the average mesh vertex error and a fixed ratio (10%) of the object diameter to quantify the similarity between the predicted pose and the ground truth pose of the object.

[0098] For the experimental details and parameter settings, the entire model (DIM and DIVCL-Net) of the present invention is optimized using the Adam optimizer with a weight decay of 5e-4 and a batch size of 64. During training, the model is trained for 70 epochs. The initial learning rate is 1e-4 and it decays every 10 epochs. For the hyperparameters, the input resolution size is set in this paper to be . All implementations are based on PyTorch.

[0099] Embodiment 2

[0100] Please refer to Figure 6 , which shows the structural schematic diagram of a hand-object estimation system proposed in the second embodiment of the present application. The system includes:

[0101] An image acquisition module 100 for acquiring hand-object interaction images;

[0102] A three-dimensional dense modeling module 200 for extracting features from the hand-object interaction image to obtain a hand deep feature map and an object deep feature map respectively; predicting hand joint points from the hand deep feature map to obtain 2D coordinates of hand vertices; performing object heatmap encoding on the object deep feature map to obtain 2D coordinates of the object; aligning the 2D coordinates of the hand vertices with the hand deep feature map to obtain hand pose alignment features; aligning the 2D coordinates of the object with the object deep feature map to obtain object pose alignment features; performing 3D feature mapping on the hand pose alignment features and the object pose alignment features to obtain hand 3D spatial features and object 3D spatial features respectively; capturing fine-grained interaction relationships between the hand 3D spatial features and the object 3D spatial features according to the cross-attention mechanism to obtain enhanced hand features and enhanced object features;

[0103] A vertex contrast learning module 300 for enhancing the similarity between hand vertices in the enhanced hand features and the difference between hand vertices in the enhanced hand features and environmental voxels through contrast learning of positive and negative samples to obtain optimized hand spatial features;

[0104] An output image module 400 for decoding the enhanced object features and the optimized hand spatial features to obtain hand-object estimation results.

[0105] A hand-object estimation system in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a network attached storage (NAS), a personal computer (PC), etc., which are not specifically limited in the embodiments of the present application.

[0106] A hand-object estimation system in an embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0107] The hand-object estimation system provided by the embodiments of the present application can implement Figure 1 each process implemented by a hand-object estimation method in the method embodiments. To avoid repetition, details are not described herein again.

[0108] Optionally, the embodiments of the present application further provide an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned hand-object estimation method embodiment and can achieve the same technical effect. To avoid repetition, details are not described herein again.

[0109] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-mentioned hand-object estimation method embodiment and can achieve the same technical effect. To avoid repetition, details are not described herein again.

[0110] Wherein, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc.

[0111] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0112] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0113] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A hand-object estimation method, characterized in that: include: Acquire hand-object interaction images; Performing feature extraction on the hand-object interaction image to obtain a hand deep feature map and an object deep feature map respectively; Predicting the hand joints on the hand deep feature map to obtain the 2D coordinates of the hand vertices; encoding the object heat map on the object deep feature map to obtain the 2D coordinates of the object; Performing posture alignment on the hand vertex 2D coordinates and the hand deep feature map to obtain a hand posture alignment feature; Performing posture alignment on the object 2D coordinates and the object deep feature map to obtain an object posture alignment feature; Performing 3D feature mapping on the hand posture alignment feature and the object posture alignment feature to obtain hand 3D spatial features and object 3D spatial features respectively; According to the cross attention mechanism, the fine-grained interaction relationship between the 3D spatial features of the hand and the 3D spatial features of the object is captured to obtain hand enhancement features and object enhancement features; By using contrastive learning of positive and negative samples, the similarity between hand vertices in the hand enhancement feature is enhanced, and the difference between the hand vertices in the hand enhancement feature and the environment voxels is enhanced, so as to obtain a hand space optimization feature; The object enhancement feature and the hand space optimization feature are decoded to obtain a hand-object estimation result.

2. A hand-object estimation method according to claim 1, characterized in that: The steps of extracting features from the hand-object interaction image to obtain a hand deep feature map and an object deep feature map respectively include: Using a Restnet50-FPN single- and double-branch hybrid network to extract shallow features of the hand-object interaction image, to obtain shallow features of the hand and shallow features of the object; The multi-scale backbone is used to perform multi-scale feature extraction on the shallow features of the hand and the shallow features of the object respectively, to obtain a deep feature map of the hand and a deep feature map of the object.

3. A hand-object estimation method according to claim 1, characterized in that: The step of predicting the hand joint points on the hand deep feature map to obtain the 2D coordinates of the hand vertices includes: The hand deep feature map is used to extract hand features through the hourglass module, residual learning is introduced through residual connection, and the hand features are processed through basic blocks to obtain three different feature vectors, which are a two-dimensional heat map, a 2D joint point coordinates, and a latent feature map.

4. The hand-object estimation method according to claim 1, characterized in that: The step of performing object heat map encoding on the object deep feature map to obtain the object 2D coordinates includes: The deep feature map of the object is passed through a convolution block and an activation function to obtain a two-dimensional heat map representation of the object, wherein the two-dimensional heat map representation of the object includes a dimension height H and a dimension width W; Calculate the average value on the dimension height H and the dimension width W to quantize the two-dimensional heat map coordinates and convert them into corresponding one-dimensional vectors respectively; Let W and H The tensors of dimensions represent the heat map indexes on the dimension width W and dimension height H respectively; The normalized one-dimensional vector is multiplied by the corresponding heat map index to obtain the object 2D coordinates.

5. The hand-object estimation method according to claim 1, characterized in that: Before the step of performing 3D feature mapping on the hand posture alignment feature and the object posture alignment feature, the method further includes: Preprocessing the hand posture alignment features and the object posture alignment features includes: Downsampling the hand pose alignment feature by a factor of four and reducing the number of hand vertices; The number of object vertices of the object pose alignment feature is reduced.

6. A hand-object estimation method according to claim 1, characterized in that: The steps of capturing the fine-grained interaction relationship between the hand 3D spatial features and the object 3D spatial features according to the cross attention mechanism to obtain the hand enhancement features and the object enhancement features include: Use a one-dimensional convolutional layer to extract the query value and feature value of the hand 3D spatial features and the object 3D spatial features respectively; Obtaining a corresponding attention weighted feature according to the query value, the feature value and the sum of the feature dimensions; A one-dimensional convolutional layer is used to fuse the hand 3D spatial features and the object 3D spatial features with the corresponding attention weighted features to obtain hand enhancement features and object enhancement features, respectively.

7. The hand-object estimation method according to claim 1, characterized in that: The steps of enhancing the similarity between hand vertices in the hand enhancement feature and enhancing the difference between the hand vertices and the environment voxels in the hand enhancement feature by using the contrast learning of positive and negative samples to obtain the hand space optimization feature include: Use environmental voxels to form an embedded representation of vertex coordinates, where each environmental voxel represents the feature information of a local area in the hand enhancement feature to capture the similarities and differences between the data. For each hand vertex, the positive sample is the predicted hand vertex coordinate value, and the negative sample is the environmental voxel. Based on the hand vertices and the environment voxels, a hand space optimization feature is obtained using a hand InfoNCE loss function, where the hand InfoNCE loss function includes a first contrast loss function between hand vertices and a second contrast loss function between hand vertices and environment voxels.

8. A hand-object estimation system, characterized in that: The system comprises: An image acquisition module is used to acquire a hand-object interaction image; A three-dimensional dense modeling module is used to extract features from the hand-object interaction image to obtain a hand deep feature map and an object deep feature map respectively; predict hand joint points on the hand deep feature map to obtain 2D coordinates of hand vertices; perform object heat map encoding on the object deep feature map to obtain object 2D coordinates; perform posture alignment on the hand vertex 2D coordinates and the hand deep feature map to obtain hand posture alignment features; perform posture alignment on the object 2D coordinates and the object deep feature map to obtain object posture alignment features; perform 3D feature mapping on the hand posture alignment features and the object posture alignment features to obtain hand 3D spatial features and object 3D spatial features respectively; and capture the fine-grained interaction relationship between the hand 3D spatial features and the object 3D spatial features according to the cross-attention mechanism to obtain hand enhancement features and object enhancement features; A vertex contrast learning module, used for enhancing the similarity between hand vertices in the hand enhancement feature and enhancing the difference between hand vertices and environmental voxels in the hand enhancement feature by contrast learning of positive and negative samples, so as to obtain a hand space optimization feature; The output image module is used to decode the object enhancement features and the hand space optimization features to obtain a hand-object estimation result.

9. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a hand-object estimation method as described in any one of claims 1 to 7 are implemented.

10. A readable storage medium, characterized in that: The readable storage medium stores programs or instructions, and when the programs or instructions are executed by the processor, the steps of a hand-object estimation method as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Three-dimensional hand-object interaction attitude estimation method and device and storage medium

    CN116580420A

  • Hand posture estimation and recognition method and device based on deep learning and medium

    CN116704554A