Hand-object interaction posture estimation method based on text prompt

Through the hand-object interaction posture estimation method based on text prompts, using the ResNet-50 network and multi-scale feature fusion technology, the problem of decreased posture estimation accuracy caused by hand-object interaction occlusion is solved, and hand-object posture estimation with higher accuracy and robustness is achieved.

CN120599665AActive Publication Date: 2025-09-05GUANGXI UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510770224.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-05
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

Existing technologies suffer from reduced accuracy in hand pose estimation due to occlusion problems, especially in hand-object interaction scenarios. The fusion of image features and text descriptions fails to effectively solve the loss of visual information caused by occlusion.

Method used

A text-cue-based hand-object interaction posture estimation method is adopted. Image features are extracted through the ResNet-50 network, and dynamic parameters are generated by combining global and local text encoders. Multi-scale feature fusion and supervision are performed, and hand-object posture prediction is performed using a multi-layer perceptron and a three-dimensional field learning module.

Benefits of technology

The accuracy and robustness of hand-object interaction posture estimation are improved, especially in complex occlusion scenarios, which can more accurately capture the posture details of hand joints and objects, and enhance the feature learning ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599665A_ABST
    Figure CN120599665A_ABST
Patent Text Reader

Abstract

The invention discloses a hand-object interaction posture estimation method based on text prompt, and relates to the technical field of computer vision and natural language processing. The method comprises the steps that a hand-object interaction posture prediction model is constructed, image data and text data are input, and the model is divided into a feature extraction layer, a feature fusion layer, a feature supervision layer and an output layer; the feature extraction layer inputs the 2D image into a ResNet-50 network to extract image features; in the feature fusion layer, global text prompts are encoded by a global text encoder and input into a multi-layer perceptron to be mapped into global dynamic parameters, and the global dynamic parameters are fused with image features to obtain global fusion features; local text prompts are encoded by a local text encoder and input into a multi-layer perceptron, the local text prompts are mapped into local dynamic parameters, and the local dynamic parameters are fused with the global fusion features to obtain harmonic features; the feature supervision layer constructs global supervision for global texts and image features, constructs joint supervision for joint names and corresponding image features, constructs pixel supervision for the joint names and corresponding heat maps in sequence, and optimizes an overall network framework; and the output layer outputs the attribute and position of each pixel in the image according to the matching degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of computer vision and natural language processing, and in particular to a method for estimating hand-object interaction posture based on text prompts. Background Art

[0002] Vision-based hand interaction can intuitively express human intentions in augmented reality operations, sign language conversations, human-computer interaction, and other areas, and has become an important means of human activity. However, vision-based hand interaction inevitably results in joint occlusions, including occlusion by other parts of the hand, occlusion by the other hand, and occlusion by objects. This loss of visual information due to occlusions can lead to a decrease in the accuracy of vision-based hand pose estimation and hand model reconstruction. Precise textual guidance can overcome the unreliability of purely visual cues in the presence of occlusions, explicitly encoding semantic priors about occluded regions and spatial interaction patterns, and providing flexible and complementary constraints to recover geometrically plausible poses.

[0003] In existing technologies, feature-level fusion is often used, such as simple splicing, addition, and multiplication between image features and text. However, this easily introduces redundant information and fails to effectively integrate the advantages of the two modalities. It does not consider the complexity of hand-object interaction and the problem that hand joints are often occluded. Moreover, the text description of the image often remains superficial, such as a simple description of the hand shape or the interaction between the hand and the object, which further leads to a significant decrease in the accuracy of hand-object estimation in the case of occlusion.

[0004] Therefore, there is an urgent need for an occlusion representation text prompt method for hand-object interaction posture estimation to improve the accuracy of hand-object estimation in occlusion situations. Summary of the Invention

[0005] Based on this, it is necessary to provide a hand-object interaction posture estimation method based on text prompts to address the above technical problems.

[0006] The present invention adopts the following technical solutions: The present invention provides a method for estimating hand-object interaction posture based on text prompts, comprising: Obtaining a 2D image of the hand-object interaction, a text prompt of the global hand-object interaction relationship, and a text prompt of local finger joint occlusion, and inputting these into a hand-object interaction posture prediction model; the hand-object interaction posture prediction model includes a feature extraction layer, a feature fusion layer, a feature supervision layer, and a feature output layer; In the feature extraction layer, the 2D image of the hand-object interaction is input into the trained ResNet-50 network to obtain image features; In the feature fusion layer, the global hand-object interaction relationship text prompt is input into the global text encoder for encoding, and the global text feature is obtained and input into the multi-layer perceptron, and mapped into the global text guided dynamic parameter; according to the global text guided dynamic parameter, the global text feature is fused with the image feature to obtain the global fusion feature; the text prompt of the local finger joint occlusion is input into the local text encoder for encoding to obtain the local text feature; the local text feature is input into the multi-layer perceptron and mapped into the local text guided dynamic parameter; the local text feature is fused with the global fusion feature through the local text guided dynamic parameter to obtain the multi-scale fusion feature; In the feature supervision layer, the global text features are fused with the image features and globally aligned with the multi-scale fusion features; the local text features are fused with the image features and locally aligned and pixel-aligned with the globally aligned multi-scale fusion features in turn to obtain aligned multi-scale fusion features; In the feature output layer, the aligned multi-scale fusion features are passed through the three-dimensional field learning module to learn the signed distance field information of the hand and object, and then through the field-guided posture regression module to obtain the final 3D posture parameters of the hand and object.

[0007] Preferably, the hand-object interaction 2D image is input into a trained ResNet-50 network to obtain image features, specifically including: Input the 2D image of hand-object interaction into the ResNet-50 network; The ResNet-50 network obtains the hand-object interaction 2D image features from the hand-object interaction 2D image; The 2D image features of hand-object interaction are input into the spatial attention mechanism and the channel attention mechanism, and the two are combined to obtain the image features.

[0008] Preferably, global alignment specifically includes: Fuse global text features with image features to obtain a global fusion feature map; The multi-scale fusion features are aligned with the global fusion feature map to obtain the globally aligned multi-scale fusion features.

[0009] Preferably, local alignment specifically includes: Process the image features through multi-head attention to obtain the processed image features; Fuse the local text features with the processed image features to obtain a local fusion feature map; The local fusion feature map is aligned with the globally aligned multi-scale fusion feature to obtain the locally aligned multi-scale fusion feature.

[0010] Preferably, pixel alignment specifically includes: Extract joint point heatmap from image features; Fuse the joint point heat map with the local text feature to obtain the fused joint text feature; Fuse the fused joint text features with the image features to obtain a joint fusion feature map; The locally aligned multi-scale fusion features are aligned with the locally aligned multi-scale fusion features to obtain the pixel-aligned multi-scale fusion features, that is, the aligned multi-scale fusion features.

[0011] Preferably, the weights in the local text editor remain unchanged during the training process of the hand-object interaction gesture prediction model.

[0012] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects: The present invention provides a method for estimating hand-object interaction posture based on text prompts, and constructs a hand-object interaction posture prediction model. A detailed text description is generated through a dual-scale structured text prompt generation module, and text semantic information is hierarchically combined with image features through a feature fusion module to enhance the features of the hand-object contact area, while suppressing background interference information, so that the model can initially focus on the area in the image that is consistent with the text description. Then, more specific occlusion text information is introduced, and by locally constraining and adjusting the fused features, it is ensured that the global semantic intention is first resolved in the fusion process. By adjusting the details through local constraints, the model can more accurately capture the occlusion of the hand joints and the hand-object interaction details. Through the feature supervision module, a multi-level supervision method is used to improve the feature learning ability of the model at different granularities, thereby improving the accuracy of hand-object posture estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0014] Figure 1 A flowchart of a method for estimating hand-object interaction posture based on text prompts provided by the present invention; Figure 2 A flowchart of text generation for a method of estimating hand-object interaction posture based on text prompts provided by the present invention; Figure 3 A schematic diagram of a posture prediction model for a hand-object interaction posture estimation method based on text prompts provided by the present invention; Figure 4 A schematic diagram of image feature extraction of a hand-object interaction posture estimation method based on text prompts provided by the present invention; Figure 5 A schematic diagram of feature fusion of a hand-object interaction posture estimation method based on text prompts provided by the present invention; Figure 6 This is a feature supervision module diagram of a hand-object interaction posture estimation method based on text prompts provided by the present invention. DETAILED DESCRIPTION

[0015] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of the present invention and corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in the specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0016] The technical solutions provided by the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0017] Figure 1 The figure is a flow chart of a method for estimating hand-object interaction posture based on text prompts in the present invention, which specifically includes the following steps: S101: Obtain a 2D image of the hand-object interaction, a text prompt of the global hand-object interaction relationship, and a text prompt of the local finger joint occlusion, and input them into a hand-object interaction posture prediction model; the hand-object interaction posture prediction model includes a feature extraction layer, a feature fusion layer, a feature supervision layer, and a feature output layer. For a schematic diagram of the model, see Figure 3 .

[0018] Specifically, the present invention first provides a text prompt template to generate global hand-object interaction text prompts and local finger joint occlusion text prompts. The prompts include the contact status between the hand and the object (whether the object is grasped or not) and the occlusion information of the finger joints. Figure 2 , which is a flowchart for text generation. In the figure, green represents self-occluded joints, red represents object-occluded joints, and light blue represents visible joints.

[0019] S102: In the feature extraction layer, see Figure 4 , is a schematic diagram of feature extraction, inputting the 2D image of hand-object interaction into the ResNet50 network to obtain image features, including: inputting the 2D image of hand-object interaction into the ResNet50 network; the ResNet50 network obtains the 2D image features of hand-object interaction; inputting the image features into the spatial attention mechanism and the channel attention mechanism to obtain enhanced image features.

[0020] Specifically, the feature extraction layer uses a network composed of ResNet50 to encode and decode the input 2D image.

[0021] S103: In the feature fusion layer, see Figure 5 , the global hand-object interaction relationship text prompt is input into the global text encoder for encoding, the global text feature is obtained and input into the multi-layer perceptron, mapped into the global text guided dynamic parameter; according to the global text guided dynamic parameter, the global text feature is fused with the image feature to obtain the global fusion feature; the text prompt occluded by the local finger joint is input into the local text encoder for encoding to obtain the local text feature; the local text feature is input into the multi-layer perceptron and mapped into the local text guided dynamic parameter; through the local text guided dynamic parameter, the local text feature is fused with the global fusion feature to obtain the multi-scale fusion feature.

[0022] Specifically, the weights within the text editor remain constant during training.

[0023] Specifically, the global text-guided dynamic parameter and the local text-guided dynamic parameter are used to adjust the scale and offset of the image feature and the global fusion feature, respectively.

[0024] Specifically, this phased fusion method avoids the conflict between high-level guidance and low-level structural adjustment, making the fusion process more consistent with the way human visual reasoning works. Specifically, in order to make full use of text information to guide model learning, the present invention uses two pre-trained visual language models, CLIP and Long-CLIP, as text encoders to encode hand-object interaction text and occluded text, respectively, to obtain corresponding text embedding features. In order to maintain language consistency, we freeze the weights of the two text encoders, which ensures the stability and consistency of text features and avoids semantic drift caused by weight updates during training. In order to make the extracted text features better adaptable to subsequent fusion tasks, we introduce a lightweight multi-layer perceptron to map text embedding features into dynamic parameters. α n and β n These two parameters will be used to modulate image features to achieve the fusion of text features and image features. Specifically, α n and β n They are used to adjust the scale and offset of image features, so as to dynamically integrate text semantic information into image features. α n and β n Modulate the image features and further optimize the fusion results so that the model can more accurately capture the occlusion of the hand joints and the details of the hand-object interaction. The specific process is as follows: Figure 5The figure below shows a schematic diagram of image-text feature fusion. By enhancing the features of the hand-object contact area while suppressing background interference, the model can initially focus on areas in the image that are consistent with the text description. The purpose of this step is to utilize global semantic information to provide a foundation for subsequent feature refinement.

[0025] S104: In the feature supervision layer, see Figure 6 , which is a schematic diagram of feature supervision, the global text feature is fused with the image feature and globally aligned with the multi-scale fusion feature; the local text feature is fused with the image feature, and locally aligned and pixel-aligned with the globally aligned multi-scale fusion feature in turn to obtain the aligned multi-scale fusion feature.

[0026] Optionally, the global alignment specifically includes: fusing global text features with image features to obtain a global fusion feature map; aligning multi-scale fusion features with the global fusion feature map to obtain globally aligned multi-scale fusion features.

[0027] Specifically, the global text features are used to supervise the multi-scale fusion features. First, the multi-scale visual features are extracted from the multi-scale fusion features and fused with the text features of the hand-object interaction. Then, the normalized fusion features are combined with the new image features to calculate F sc The similarity between the hand and the object is calculated and the model parameters are adjusted so that the image features can better reflect the semantic information in the text description. This alignment strategy ensures that the model can accurately capture the macro-interaction relationship between the hand and the object when processing the entire image, laying the foundation for subsequent detailed analysis.

[0028] Optionally, local alignment specifically includes: processing image features through multi-head attention to obtain processed image features; fusing local text features with processed image features to obtain a local fusion feature map; aligning the local fusion feature map with the globally aligned multi-scale fusion features to obtain locally aligned multi-scale fusion features.

[0029] Specifically, at the local joint level, we focus on feature learning of hand joints. By aligning the hand joint text with the features of the corresponding joints of the multi-scale fusion features, the model can more accurately identify and locate the hand joints. Specifically, we generate corresponding text features for each hand joint, and extract the local features of the joint area from the image features and fuse them into the image features. Then, we calculate the similarity between the fused image features and the multi-scale fusion features, and optimize them using the contrast loss function. This enables the model to deeply understand the local features of the hand joints and accurately capture subtle changes in hand posture even in complex hand-object interaction scenarios.

[0030] Optionally, pixel alignment specifically includes: extracting a joint point heat map from image features; fusing the joint point heat map with local text features to obtain a fused joint text feature; fusing the fused joint text feature with the image feature to obtain a joint fusion feature map; aligning the locally aligned multi-scale fusion feature with the locally aligned multi-scale fusion feature to obtain a pixel-aligned multi-scale fusion feature, that is, an aligned multi-scale fusion feature.

[0031] Specifically, at the pixel level, we use 2D label heatmaps to further refine the model’s ability to locate image features. Each hand joint has a corresponding heatmap ( Figure 6 The joint position on the right is highlighted). By learning these heat maps, the model can accurately know the specific position of each joint in the image. In the specific operation, we combine the heat map information with the image features and use the multi-head self-attention mechanism to enhance the expressiveness of the features. Then, the matching degree of the multi-scale fusion features and the joint text features in the image features is calculated in the embedding space, and compared with the target heat map to optimize the model parameters. In this way, the model can also accurately predict the attributes and position of each pixel at the pixel level, thereby achieving high-precision estimation of the hand-object posture. The specific process is as follows: Figure 6 The following is a schematic diagram of the text supervision module.

[0032] S105: In the feature output layer, the aligned multi-scale fusion features are passed through the three-dimensional field learning module to learn the signed distance field information of the hand and the object, and then passed through the field-guided posture regression module to obtain the final 3D posture parameters of the hand and the object.

[0033] Optionally, in the feature output layer, the attributes and position of each pixel in the image are output according to the matching degree between the enhanced image features and the local alignment features.

[0034] In this paper, a new image-text fusion method is first proposed, which successfully integrates text semantic information with image features in a deep and precise manner. By introducing a dynamic parameter modulation mechanism, the model not only enhances its focus on key areas of hand-object interaction in the image, but also effectively suppresses interference from irrelevant information such as background. In the task of hand-object pose estimation, this method enables the model to more accurately capture the occlusion of hand joints and the pose details of the object, thereby significantly improving the accuracy and robustness of hand-object pose estimation, especially in complex scenes and under occlusion.

[0035] Furthermore, the present invention uses global hand-object interaction text, local joint label text, and 2D heatmaps to supervise the fused image features, achieving multi-level feature learning from the overall to the details, from semantic understanding to precise positioning. This not only significantly improves the accuracy and robustness of hand-object pose estimation, but also more accurately captures the occlusion of hand joints and the pose details of objects in complex scenes and occlusions, thereby achieving higher-quality hand-object pose estimation.

[0036] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present invention.

Claims

1. A method for estimating hand-object interaction posture based on text prompts, characterized in that: include: Obtaining a 2D image of the hand-object interaction, a text prompt of the global hand-object interaction relationship, and a text prompt of local finger joint occlusion, and inputting these into a hand-object interaction posture prediction model; the hand-object interaction posture prediction model includes a feature extraction layer, a feature fusion layer, a feature supervision layer, and a feature output layer; In the feature extraction layer, the 2D image of the hand-object interaction is input into the trained ResNet-50 network to obtain image features; In the feature fusion layer, the global hand-object interaction relationship text prompt is input into the global text encoder for encoding, and the global text feature is obtained and input into the multi-layer perceptron, and mapped into the global text guided dynamic parameter; according to the global text guided dynamic parameter, the global text feature is fused with the image feature to obtain the global fusion feature; the text prompt of the local finger joint occlusion is input into the local text encoder for encoding to obtain the local text feature; the local text feature is input into the multi-layer perceptron and mapped into the local text guided dynamic parameter; the local text feature is fused with the global fusion feature through the local text guided dynamic parameter to obtain the multi-scale fusion feature; In the feature supervision layer, the global text feature is fused with the image feature and globally aligned with the multi-scale fusion feature; the local text feature is fused with the image feature and locally aligned and pixel-aligned with the globally aligned multi-scale fusion feature in turn to obtain the aligned multi-scale fusion feature; In the feature output layer, the aligned multi-scale fusion features are passed through the three-dimensional field learning module to learn the signed distance field information of the hand and object, and then through the field-guided posture regression module to obtain the final 3D posture parameters of the hand and object.

2. The method for estimating hand-object interaction posture based on text prompts according to claim 1, characterized in that: The hand-object interaction 2D image is input into the trained ResNet-50 network to obtain image features, specifically including: Input the 2D image of hand-object interaction into the ResNet-50 network; The ResNet-50 network obtains the hand-object interaction 2D image features from the hand-object interaction 2D image; The 2D image features of hand-object interaction are input into the spatial attention mechanism and the channel attention mechanism, and the two are combined to obtain the image features.

3. The method for estimating hand-object interaction posture based on text prompts according to claim 1, characterized in that: The global alignment specifically includes: Fuse global text features with image features to obtain a global fusion feature map; The multi-scale fusion features are aligned with the global fusion feature map to obtain the globally aligned multi-scale fusion features.

4. The method for estimating hand-object interaction posture based on text prompts according to claim 1, wherein: The local alignment specifically includes: Process the image features through multi-head attention to obtain the processed image features; Fuse the local text features with the processed image features to obtain a local fusion feature map; The local fusion feature map is aligned with the globally aligned multi-scale fusion feature to obtain the locally aligned multi-scale fusion feature.

5. The method for estimating hand-object interaction posture based on text prompts according to claim 1, characterized in that: The pixel alignment specifically includes: Extract joint point heatmap from image features; Fuse the joint point heat map with the local text feature to obtain the fused joint text feature; Fuse the fused joint text features with the image features to obtain a joint fusion feature map; The locally aligned multi-scale fusion features are aligned with the locally aligned multi-scale fusion features to obtain the pixel-aligned multi-scale fusion features, that is, the aligned multi-scale fusion features.

6. The method for estimating hand-object interaction posture based on text prompts according to claim 1, characterized in that: The weights in the local text editor remain unchanged during the training process of the hand-object interaction posture prediction model.

Citation Information

Patent Citations

  • Text vehicle re-identification method based on multi-scale and multi-view feature alignment

    CN116704452A

  • Text-guided multi-person attitude estimation method based on step-by-step joint processing

    CN117727071A

  • Method for predicting whole-body posture key points of human body based on vision-language level alignment relationship

    CN118397704A

  • Attitude estimation method and system based on deep learning

    CN119006598A

  • Image-text retrieval method based on comparative learning and modal fusion

    CN119441512A