Attitude estimation methods, devices and electronic equipment

By extracting features from images and point cloud data, and using multilayer perceptrons and neural networks for feature fusion and interactive learning, the problem of inaccurate pose estimation when objects or people are occluded is solved, and accurate pose estimation is achieved under occlusion conditions.

CN115063887BActive Publication Date: 2025-10-31BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210707482.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-21
Publication Date
2025-10-31
Estimated Expiration
2042-06-21

AI Technical Summary

Technical Problem

In existing technologies, when an object or human body is occluded, the pose estimation is inaccurate, resulting in incomplete pose information.

Method used

By acquiring point cloud data of the image to be identified and the object in a standardized pose, the image features of the human body and the object are extracted. Multilayer perceptron and neural network are used for feature fusion and interactive learning to obtain the pose features of the human body and the object, thus compensating for the incomplete pose information caused by occlusion.

Benefits of technology

It enables accurate estimation of the pose of objects or people regardless of whether they are occluded, thus improving the accuracy of pose estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115063887B_ABST
    Figure CN115063887B_ABST
Patent Text Reader

Abstract

This invention provides a pose estimation method, apparatus, and electronic device. It extracts human image features and object image features from point cloud data of the image to be identified and the object's standardized pose. Pose information is then extracted based on these human and object image features to obtain human pose features and object pose features. These human and object pose features contain interaction information between the human body and the object. The human and object pose features are input into a multilayer perceptron to obtain predicted pose information. By utilizing the interaction information in the human and object pose features, the invention compensates for the incomplete pose information caused by occlusion of the object or human body, thereby estimating the pose of the object and human body. This achieves accurate pose estimation of objects and humans regardless of whether they are occluded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to an attitude estimation method, apparatus, and electronic device. Background Technology

[0002] The main methods for object pose estimation are based on observed object point cloud data. Algorithms calculate the pose by matching feature points with significant geometric features in the point cloud data, and by using neural networks to directly predict the object's pose based on the observed point cloud or image.

[0003] However, in existing technologies, changes in the pose of objects such as cushions, chairs, cabinet doors, or drawers are usually caused by human interaction. In such cases, people often severely occlude the object, resulting in incomplete point cloud data and inaccurate pose estimation.

[0004] Therefore, a pose estimation method is proposed to achieve accurate estimation of the pose of objects and people regardless of whether they are occluded. Summary of the Invention

[0005] This invention provides a posture estimation method to address the shortcomings of existing technologies where incomplete posture information is caused by occlusion of objects or human bodies, making it impossible to accurately estimate the posture of objects. This method achieves the effect of accurately estimating the posture of objects and human bodies regardless of whether they are occluded.

[0006] This invention provides an attitude estimation method, comprising:

[0007] Acquire point cloud data of the image to be identified and the object in the standard pose;

[0008] Human image features and object image features are extracted from the image to be identified;

[0009] Pose information is extracted based on the human image features and the object image features to obtain human pose features and object pose features, wherein the human pose features and the object pose features contain interaction information between the human body and the object.

[0010] The human posture features and the object posture features are input into a multilayer perceptron to obtain predicted posture information.

[0011] According to a pose estimation method provided by the present invention, human image features and object image features are extracted based on the point cloud data of the image to be identified and the object's canonical pose, including:

[0012] The image to be identified is input into the first neural network, which outputs human image features.

[0013] The object's standard pose point cloud data is input into the second neural network, which outputs the object's geometric features.

[0014] The image to be identified is input into a third neural network, which outputs the initial image features of the object.

[0015] The object image features are obtained by performing cross-attention feature fusion based on the object's geometric features and the object's initial image features.

[0016] According to a pose estimation method provided by the present invention, human image features and object image features are extracted from the image to be identified, including:

[0017] Pose information is extracted based on the human image features and the object image features to obtain human pose features and object pose features, including:

[0018] Based on the human image features and the object image features, feature interaction learning based on mutual attention is performed to obtain interactive object image features and interactive human image features.

[0019] The interactive human image features and the interactive object image features are subjected to self-attention-based feature reinforcement learning to obtain human pose features and object pose features.

[0020] According to a pose estimation method provided by the present invention, the human pose features and the object pose features are input into a multilayer perceptron to obtain predicted pose information, including:

[0021] The human posture features and the object posture features are respectively input into a multilayer perceptron, and the corresponding predicted human posture information and predicted object posture information are output.

[0022] The present invention also provides an attitude estimation device, comprising:

[0023] The acquisition unit is used to acquire point cloud data of the image to be identified and the object's standard pose.

[0024] The image feature unit is used to extract human image features and object image features based on the image to be identified and the object's standard pose point cloud data.

[0025] The pose feature unit is used to extract pose information based on the human image features and the object image features to obtain human pose features and object pose features, wherein the human pose features and the object pose features contain interaction information between the human body and the object.

[0026] The sensing unit is used to input the human posture features and the object posture features into the multilayer perceptron to obtain predicted posture information.

[0027] According to a pose estimation device provided by the present invention, the image feature unit is specifically used for:

[0028] The image to be identified is input into the first neural network, which outputs human image features.

[0029] The object's standard pose point cloud data is input into the second neural network, which outputs the object's geometric features.

[0030] The image to be identified is input into a third neural network, which outputs the initial image features of the object.

[0031] The object image features are obtained by performing cross-attention feature fusion based on the object's geometric features and the object's initial image features.

[0032] According to the attitude estimation device provided by the present invention, the attitude feature unit is specifically used for:

[0033] Based on the human image features and the object image features, feature interaction learning based on mutual attention is performed to obtain interactive object image features and interactive human image features.

[0034] The interactive human image features and the interactive object image features are subjected to self-attention-based feature reinforcement learning to obtain human pose features and object pose features.

[0035] According to the attitude estimation device provided by the present invention, the sensing unit is specifically used for:

[0036] The human posture features and the object posture features are respectively input into a multilayer perceptron, and the corresponding predicted human posture information and predicted object posture information are output.

[0037] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the above-described attitude estimation methods.

[0038] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described attitude estimation methods.

[0039] The pose estimation method, apparatus, and electronic device provided by this invention extract human image features and object image features from point cloud data of the image to be identified and the object's standardized pose. Pose information is then extracted based on these human and object image features to obtain human pose features and object pose features. These human and object pose features contain interaction information between the human body and the object. The human and object pose features are then input into a multilayer perceptron to obtain predicted pose information. By utilizing the interaction information contained in the human and object pose features, the invention compensates for the incomplete pose information caused by occlusion of the object or human body, thereby estimating the pose of the object and human body. This achieves accurate pose estimation of objects and humans regardless of whether they are occluded. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0041] Figure 1 This is a flowchart illustrating the attitude estimation method provided by the present invention;

[0042] Figure 2 This is a schematic diagram of the attitude estimation device provided by the present invention;

[0043] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0045] This invention provides an attitude estimation method, such as... Figure 1 As shown, it includes:

[0046] S11. Obtain point cloud data of the image to be identified and the standard pose of the object.

[0047] S12. Extract human image features and object image features based on the point cloud data of the image to be identified and the object's standard pose.

[0048] S13. Extract posture information based on the human image features and the object image features to obtain human posture features and object posture features, wherein the human posture features and the object posture features contain interaction information between the human body and the object.

[0049] S14. Input the human body posture features and the object posture features into a multilayer perceptron to obtain predicted posture information.

[0050] Specifically, image features can be extracted from the acquired image to be identified to obtain human image features. Object image features can be extracted from the image to be identified and the canonical pose point cloud data of the object. Human image features represent the image features of each part of the human body, while object image features represent the image features of each component of the object. It should be noted that the object described here and below can consist of one component or multiple components. The canonical pose point cloud data represents the point cloud of each component of the object in its canonical pose, i.e., its geometric shape. A canonical pose refers to the posture of an object in a specific state, such as a microwave oven with the door closed, or an office chair with the backrest upright and the seat adjusted to its highest position. Human image features can determine the positional information of each part of the human body in the image, and object image features can determine the positional information of each component of the object in the image.

[0051] Pose information is extracted from human and object image features. Human image features are transformed into information representing human pose, and object image features in the image to be identified are transformed into information representing object pose. The human and object pose features contain interaction information between the human body and the object.

[0052] Using a multilayer perceptron, posture prediction is performed based on the input human or object posture features to obtain predicted posture information.

[0053] In this embodiment of the invention, human image features and object image features are extracted from the point cloud data of the image to be identified and the object's standardized pose. Pose information is then extracted based on these human and object image features to obtain human pose features and object pose features. These human and object pose features contain interaction information between the human body and the object. The human and object pose features are then input into a multilayer perceptron to obtain predicted pose information. By utilizing the interaction information contained in the human and object pose features, the incomplete pose information caused by occlusion of the object or human body is compensated for, thereby estimating the pose of the object and human body. This achieves accurate pose estimation of objects and humans regardless of whether they are occluded.

[0054] According to the attitude estimation method provided by the present invention, step S12 includes:

[0055] S121. Input the image to be identified into the first neural network and output human image features.

[0056] S122. Input the object's standard pose point cloud data into the second neural network and output the object's geometric features;

[0057] S123. Input the image to be identified into the third neural network and output the initial image features of the object;

[0058] S124. Perform cross-attention feature fusion based on the geometric features of the object and the initial image features of the object to obtain the object image features.

[0059] Specifically, to ensure that the initial image features of the obtained object contain richer information, a first neural network, a second neural network, and a third neural network with different structures can be used. The first neural network, the second neural network, and the third neural network are all pre-trained neural networks. Pre-training is carried out using image samples to be identified. The training ends when the loss value of the neural network reaches a preset condition by calculating the loss value of the neural network.

[0060] For step S121, in one example, the first neural network can employ a PARE (Part Attention Regressor for 3D Human Body Estimation) neural network. The PARE neural network is a 3D human pose estimation neural network submitted to ICCV (International Conference on Computer Vision) in 2021, aiming to solve the occlusion problem in 3D human pose estimation. The image to be recognized is input into the PARE neural network, which extracts human image features from the image to be recognized, obtaining human image features that contain human pose information.

[0061] For step S122, in one example, the second neural network can be the PointNet neural network proposed in 2017. The PointNet network directly takes the image containing the point cloud as input and outputs the feature information of the entire input or the feature information of each point in the input. In a preferred example, the object's normalized pose point cloud data can be input into a PointNet neural network with shared weights. That is, the object's normalized pose point cloud data of object A is divided into point cloud data of a predetermined number of parts such as B and C when object A is in a normalized pose. The point cloud data of each part is sequentially input into the PointNet neural network with fixed weights to obtain the geometric features corresponding to the predetermined number of parts. The set of these geometric features is taken as the geometric features of object A, i.e., the object's geometric features.

[0062] For step S123, in one example, the third neural network can be a ResNet (Deep residual network) neural network. The image to be recognized is input into the ResNet neural network, and the ResNet neural network extracts the image features of the object to be recognized to obtain the initial image features of the object.

[0063] In step S124, both the object's geometric features and the initial image features are features extracted specifically for the object. To more comprehensively represent the object's features, a cross-attention feature fusion is performed on the object's geometric features and the initial image features to obtain the object's image features. In one example, the object's geometric features and the initial image features can be input into a cross-attention layer. The cross-attention layer projects the object's geometric features onto the initial image features, achieving cross-attention feature fusion and obtaining the object's image features. The parameters of the cross-attention layer can be set according to actual needs. The object's image features contain the object's pose information. However, if the object in the image to be identified is occluded by a human body, the obtained pose information is incomplete and cannot be used to predict the object's pose.

[0064] In this embodiment of the invention, a first neural network extracts human image features from the image to be recognized, obtaining human image features that include human posture information. A second neural network extracts geometric features from the object's standard posture point cloud data, obtaining object geometric features. A third neural network extracts object image features from the image to be recognized, obtaining object image features. Based on cross-attention feature fusion, the object geometric features and object image features are fused to obtain more comprehensive object features, i.e., object image features. This facilitates subsequent posture information extraction and prediction based on human image features and object image features to obtain more accurate predicted posture information.

[0065] According to the attitude estimation method provided by the present invention, step S13 includes:

[0066] S131. Perform feature interaction learning based on mutual attention according to the human image features and the object image features to obtain interactive object image features and interactive human image features.

[0067] Specifically, human image features and object image features are features that express some attributes and postures of the corresponding human body and object, and they do not have correlation or interactivity. In order to strengthen the interaction between the human body and the object, feature interaction learning based on mutual attention can be carried out based on human image features and object image features.

[0068] In one example, human image features and object image features can be input into a mutual attention layer. The mutual attention layer performs feature interaction learning based on mutual attention on the human image features and object image features, which strengthens the correlation and interaction between the human image features and object image features, resulting in interactive object image features and interactive human image features. At this point, the interactive object image features and interactive human image features have interactive information between the object and the human body. The parameters of the mutual attention layer can be set according to actual needs.

[0069] S132. Perform self-attention-based feature reinforcement learning on the interactive human image features and the interactive object image features to obtain human posture features and object posture features.

[0070] Specifically, after enhancing the correlation and interactivity of human image features and object image features, the resulting interactive human image features and interactive object image features contain interactive information between the human body and the object. However, in order to prevent the loss of information representing one's own posture due to excessive enhancement of correlation and interactivity, it is necessary to perform feature reinforcement learning based on self-attention on the interactive human image features and interactive object image features to mitigate the unexpected impact brought about by feature interaction learning based on mutual attention.

[0071] In one example, interactive human image features and interactive object image features can be input into a self-attention layer. The self-attention layer then performs feature reinforcement learning on the information representing the self-features in the interactive human and object image features to obtain human pose features and object pose features. At this point, the human pose features and object pose features are interactive while also preserving their own pose information. That is, the human pose features contain information about the interaction between the human and the object, as well as information about the human pose; similarly, the object pose features contain information about the interaction between the human and the object, as well as information about the object pose.

[0072] In this embodiment of the invention, interactive object image features and interactive human image features are obtained by performing mutual attention-based feature interaction learning based on human image features and object image features. This strengthens the interaction between the human body and the object, enriching them with relevant and interactive information. Simultaneously, to prevent over-enhancing relevance and interactivity from causing a loss of information representing the user's own posture, self-attention-based feature reinforcement learning is performed on the interactive human image features and interactive object image features to mitigate the unexpected impact of mutual attention-based feature interaction learning. The resulting human posture features and object posture features contain interactive information representing relevance and interactivity while also ensuring the completeness of their own posture information.

[0073] According to the attitude estimation method provided by the present invention, step S14 specifically includes:

[0074] The human posture features and the object posture features are respectively input into a multilayer perceptron, and the corresponding predicted human posture information and predicted object posture information are output.

[0075] Specifically, human pose features and object pose features can be input into a multilayer perceptron (MLP). The MLP then predicts the poses of the human body and objects from the human pose features and object pose features, obtaining predicted human pose information and predicted object pose information. In one example, human pose features and object pose features can be input into a MLP with fixed weights. The MLP predicts the human pose from the human pose features and outputs the predicted human pose information. The MLP predicts the object pose from the object pose features and outputs the predicted object pose information.

[0076] In this embodiment of the invention, a multilayer perceptron is used to predict the posture of the human body and the object, thereby obtaining the corresponding predicted posture information of the human body and the predicted posture information of the object. Since both the human body posture features and the object posture features contain the interaction information between the human body and the object, the posture of the object and the human body can be accurately estimated regardless of whether they are occluded.

[0077] The attitude estimation device provided by the present invention is described below. The attitude estimation device described below and the attitude estimation method described above can be referred to in correspondence.

[0078] The present invention also provides an attitude estimation device, such as Figure 2 As shown, it includes:

[0079] The acquisition unit 21 is used to acquire point cloud data of the image to be identified and the standard pose of the object.

[0080] Image feature unit 22 is used to extract human image features and object image features based on the image to be identified and the object's standard pose point cloud data.

[0081] The posture feature unit 23 is used to extract posture information based on the human image features and the object image features to obtain human posture features and object posture features, wherein the human posture features and the object posture features contain interaction information between the human body and the object.

[0082] The sensing unit 24 is used to input the human posture features and the object posture features into the multilayer perceptron to obtain predicted posture information.

[0083] In this embodiment of the invention, human image features and object image features are extracted from the point cloud data of the image to be identified and the object's standardized pose. Pose information is then extracted based on these human and object image features to obtain human pose features and object pose features. These human and object pose features contain interaction information between the human body and the object. The human and object pose features are then input into a multilayer perceptron to obtain predicted pose information. By utilizing the interaction information contained in the human and object pose features, the incomplete pose information caused by occlusion of the object or human body is compensated for, thereby estimating the pose of the object and human body. This achieves accurate pose estimation of objects and humans regardless of whether they are occluded.

[0084] According to a pose estimation device provided by the present invention, the image feature unit 22 is specifically used for:

[0085] The image to be identified is input into the first neural network, which outputs human image features.

[0086] The object's standard pose point cloud data is input into the second neural network, which outputs the object's geometric features.

[0087] The image to be identified is input into a third neural network, which outputs the initial image features of the object.

[0088] The object image features are obtained by performing cross-attention feature fusion based on the object's geometric features and the object's initial image features.

[0089] According to the attitude estimation device provided by the present invention, the attitude feature unit 23 is specifically used for:

[0090] Based on the human image features and the object image features, feature interaction learning based on mutual attention is performed to obtain interactive object image features and interactive human image features.

[0091] The interactive human image features and the interactive object image features are subjected to self-attention-based feature reinforcement learning to obtain human pose features and object pose features.

[0092] According to the attitude estimation device provided by the present invention, the sensing unit 24 is specifically used for:

[0093] The human posture features and the object posture features are respectively input into a multilayer perceptron, and the corresponding predicted human posture information and predicted object posture information are output.

[0094] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include a processor 310, a communication interface 320, a memory 330, and a communication bus 340, wherein the processor 310, the communication interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 can call logical instructions in the memory 330 to execute a pose estimation method, which includes: acquiring point cloud data of a to-be-identified image and an object's standardized pose; extracting human image features and object image features based on the point cloud data of the to-be-identified image and the object's standardized pose; extracting pose information based on the human image features and the object image features to obtain human pose features and object pose features, wherein the human pose features and the object pose features contain interaction information between the human body and the object; and inputting the human pose features and the object pose features into a multilayer perceptron to obtain predicted pose information.

[0095] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0096] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the pose estimation methods provided by the above methods. The method includes: acquiring point cloud data of an image to be identified and a canonical pose of an object; extracting human image features and object image features based on the point cloud data of the image to be identified and the canonical pose of the object; extracting pose information based on the human image features and the object image features to obtain human pose features and object pose features, wherein the human pose features and the object pose features contain interaction information between the human body and the object; and inputting the human pose features and the object pose features into a multilayer perceptron to obtain predicted pose information.

[0097] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the pose estimation method provided by the methods described above. The method includes: acquiring point cloud data of an image to be identified and a canonical pose of an object; extracting human image features and object image features based on the point cloud data of the image to be identified and the canonical pose of the object; extracting pose information based on the human image features and the object image features to obtain human pose features and object pose features, wherein the human pose features and the object pose features contain interaction information between the human body and the object; and inputting the human pose features and the object pose features into a multilayer perceptron to obtain predicted pose information.

[0098] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0099] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A pose estimation method, characterized in that, include: Acquire point cloud data of the image to be identified and the object in the standard pose; Human image features and object image features are extracted based on the point cloud data of the image to be identified and the object's standard pose. Pose information is extracted based on the human image features and the object image features to obtain human pose features and object pose features, wherein the human pose features and the object pose features contain interaction information between the human body and the object. The human posture features and the object posture features are input into a multilayer perceptron to obtain predicted posture information; Pose information is extracted based on the human image features and the object image features to obtain human pose features and object pose features, including: Based on the human image features and the object image features, feature interaction learning based on mutual attention is performed to obtain interactive object image features and interactive human image features. The interactive human image features and the interactive object image features are subjected to self-attention-based feature reinforcement learning to obtain human pose features and object pose features.

2. The attitude estimation method according to claim 1, characterized in that, Based on the point cloud data of the image to be identified and the object's standardized pose, human image features and object image features are extracted, including: The image to be identified is input into the first neural network, which outputs human image features. The object's standard pose point cloud data is input into the second neural network, which outputs the object's geometric features. The image to be identified is input into a third neural network, which outputs the initial image features of the object. The object image features are obtained by performing cross-attention feature fusion based on the object's geometric features and the object's initial image features.

3. The attitude estimation method according to claim 1, characterized in that, The human posture features and the object posture features are input into a multilayer perceptron to obtain predicted posture information, including: The human posture features and the object posture features are respectively input into a multilayer perceptron, and the corresponding predicted human posture information and predicted object posture information are output.

4. An attitude estimation device, characterized in that, include: The acquisition unit is used to acquire point cloud data of the image to be identified and the object's standard pose. The image feature unit is used to extract human image features and object image features based on the image to be identified and the object's standard pose point cloud data. The pose feature unit is used to extract pose information based on the human image features and the object image features to obtain human pose features and object pose features, wherein the human pose features and the object pose features contain interaction information between the human body and the object. A perception unit is used to input the human posture features and the object posture features into a multilayer perceptron to obtain predicted posture information; The posture feature unit is specifically used for: Based on the human image features and the object image features, feature interaction learning based on mutual attention is performed to obtain interactive object image features and interactive human image features. The interactive human image features and the interactive object image features are subjected to self-attention-based feature reinforcement learning to obtain human pose features and object pose features.

5. The attitude estimation device according to claim 4, characterized in that, The image feature unit is specifically used for: The image to be identified is input into the first neural network, which outputs human image features. The object's standard pose point cloud data is input into the second neural network, which outputs the object's geometric features. The image to be identified is input into a third neural network, which outputs the initial image features of the object. The object image features are obtained by performing cross-attention feature fusion based on the object's geometric features and the object's initial image features.

6. The attitude estimation device according to claim 4, characterized in that, The sensing unit is specifically used for: The human posture features and the object posture features are respectively input into a multilayer perceptron, and the corresponding predicted human posture information and predicted object posture information are output.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the attitude estimation method as described in any one of claims 1 to 3.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the attitude estimation method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Pose estimation method based on region-level feature fusion

    CN114155406A

  • Character interaction relation identification method and device, and electronic equipment

    CN114170688A