Parameterized three-dimensional human body posture estimation method, device, equipment and medium

Through the parameterized three-dimensional human posture estimation method, the cross-attention module of topological information enhancement and Transformer structure is used to optimize the human node characteristics, solve the positioning errors when obstructing the posture, and realize the accurate estimation of the three-dimensional posture.

CN120544259APending Publication Date: 2025-08-26SHANGHAI INST OF MICROSYSTEM & INFORMATION TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510521456.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, when facing large-area human body occlusion, it is difficult to accurately estimate the three-dimensional human body posture, especially when the occlusion posture or unpredictable posture, there is a problem of positioning errors.

Method used

The parameterized three-dimensional human posture estimation method is adopted, and the image encoder, topology enhancement decoder and human model parameter regression module are used to optimize the human node characteristics and accurately estimate the shape and posture parameters through the image encoder, topology enhancement module and the cross attention module of the Transformer structure.

Benefits of technology

It enhances the ability to capture joint topology information, improves the expression and estimation accuracy of three-dimensional pose distribution, and solves the problem of positioning errors in occlusion postures and poses without poses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544259A_ABST
    Figure CN120544259A_ABST
Patent Text Reader

Abstract

The invention relates to a parameterized three-dimensional human body posture estimation method and device, equipment and a medium, and the method comprises the steps: obtaining a human body posture image; inputting the human body posture image into a human body posture estimation model to obtain a shape parameter and a posture parameter of a human body posture; the human body posture estimation model comprises an image encoder module which is used for carrying out feature encoding on a human body posture image to obtain a feature map after global context feature enhancement; the topology enhancement decoder module is used for constraining the optimization direction of human body node features and optimizing the human body node features based on the feature graph after global context feature enhancement; and the human body model parameter regression module is used for obtaining shape parameters and posture parameters of human body postures according to output regression of the topology enhancement decoder module. The human body model parameters can be accurately estimated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human body posture estimation, and in particular to a parameterized three-dimensional human body posture estimation method, device, equipment and medium. Background Art

[0002] Given the need for human visual perception in biomimetic robots, perceiving the human body's position or posture is a fundamental task, fundamental to robots' understanding and perception of human behavior. With the rapid development of deep learning, significant progress has been made in estimating the 3D pose of the human body. To address issues such as pose localization errors, many methods employ various model structures (such as transformers or CLIP) to extract effective features and achieve more accurate 3D poses. However, these methods rely on the design of complex network models and lack the ability to understand the pose itself.

[0003] Reconstructing the correct 3D human pose from a 2D image when large areas of the body are occluded is a highly uncertain problem, resulting in multiple plausible 3D poses matching the same image. Current methods addressing large-area occlusion often focus on finding a single optimal solution for a given input by optimizing the model structure, overlooking the inverse problem with multiple feasible solutions. In the field of human pose estimation, some recent methods have advocated generating multiple predictions for each input. However, these methods typically rely on aggregated predictions, resulting in a limited number of outputs and failing to fully represent the 3D pose distribution. Previously, some parameterized methods used convolutional neural meshes to estimate the shape parameters β and pose parameters θ of a parametric human model, generating 3D coordinate representations of human joints and a complete human mesh model. These parameterized methods, by modeling the uncertainty of the pose parameters θ, can fully represent the 3D pose distribution. However, this process often involves highly nonlinear mappings, making it difficult for convolutional neural meshes to accurately estimate the shape parameters β and pose parameters θ, thus impacting pose estimation performance. These methods still suffer from localization errors when faced with occluded poses or poses that have never been seen before. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a parameterized three-dimensional human body posture estimation method, device, equipment and medium, which can accurately estimate human body model parameters.

[0005] The present invention solves the technical problem by providing a parameterized three-dimensional human body posture estimation method, comprising the following steps:

[0006] Acquire human body posture images;

[0007] Inputting the human body posture image into a human body posture estimation model to obtain shape parameters and posture parameters of the human body posture;

[0008] Wherein, the human body posture estimation model includes:

[0009] An image encoder module is used to perform feature encoding on the human posture image to obtain a feature map after global context feature enhancement;

[0010] A topology enhancement decoder module, configured to constrain the optimization direction of human node features and optimize human node features based on the feature map enhanced with the global context features;

[0011] The human body model parameter regression module is used to regress the shape parameters and posture parameters of the human body posture according to the output of the topology enhancement decoder module.

[0012] The image encoder module includes:

[0013] The CNN backbone network part is used to extract features from the human posture image to obtain a feature map;

[0014] A sequence generation part is used to perform dimensionality reduction processing on the feature map and add position coding features to generate a feature sequence;

[0015] The Transformer encoding part is used to encode the feature sequence to obtain a feature map after global context feature enhancement.

[0016] The Transformer encoding part consists of several layers of standard Transformer encoder layers.

[0017] The topology enhancement decoder module includes:

[0018] The topological information enhancement part is used to obtain the human body node features based on the human body kinematic topological structure;

[0019] A cross self-attention part is used to extract relevant semantic information from the feature map enhanced with the global context feature using the human node feature as a query sequence, thereby optimizing the human node feature;

[0020] The feedforward neural network part is used to perform feedforward neural network calculation on the optimized human body node features and output them to the human body model parameter regression module.

[0021] The topology information enhancement part includes:

[0022] The joint GCN unit is used to take a pose query vector as input and use a graph convolutional neural network to adaptively model the complex semantic relationships between neighboring nodes to constrain the optimization direction of human node features, thereby obtaining a pose query vector that captures the kinematic topology between joints.

[0023] The feature interaction unit is used to interact the information between the kinematic topological posture query vector and the shape query vector with captured joints to obtain human body node features.

[0024] The joint GCN unit is represented as: P′ Q =P Q +σ(W1P Q +W2P Q A), P′ Q is a pose query vector that captures the kinematic topology between joints, P Q is the posture query vector, σ() is the activation function, W1, W2 are trainable transformation weights, and A is the adjacency matrix, which is used to inject the human body kinematic topology into the graph convolutional neural network; the feature interaction unit is represented as: H PS =Attention([P′ Q ,S Q ]), H PS is the human node feature, S Q is the shape query vector, Attention() represents the self-attention calculation, and [] represents the connection operation.

[0025] The human body model parameter regression module includes:

[0026] A shape parameter regression head, consisting of three linear transformation layers, is used to regress the shape parameters of the human body posture based on the output of the topology enhancement decoder module;

[0027] The posture parameter regression head part consists of three semantic graph convolution layers and a 1×1 convolution projection layer, which is used to regress the posture parameters of the human body according to the output of the topology enhancement decoder module.

[0028] The technical solution adopted by the present invention to solve the technical problem is to provide a parameterized three-dimensional human posture estimation device, comprising:

[0029] An acquisition module, used to acquire human posture images;

[0030] an estimation module, configured to input the human body posture image into a human body posture estimation model to obtain shape parameters and posture parameters of the human body posture;

[0031] Wherein, the human body posture estimation model includes:

[0032] An image encoder module is used to perform feature encoding on the human posture image to obtain a feature map after global context feature enhancement;

[0033] A topology enhancement decoder module, configured to constrain the optimization direction of human node features and optimize human node features based on the feature map enhanced with the global context features;

[0034] The human body model parameter regression module is used to regress the output of the topology enhancement decoder module to obtain the shape parameters and posture parameters of the human body posture.

[0035] The technical solution adopted by the present invention to solve its technical problem is: to provide an electronic device, including a memory, a processor and a computer program stored in the memory and capable of running on the processor, and when the processor executes the computer program, the steps of the above-mentioned parameterized three-dimensional human body posture estimation method are implemented.

[0036] The technical solution adopted by the present invention to solve its technical problem is: providing a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned parameterized three-dimensional human body posture estimation method are implemented.

[0037] Beneficial effects

[0038] Due to the adoption of the above-mentioned technical solution, the present invention has the following advantages and positive effects compared with the prior art: the present invention enhances the ability of posture to capture joint topology information by using a topological information enhancement module, thereby constraining the optimization direction of the joint points and enhancing the feature expression of the joint points. At the same time, the present invention utilizes the powerful cross-attention module in the Transformer structure to extract relevant semantic information from the image coding features, further optimize the feature expression of the nodes, enhance the ability to accurately estimate the human body model parameters, and finally accurately estimate the human body model parameters through the human body model parameter regression part, so as to fully and accurately express the three-dimensional posture distribution. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flow chart of a parameterized three-dimensional human body pose estimation method according to a first embodiment of the present invention;

[0040] Figure 2 is an architectural diagram of a human posture estimation model in a first embodiment of the present invention;

[0041] Figure 3 It is a schematic structural diagram of the topology information enhancement part in the first embodiment of the present invention. DETAILED DESCRIPTION

[0042] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.

[0043] The first embodiment of the present invention relates to a parameterized 3D human body pose estimation method, such as Figure 1 As shown, the following steps are included:

[0044] Step 1, obtaining a human body posture image;

[0045] Step 2: inputting the human body posture image into a human body posture estimation model to obtain shape parameters and posture parameters of the human body posture.

[0046] like Figure 2 As shown, the human body posture estimation model in this embodiment includes:

[0047] An image encoder module is used to perform feature encoding on the human posture image to obtain a feature map after global context feature enhancement;

[0048] A topology enhancement decoder module, configured to constrain the optimization direction of human node features and optimize human node features based on the feature map enhanced with the global context features;

[0049] The human body model parameter regression module is used to regress the shape parameters and posture parameters of the human body posture according to the output of the topology enhancement decoder module.

[0050] The image encoder module in this embodiment includes a CNN backbone network part, a sequence generation part, and a Transformer encoding part. The CNN backbone network part is used to extract features from the human posture image to obtain a feature map F∈R H×W×C , where H, W, and C represent the height, width, and number of channels, respectively. The sequence generation part is used to perform dimensionality reduction processing on the feature map F using a 1×1 convolution layer, and add position encoding features to generate a feature sequence X that meets the input of the Transformer encoding part. The Transformer encoding part is used to encode the feature sequence to obtain the feature map F after global context feature enhancement. c In this embodiment, the Transformer encoding part is composed of several layers of standard Transformer encoder layers.

[0051] The topology enhancement decoder module in this embodiment includes a topology information enhancement part, a cross self-attention part and a feedforward neural network part.

[0052] Among them, the topological information enhancement part is used to obtain human node features based on the human body dynamics topological structure, such as Figure 3 As shown, it includes a joint GCN unit and a feature interaction unit. The joint GCN unit is used to take the posture query vector as input, and use the graph convolutional neural network to adaptively model the complex semantic relationship between neighboring nodes to constrain the optimization direction of the human body node features, and obtain a kinematic topological posture query vector that captures the kinematic topology between joints; the feature interaction unit is composed of a self-attention layer, which is used to interact the information between the kinematic topological posture query vector and the shape query vector that captures the kinematic topology between joints to obtain the human body node features. Specifically, the posture query vector P Q First, it is enhanced by the joint GCN unit, and then through the residual connection to obtain the kinematic topological pose query vector P that captures the joints. Q ′, the joint GCN unit can be expressed as:

[0053] P′ Q =P Q +σ(W1P Q +W2P Q A);

[0054] Among them, W1, W2 are trainable transformation weights, σ() is the activation function, and A is the adjacency matrix, through which the human body kinematic topology can be injected into the graph convolutional neural network. Then the pose query vector P with the kinematic topology between the joints is captured. Q ′ and shape query vector S Q are connected together and fed into the feature interaction unit, which captures the pose query vector P with the kinematic topology between joints through the self-attention mechanism. Q ′ and shape query vector S Q Interact and output human node features H PS , the feature interaction unit can be expressed as:

[0055] H PS =Attention([P′ Q ,S Q ]);

[0056] Among them, [] represents the connection operation, and Attention() represents the self-attention calculation.

[0057] The cross self-attention part in this embodiment is used to transform the human node feature H PS The feature map F after the query sequence is enhanced from the global context features c Extract relevant semantic information and optimize the human node features. The cross self-attention part can be expressed as:

[0058]

[0059] in, For the optimized human node features, the corresponding positions of Q, K, and V in Attention(Q, K, V) represent the query vector, value vector, and key vector.

[0060] The feedforward neural network part in this embodiment is used to perform feedforward neural network calculation on the optimized human node features and output them to the human body model parameter regression module, which can be expressed as:

[0061]

[0062] Among them, H′ PS is the output of the feedforward neural network part, and FFN() represents the feedforward neural network calculation.

[0063] The human body model parameter regression module in this embodiment includes a shape parameter regression head part and a posture parameter regression head part. The shape parameter regression head part consists of three layers of linear transformation layers, defined as a function SRH(), which is used to regress the shape parameter β of the human body posture according to the output of the topology enhancement decoder module, which can be expressed as: β = SRH(H′ PS ); The posture parameter regression head consists of three semantic graph convolution layers and a 1×1 convolution projection layer, which is defined as a function PRH(), which is used to regress the posture parameter θ of the human posture according to the output of the topology enhancement decoder module, which can be expressed as: θ = PRH(H′ PS By leveraging the advantages of the semantic graph convolutional layer, the pose parameter regression head can guide the pose to a more accurate space constrained by the human body topology.

[0064] Finally, the human body parametric model SMPL is driven by the shape parameters β and posture parameters θ of the human body posture to generate the three-dimensional coordinate representation of the human body joints and the complete human body mesh model.

[0065] It is not difficult to find that the present invention enhances the ability of posture to capture joint topology information by using a topological information enhancement module, thereby constraining the optimization direction of the joint points and enhancing the feature expression of the joint points. At the same time, the powerful cross-attention module in the Transformer structure is used to extract relevant semantic information from the image coding features, further optimize the feature expression of the nodes, and enhance the ability to accurately estimate the parameters of the human body model. Finally, the human body model parameter regression module is used to accurately estimate the parameters of the human body model, namely the shape parameter β and posture parameter θ of the human body posture. Compared with existing 3D human body posture estimation technology, the method provided by the present invention can accurately estimate the parameters of the human body parameter model while fully expressing the three-dimensional posture distribution. It provides an effective solution to the problem of positioning errors when facing occluded postures or postures that have not appeared before.

[0066] A second embodiment of the present invention relates to a parameterized 3D human body pose estimation device, comprising:

[0067] An acquisition module, used to acquire human posture images;

[0068] an estimation module, configured to input the human body posture image into a human body posture estimation model to obtain shape parameters and posture parameters of the human body posture;

[0069] Wherein, the human body posture estimation model includes:

[0070] An image encoder module is used to perform feature encoding on the human posture image to obtain a feature map after global context feature enhancement;

[0071] A topology enhancement decoder module, configured to constrain the optimization direction of human node features and optimize human node features based on the feature map enhanced with the global context features;

[0072] The human body model parameter regression module is used to regress the output of the topology enhancement decoder module to obtain the shape parameters and posture parameters of the human body posture.

[0073] The image encoder module includes:

[0074] The CNN backbone network part is used to extract features from the human posture image to obtain a feature map;

[0075] A sequence generation part is used to perform dimensionality reduction processing on the feature map and add position coding features to generate a feature sequence;

[0076] The Transformer encoding part is used to encode the feature sequence to obtain a feature map after global context feature enhancement.

[0077] The Transformer encoding part consists of several layers of standard Transformer encoder layers.

[0078] The topology enhancement decoder module includes:

[0079] The topological information enhancement part is used to obtain the human body node features based on the human body kinematic topological structure;

[0080] A cross self-attention part is used to extract relevant semantic information from the feature map enhanced with the global context feature using the human node feature as a query sequence, thereby optimizing the human node feature;

[0081] The feedforward neural network part is used to perform feedforward neural network calculation on the optimized human body node features and output them to the human body model parameter regression module.

[0082] The topology information enhancement part includes:

[0083] The joint GCN unit is used to take a pose query vector as input and use a graph convolutional neural network to adaptively model the complex semantic relationships between neighboring nodes to constrain the optimization direction of human node features, thereby obtaining a pose query vector that captures the kinematic topology between joints.

[0084] The feature interaction unit is used to interact the information between the kinematic topological posture query vector and the shape query vector with captured joints to obtain human body node features.

[0085] The joint GCN unit is represented as: P′ Q =P Q +σ(W1P Q +W2P Q A), P′ Q is a pose query vector that captures the kinematic topology between joints, P Q is the posture query vector, σ() is the activation function, W1, W2 are trainable transformation weights, and A is the adjacency matrix, which is used to inject the human body kinematic topology into the graph convolutional neural network; the feature interaction unit is represented as: H PS =Attention([P′ Q ,S Q ]), H PS is the human node feature, S Q is the shape query vector, Attention() represents the self-attention calculation, and [] represents the connection operation.

[0086] The human body model parameter regression module includes:

[0087] A shape parameter regression head, consisting of three linear transformation layers, is used to regress the shape parameters of the human body posture based on the output of the topology enhancement decoder module;

[0088] The posture parameter regression head part consists of three semantic graph convolution layers and a 1×1 convolution projection layer, which is used to regress the posture parameters of the human body according to the output of the topology enhancement decoder module.

[0089] A third embodiment of the present invention relates to an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the parameterized three-dimensional human body posture estimation method of the first embodiment are implemented.

[0090] A fourth embodiment of the present invention relates to a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the parameterized three-dimensional human pose estimation method of the first embodiment are implemented.

[0091] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) that contain computer-usable program code.

[0092] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0093] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction method, which is implemented in the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0095] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A parameterized three-dimensional human body pose estimation method, characterized in that: The following steps are involved: Acquire human body posture images; Inputting the human body posture image into a human body posture estimation model to obtain shape parameters and posture parameters of the human body posture; Wherein, the human body posture estimation model includes: An image encoder module is used to perform feature encoding on the human posture image to obtain a feature map after global context feature enhancement; A topology enhancement decoder module, configured to constrain the optimization direction of human node features and optimize human node features based on the feature map enhanced with the global context features; The human body model parameter regression module is used to regress the output of the topology enhancement decoder module to obtain the shape parameters and posture parameters of the human body posture.

2. The parameterized 3D human pose estimation method according to claim 1, wherein: The image encoder module includes: The CNN backbone network part is used to extract features from the human posture image to obtain a feature map; The sequence generation part is used to perform dimensionality reduction processing on the feature map and add position coding features to generate a feature sequence; the Transformer encoding part is used to encode the feature sequence to obtain a feature map after global context feature enhancement.

3. The parameterized 3D human pose estimation method according to claim 2, wherein: The Transformer encoding part consists of several layers of standard Transformer encoder layers.

4. The parameterized 3D human pose estimation method according to claim 1, wherein: The topology enhancement decoder module includes: The topological information enhancement part is used to obtain the human body node features based on the human body kinematic topological structure; A cross self-attention part is used to extract relevant semantic information from the feature map enhanced with the global context feature using the human node feature as a query sequence, thereby optimizing the human node feature; The feedforward neural network part is used to perform feedforward neural network calculation on the optimized human body node features and output them to the human body model parameter regression module.

5. The parameterized 3D human body pose estimation method according to claim 4, wherein: The topology information enhancement part includes: The joint GCN unit is used to take a pose query vector as input and use a graph convolutional neural network to adaptively model the complex semantic relationships between neighboring nodes to constrain the optimization direction of human node features, thereby obtaining a pose query vector that captures the kinematic topology between joints. The feature interaction unit is used to interact the information between the kinematic topological posture query vector and the shape query vector with captured joints to obtain human body node features.

6. The parameterized 3D human body pose estimation method according to claim 5, characterized in that: The joint GCN unit is represented as: P′ Q =P Q +σ(W1P Q +W2P Q A), P′ Q is a pose query vector that captures the kinematic topology between joints, P Q is the posture query vector, σ() is the activation function, W1, W2 are trainable transformation weights, and A is the adjacency matrix, which is used to inject the human body kinematic topology into the graph convolutional neural network; the feature interaction unit is represented as: H PS =Attention([P′ Q ,S Q ]), H PS is the human node feature, S Q is the shape query vector, Attention() represents the self-attention calculation, and [] represents the connection operation.

7. The parameterized 3D human pose estimation method according to claim 1, wherein: The human body model parameter regression module includes: A shape parameter regression head, consisting of three linear transformation layers, is used to regress the shape parameters of the human body posture based on the output of the topology enhancement decoder module; The posture parameter regression head part consists of three semantic graph convolution layers and a 1×1 convolution projection layer, which is used to regress the posture parameters of the human body according to the output of the topology enhancement decoder module.

8. A parameterized 3D human body posture estimation device, characterized in that: include: An acquisition module, used to acquire human posture images; an estimation module, configured to input the human body posture image into a human body posture estimation model to obtain shape parameters and posture parameters of the human body posture; Wherein, the human body posture estimation model includes: An image encoder module is used to perform feature encoding on the human posture image to obtain a feature map after global context feature enhancement; A topology enhancement decoder module, configured to constrain the optimization direction of human node features and optimize human node features based on the feature map enhanced with the global context features; The human body model parameter regression module is used to regress the shape parameters and posture parameters of the human body posture according to the output of the topology enhancement decoder module.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the parameterized three-dimensional human body pose estimation method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the parameterized three-dimensional human pose estimation method according to any one of claims 1 to 7 are implemented.