Deeply enhanced three-dimensional human pose estimation method, device, equipment and medium
By extracting depth features from images and 2D human pose, and using an adaptive feature enhancement network and a self-attention mechanism for feature fusion, the depth ambiguity problem in 3D human pose estimation in existing technologies is solved, and the accuracy of 3D human pose estimation is improved.
Patent Information
- Application Number
- CN202510586856.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-05-08
AI Technical Summary
Existing technologies suffer from depth ambiguity when mapping two-dimensional human poses to three-dimensional human coordinates, which affects the accuracy of three-dimensional human pose estimation.
By extracting depth features, 2D human pose, and planar features from images, an adaptive feature enhancement network is used to regress the uncertainty of human joints, perform feature alignment and adaptive sampling, and combine a self-attention mechanism for feature fusion to estimate 3D human pose.
It effectively mitigates the effects of depth blur, improves the accuracy of 3D human pose estimation, ensures the rational use of depth information, and avoids noise interference.
Smart Images

Figure CN120108044B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to a depth-enhanced three-dimensional human pose estimation method, apparatus, device, and medium. Background Technology
[0002] 3D human pose estimation not only supports complex application scenarios but also demonstrates enormous potential in multiple fields. For example, in areas such as intelligent surveillance, virtual reality, augmented reality, and healthcare, accurate human pose estimation can help improve the intelligence level of the system. 3D human pose estimation aims to infer the positions of various joints of the human body in three-dimensional space from images or videos. The current mainstream implementation method is a two-stage lifting approach. In the first stage, an existing 2D human pose estimator is used to estimate the 2D human coordinates. In the second stage, the 2D human coordinates identified in the first stage are used as input to lift the 2D human coordinates to 3D human coordinates. However, when mapping 2D human coordinates to 3D human coordinates, there is an inherent depth ambiguity problem, which affects the accuracy of 3D human pose estimation.
[0003] Therefore, existing technologies still need to be improved and enhanced. Summary of the Invention
[0004] The technical problem to be solved by this application is to provide a depth-enhanced three-dimensional human pose estimation method, device, equipment and medium to address the shortcomings of the existing technology.
[0005] To address the aforementioned technical problems, the first aspect of this application provides a depth-enhanced three-dimensional human pose estimation method, wherein the depth-enhanced three-dimensional human pose estimation method specifically includes:
[0006] Extract depth features, 2D human pose, and planar features from the image to be estimated;
[0007] The uncertainty of human joints is regressed based on the depth features, and the uncertainty of human joints is applied to the depth features to obtain enhanced depth features.
[0008] The enhanced depth features, 2D human pose, and planar features are fused to obtain fused features;
[0009] The three-dimensional human pose corresponding to the image to be estimated is estimated based on the fused features.
[0010] The depth-enhanced 3D human pose estimation method, wherein the extraction of depth features, 2D human pose, and planar features from the image to be estimated specifically includes:
[0011] Input the image to be estimated into the depth estimator and the 2D pose estimator, respectively;
[0012] The depth estimator extracts the depth features of the image to be estimated, and the two-dimensional human pose and planar features of the image to be estimated are extracted by the two-dimensional pose estimator.
[0013] The depth-enhanced 3D human pose estimation method, after extracting the depth features, 2D human pose, and planar features of the image to be estimated, further includes:
[0014] An alignment operation is performed on the depth features, the two-dimensional human pose, and the planar features, wherein the alignment operation involves sampling the depth features and the planar features using the two-dimensional human pose.
[0015] The depth-enhanced 3D human pose estimation method, wherein after aligning the depth features, the 2D human pose, and the planar features, the method further includes:
[0016] Determine the depth offset and depth weight based on the aligned depth features, and then adaptively sample the aligned depth features based on the depth offset and depth weight.
[0017] The depth-enhanced 3D human pose estimation method, wherein after aligning the depth features, the 2D human pose, and the planar features, the method further includes:
[0018] The plane offset and plane weight are determined based on the aligned plane features, and the aligned plane features are adaptively sampled based on the plane offset and plane weight.
[0019] The depth-enhanced 3D human pose estimation method, wherein the uncertainty in regressing human joints based on the depth features specifically includes:
[0020] The depth features are input into an adaptive feature enhancement network, which outputs the uncertainty of the human joints corresponding to the image to be estimated and the depth value of the human joints. The adaptive feature enhancement network is described as a Bayesian neural network.
[0021] The adaptive feature enhancement network applies the uncertainty of the human joints to the depth features to obtain enhanced depth features.
[0022] The depth-enhanced 3D human pose estimation method, wherein fusing the enhanced depth features, 2D human pose, and planar features to obtain fused features specifically includes:
[0023] The enhanced depth features, 2D human pose, and planar features are stitched together to obtain the stitched features;
[0024] Based on the splicing features, the fusion features are determined through a self-attention mechanism.
[0025] A second aspect of this application provides a depth-enhanced three-dimensional human pose estimation device, wherein the depth-enhanced three-dimensional human pose estimation device specifically includes:
[0026] The feature extraction module is used to extract depth features, two-dimensional human pose, and planar features from the image to be estimated.
[0027] The feature enhancement module is used to regress the uncertainty of human joints based on the depth features, and to obtain enhanced depth features by applying the uncertainty of human joints to the depth features.
[0028] The feature fusion module is used to fuse the enhanced depth features, two-dimensional human pose, and planar features to obtain fused features;
[0029] The human pose estimation module is used to estimate the three-dimensional human pose corresponding to the image to be estimated based on the fused features.
[0030] A third aspect of this application provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps in the depth-enhanced 3D human pose estimation method as described above.
[0031] A fourth aspect of this application provides a terminal device, which includes: a processor and a memory;
[0032] The memory stores a computer-readable program that can be executed by the processor;
[0033] When the processor executes the computer-readable program, it implements the steps in any of the depth-enhanced 3D human pose estimation methods described above.
[0034] Beneficial Effects: Compared with existing technologies, this application provides a depth-enhanced 3D human pose estimation method, apparatus, device, and medium. The method includes extracting depth features, 2D human pose, and planar features from the image to be estimated; regressing the uncertainty of human joints based on the depth features, and applying the uncertainty of human joints to the depth features to obtain enhanced depth features; fusing the enhanced depth features, 2D human pose, and planar features to obtain fused features; and estimating the 3D human pose of the image to be estimated based on the fused features. This application first captures the depth features in human pose by accurately extracting and explicitly representing the depth features, and then adaptively enhances the depth features using uncertainty. This ensures that effective depth information is used reasonably and avoids noise in the depth features from interfering with the pose estimation task, effectively mitigating the impact of depth blur and improving the accuracy of 3D human pose estimation. Attached Figure Description
[0035] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 A flowchart of a depth-enhanced 3D human pose estimation method provided in an embodiment of this application.
[0037] Figure 2 This is an example diagram of a three-dimensional human body pose.
[0038] Figure 3 A schematic diagram of the principle of the depth-enhanced three-dimensional human pose estimation device provided in the embodiments of this application.
[0039] Figure 4 A schematic block diagram of the terminal device provided in the embodiments of this application. Detailed Implementation
[0040] This application provides a depth-enhanced three-dimensional human pose estimation method, apparatus, device, and medium. To make the objectives, technical solutions, and effects of this application clearer and more explicit, the following detailed description is provided with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining this application and are not intended to limit this application.
[0041] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0042] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0043] It should be understood that the sequence number and size of each step in this embodiment do not imply the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of this application embodiment.
[0044] The application content will be further explained below with reference to the accompanying drawings and the description of the embodiments.
[0045] This embodiment provides a depth-enhanced 3D human pose estimation method. This method utilizes a trained 3D human pose estimation model. Specifically, the image to be estimated is input into the 3D human pose estimation model, and the 3D human pose is input through the model. The 3D human pose estimation model can be a deep learning model and specifically includes a feature extraction module, a feature enhancement module, a feature fusion module, and a human pose estimation module. The feature extraction module extracts depth features, 2D human pose, and planar features from the image to be estimated. The feature enhancement module regresses the uncertainty of human joints based on the depth features and applies this uncertainty to the depth features to obtain enhanced depth features. The feature fusion module fuses the enhanced depth features, 2D human pose, and planar features to obtain fused features. The human pose estimation module estimates the 3D human pose corresponding to the image to be estimated based on the fused features.
[0046] To illustrate in detail the depth-enhanced 3D human pose estimation method provided in the embodiments of this application, the specific implementation process of the depth-enhanced 3D human pose estimation method is described below in the form of flowchart steps. The specific process of each module in the 3D human pose estimation model implementing its corresponding function can be referred to the specific description of its corresponding step in the flowchart steps, so as to illustrate the specific implementation process of the 3D human pose estimation model in realizing 3D human pose estimation.
[0047] like Figure 1 As shown, the depth-enhanced 3D human pose estimation method provided in this application specifically includes:
[0048] S10. Extract depth features, two-dimensional human pose, and planar features from the image to be estimated.
[0049] Specifically, the image to be estimated can be an RGB image acquired by an image acquisition device, or an RGB image transmitted from an external device, for example, a monocular RGB image. Depth features reflect the depth information in the image to be estimated that is key to the human body, and are used to assist in the mapping from 2D human pose to 3D human pose, mitigating depth ambiguity by explicitly providing depth information as a reference. Planar features reflect the planar information in the image to be estimated that is key to the human body, and are used to assist in providing background context information. 2D human pose reflects the 2D human pose information extracted from the image to be estimated. The depth features, 2D human pose, and planar features can be obtained by a feature extraction module that extracts features from the image to be estimated. This feature extraction module can include a depth estimator and a 2D pose estimator, extracting depth features through the depth estimator and extracting 2D human pose and planar features through the 2D pose estimator.
[0050] Based on this, the extraction of depth features, two-dimensional human pose, and planar features from the image to be estimated specifically includes:
[0051] Input the image to be estimated into the depth estimator and the 2D pose estimator, respectively;
[0052] The depth estimator extracts the depth features of the image to be estimated, and the two-dimensional human pose and planar features of the image to be estimated are extracted by the two-dimensional pose estimator.
[0053] Specifically, a depth estimator is used to extract depth features from the image to be estimated, and a 2D pose estimator is used to estimate the 2D pose of the image to be estimated, thereby extracting the 2D human pose and planar features of the image. In other words, the feature extraction model in the 3D human pose estimation model can include a depth estimator and a 2D pose estimator. The image to be estimated is input to both the depth estimator and the 2D pose estimator in the feature extraction model. The depth estimator outputs depth features, and the 2D pose estimator outputs the 2D human pose and planar features. For example, for a monocular RGB image... , Indicates the image height. The value represents the image width, and 3 represents the number of channels in the image. Pixel-level depth features are extracted using a depth estimator. Two-dimensional human pose is extracted using a two-dimensional pose estimator. and planar features, among which, 2 represents the number of joints in the human body, 2 represents the two-dimensional coordinates of the joints, and C represents the feature dimension.
[0054] Furthermore, after obtaining the depth features, 2D human pose, and planar features, to avoid unnecessary computational overhead from global processing of 3D human pose based on these features, an alignment operation can be performed on the depth features, 2D human pose, and planar features. This alignment operation involves sampling the depth features and planar features using the 2D human pose, and can be expressed as:
[0055] ,
[0056] ,
[0057] in, This represents the aligned planar features. This represents the aligned depth features. Indicates a sampling operation. Representing planar features, Representing depth features, Represents the posture of a two-dimensional human body.
[0058] This application embodiment samples depth and planar features by two-dimensional human pose, which can extract rich depth and planar information related to human joints from the depth and planar features, and filter out background information unrelated to human joints contained in the depth and planar features, thereby avoiding interference from background noise in the image to be estimated.
[0059] Furthermore, after aligning the depth features, 2D human pose, and planar features, the aligned depth features, 2D human pose, and planar features can be directly used to estimate the 3D human pose. However, in real-world scenarios, the 2D pose estimator used to extract the 2D human pose inevitably introduces errors, causing the extracted joint coordinates in the 2D human pose to be non-true values, resulting in the loss of information about the true joints in the features. To address this, in this embodiment, after aligning the depth features, 2D human pose, and planar features, adaptive sampling can be performed on the aligned depth features and / or aligned planar features. Furthermore, the execution process of the adaptive sampling operation for depth features and planar features can be the same: the adaptive sampling operation can first determine the offset and weight based on the original features, then sample the original features based on the offset, and finally weight the sampled features with the original features based on the weight. The difference between the two lies in the original features. When performing adaptive sampling on depth features, the original features are depth features, while when performing adaptive sampling on planar features, the original features are planar features. That is, the execution process of the adaptive sampling operation for depth features is to determine the depth offset and depth weight based on the aligned depth features, and then perform adaptive sampling on the aligned depth features based on the depth offset and depth weight. The execution process of the adaptive sampling operation for planar features is to determine the planar offset and planar weight based on the aligned planar features, and then perform adaptive sampling on the aligned planar features based on the planar offset and planar weight.
[0060] Based on this, taking deep features as an example, the execution process of the adaptive sampling operation can be represented as follows:
[0061] ,
[0062] ,
[0063] ,
[0064] in, Indicates depth weights, Represents a linear mapping. Indicates the depth offset. This represents the depth features after adaptive sampling.
[0065] In addition, it should be noted that, firstly This is determined by the softmax normalization performed after the linear mapping. This is obtained by performing the tanh activation function after linear mapping, and is used here for... This is used to represent the process. Secondly, when adaptively sampling depth features and / or planar features, it can be done once or multiple times. Here, we will use one adaptive sampling as an example for explanation. In practical applications, the number of adaptive samplings can be determined according to actual needs, and no specific limit is imposed here.
[0066] In this embodiment, after aligning the depth features, 2D human pose, and planar features, adaptive sampling is performed on the aligned depth features and / or planar features. This results in the adaptively sampled depth features and / or planar features containing more complete feature information related to human joints, avoiding the loss of information about real human joints in the depth features and / or planar features caused by sampling using 2D human pose. This improves the accuracy of the 3D human pose determined subsequently based on depth features, 2D human pose, and planar features.
[0067] S20. Based on the depth features, regress the uncertainty of human joints, and apply the uncertainty of human joints to the depth features to obtain enhanced depth features.
[0068] Specifically, the uncertainty of human joints is used to reflect the reliability of depth features. This uncertainty includes the uncertainty of each human joint in the image to be estimated; that is, the regression of human joint uncertainty based on depth features results in an uncertainty sequence, which includes the uncertainty of each human joint in the image to be estimated. The uncertainty of human joints can be determined by regressing depth features, i.e., using depth features as input to a regression head and outputting the uncertainty of human joints. When alignment and adaptive sampling operations are performed on the depth features, the depth features input to the regression head are the adaptively sampled depth features; when only alignment is performed, the depth features input to the regression head are the aligned depth features.
[0069] In one implementation, the feature enhancement module in the 3D human pose estimation model can employ an adaptive feature enhancement network, which outputs enhanced depth features. Accordingly, the step of regressing the uncertainty of human joints based on the depth features, and applying the uncertainty of human joints to the depth features to obtain enhanced depth features, specifically includes:
[0070] The depth features are input into an adaptive feature enhancement network, and the adaptive feature enhancement network outputs the uncertainty of the human joint points corresponding to the image to be estimated and the depth values of the human joint points.
[0071] The adaptive feature enhancement network applies the uncertainty of the human joints to the depth features to obtain enhanced depth features.
[0072] Specifically, the adaptive feature enhancement network is used to determine the uncertainty of human joints corresponding to the image to be estimated. At the same time, the adaptive feature enhancement network can also output the depth value of the human joint and apply the uncertainty of the human joint to the depth feature to obtain enhanced depth feature, and control the participation of the depth feature in feature fusion. The depth value of the human joint is a depth value sequence, which includes the depth value of each human joint in the image to be estimated.
[0073] Taking the depth features as adaptively sampled depth features as an example, the process of determining the enhanced depth features can be represented as follows:
[0074] ,
[0075] ,
[0076] ,
[0077] in, This represents the depth value of a human joint. This indicates the uncertainty of the human body's joint points. This represents a multilayer perceptron. This indicates enhanced deep features.
[0078] Furthermore, this application requires identifying which parts of the depth features have higher noise levels, necessitating the prediction of uncertainty for each human joint. To this end, the adaptive feature enhancement network is formulated as a Bayesian neural network. The Bayesian neural network obtains the Gaussian distribution of the depth features to derive the depth value (i.e., the mean of the Gaussian distribution) and the uncertainty (i.e., the variance of the Gaussian distribution). This allows the adaptive feature enhancement network to estimate the Gaussian distribution of the human body, not just predict the absolute coordinates of the human joints. During the training of the adaptive feature enhancement network, the uncertainty of the human joints is used to represent numerical stability, and a Bayesian loss constraint is employed (reducing uncertainty S while making μ closer to the true depth value) to make the depth distribution of the joints closer to the true value. The Bayesian loss constraint can be expressed as:
[0079] ,
[0080] in, This represents the loss function used during the training of the adaptive feature enhancement network. This indicates the number of training samples included in the training batch. Indicates the first The annotation depth value of each training sample. Indicates the first The predicted depth value for each training sample. Indicates the first The prediction uncertainty of each training sample.
[0081] S30. The enhanced depth features, two-dimensional human pose, and planar features are fused to obtain fused features.
[0082] Specifically, the fused features are obtained through multimodal information interaction among enhanced depth features, 2D human pose, and planar features. This multimodal information interaction can be achieved through an attention mechanism. In other words, the feature fusion module in the aforementioned 3D human pose estimation model can be configured with an attention mechanism. The fused features are obtained by inputting the enhanced depth features, 2D human pose, and planar features into the feature fusion module, and the feature fusion module outputs the fused features.
[0083] For example, the feature fusion of the enhanced depth features, 2D human pose, and planar features to obtain fused features specifically includes:
[0084] The enhanced depth features, 2D human pose, and planar features are stitched together to obtain the stitched features;
[0085] Based on the splicing features, the fusion features are determined through a self-attention mechanism.
[0086] Specifically, concatenating the enhanced depth features, 2D human pose, and planar features refers to merging the tokens of these features into a single transformer encoder. This encoder simulates the interaction between different features to achieve the interactive information between the enhanced depth features, 2D human pose, and planar features. The process of determining the fused features can be represented as follows:
[0087] ,
[0088] ,
[0089] in, Indicates splicing characteristics, Indicates fusion characteristics, This indicates a splicing operation. This indicates that the focus is on the bulls.
[0090] S40. Estimate the three-dimensional human pose corresponding to the image to be estimated based on the fused features.
[0091] Specifically, after obtaining the fused features, the 3D human pose can be predicted using these features, for example, such as... Figure 2The three-dimensional human pose is shown. The three-dimensional human pose can be estimated using a multilayer perceptron; that is, the human pose estimation module in the above three-dimensional human pose estimation model can use a multilayer perceptron. The process of estimating the three-dimensional human pose based on fused features by the human pose estimation module can be represented as follows:
[0092] ,
[0093] in, Represents the three-dimensional human posture.
[0094] In summary, this embodiment provides a depth-enhanced 3D human pose estimation method. The method includes extracting depth features, 2D human pose, and planar features from the image to be estimated; regressing the uncertainty of human joints based on the depth features, and applying the uncertainty of human joints to the depth features to obtain enhanced depth features; fusing the enhanced depth features, 2D human pose, and planar features to obtain fused features; and estimating the 3D human pose of the image to be estimated based on the fused features. This application first captures depth features in human pose by accurately extracting and explicitly representing them, and then adaptively enhances the depth features using uncertainty. This ensures that effective depth information is used appropriately and avoids noise in the depth features from interfering with the pose estimation task, effectively mitigating the impact of depth blur and improving the accuracy of 3D human pose estimation.
[0095] Based on the aforementioned depth-enhanced 3D human pose estimation method, this embodiment provides a depth-enhanced 3D human pose estimation device, such as... Figure 3 As shown, the depth-enhanced 3D human pose estimation device specifically includes:
[0096] The feature extraction module 100 is used to extract depth features, two-dimensional human pose, and planar features of the image to be estimated.
[0097] The feature enhancement module 200 is used to regress the uncertainty of human joints based on the depth features, and to obtain enhanced depth features by applying the uncertainty of human joints to the depth features.
[0098] The feature fusion module 300 is used to fuse the enhanced depth features, two-dimensional human pose, and planar features to obtain fused features;
[0099] The human pose estimation module 400 is used to estimate the three-dimensional human pose corresponding to the image to be estimated based on the fused features.
[0100] Based on the depth-enhanced 3D human pose estimation method described above, this embodiment provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps in the depth-enhanced 3D human pose estimation method as described in the above embodiment.
[0101] Based on the aforementioned depth-enhanced 3D human pose estimation method, this application also provides a terminal device, such as... Figure 4 As shown, it includes at least one processor 20; a display screen 21; and a memory 22, and may also include a communications interface 23 and a bus 24. The processor 20, display screen 21, memory 22, and communications interface 23 can communicate with each other via the bus 24. The display screen 21 is configured to display a preset user guide interface in the initial setup mode. The communications interface 23 can transmit information. The processor 20 can invoke logical instructions in the memory 22 to execute the methods described in the above embodiments.
[0102] Furthermore, the logical instructions in the aforementioned memory 22 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.
[0103] The memory 22, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, such as program instructions or modules corresponding to the methods in the embodiments of this disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, thereby implementing the methods in the above embodiments.
[0104] The memory 22 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 22 may include high-speed random access memory (RAM) and non-volatile memory. Examples include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, as well as transient storage media.
[0105] Furthermore, the specific process of loading and executing multiple instruction processors in the aforementioned storage medium and terminal device has been described in detail in the above method, and will not be repeated here.
[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A depth-enhanced 3D human pose estimation method, characterized in that, The depth-enhanced 3D human pose estimation method specifically includes: Extract depth features, two-dimensional human pose, and planar features from the image to be estimated, and perform an alignment operation on the depth features, the two-dimensional human pose, and the planar features, wherein the alignment operation is to sample the depth features and the planar features using the two-dimensional human pose; The uncertainty of human joints is regressed based on the depth features, and the uncertainty of human joints is applied to the depth features to obtain enhanced depth features, wherein the uncertainty of human joints is used to control the participation of depth features in feature fusion. The enhanced depth features, 2D human pose, and planar features are fused to obtain fused features; Estimate the 3D human pose corresponding to the image to be estimated based on the fused features; Specifically, the uncertainty in regressing human joints based on the depth features includes: The depth features are input into an adaptive feature enhancement network, and the adaptive feature enhancement network outputs the uncertainty of the human joints corresponding to the image to be estimated and the depth value of the human joints. The adaptive feature enhancement network is described as a Bayesian neural network. The adaptive feature enhancement network applies the uncertainty of the human joints to the depth features to obtain enhanced depth features.
2. The depth-enhanced 3D human pose estimation method according to claim 1, characterized in that, The extraction of depth features, two-dimensional human pose, and planar features from the image to be estimated specifically includes: Input the image to be estimated into the depth estimator and the 2D pose estimator, respectively; The depth estimator extracts the depth features of the image to be estimated, and the two-dimensional human pose and planar features of the image to be estimated are extracted by the two-dimensional pose estimator.
3. The depth-enhanced 3D human pose estimation method according to claim 1, characterized in that, After aligning the depth features, the two-dimensional human pose, and the planar features, the method further includes: Determine the depth offset and depth weight based on the aligned depth features, and then adaptively sample the aligned depth features based on the depth offset and depth weight.
4. The depth-enhanced three-dimensional human pose estimation method according to claim 1 or 3, characterized in that, After aligning the depth features, the two-dimensional human pose, and the planar features, the method further includes: The plane offset and plane weight are determined based on the aligned plane features, and the aligned plane features are adaptively sampled based on the plane offset and plane weight.
5. The depth-enhanced 3D human pose estimation method according to claim 1, characterized in that, The process of fusing the enhanced depth features, 2D human pose, and planar features to obtain fused features specifically includes: The enhanced depth features, 2D human pose, and planar features are stitched together to obtain the stitched features; Based on the splicing features, the fusion features are determined through a self-attention mechanism.
6. A depth-enhanced three-dimensional human pose estimation device, characterized in that, The depth-enhanced 3D human pose estimation device specifically includes: The feature extraction module is used to extract depth features, two-dimensional human pose, and planar features of the image to be estimated, and to perform an alignment operation on the depth features, the two-dimensional human pose, and the planar features, wherein the alignment operation is to sample the depth features and the planar features using the two-dimensional human pose; The feature enhancement module is used to regress the uncertainty of human joints based on the deep features, and to obtain enhanced deep features by applying the uncertainty of human joints to the deep features, wherein the uncertainty of human joints is used to control the participation of deep features in feature fusion. The feature fusion module is used to fuse the enhanced depth features, two-dimensional human pose, and planar features to obtain fused features; A human pose estimation module is used to estimate the three-dimensional human pose corresponding to the image to be estimated based on the fused features. Specifically, the uncertainty in regressing human joints based on the depth features includes: The depth features are input into an adaptive feature enhancement network, and the adaptive feature enhancement network outputs the uncertainty of the human joints corresponding to the image to be estimated and the depth value of the human joints. The adaptive feature enhancement network is described as a Bayesian neural network. The adaptive feature enhancement network applies the uncertainty of the human joints to the depth features to obtain enhanced depth features.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the steps in the depth-enhanced three-dimensional human pose estimation method as described in any one of claims 1-5.
8. A terminal device, characterized in that, include: Processor and memory; The memory stores a computer-readable program that can be executed by the processor; When the processor executes the computer-readable program, it implements the steps in the depth-enhanced three-dimensional human pose estimation method as described in any one of claims 1-5.
Citation Information
Patent Citations
Monocular three-dimensional human body posture estimation method and device based on diffusion model and feature fusion
CN119810917A