Depth-enhanced three-dimensional human body posture estimation method and device, equipment and medium

By extracting and enhancing depth features and combining two-dimensional human posture and plane features for feature fusion, the depth ambiguity problem in three-dimensional human posture estimation is solved, and the accuracy of the estimation is improved.

CN120108044AActive Publication Date: 2025-06-06PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510586856.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing three-dimensional human posture estimation method has depth ambiguity problems when mapping two-dimensional human coordinates to three-dimensional human coordinates, which affects the accuracy of the estimation.

Method used

The depth features, two-dimensional human postures and plane features of the image to be estimated are extracted, and the uncertainty of the human body's joint nodes are returned based on the depth features, and they are applied to the depth features to obtain enhanced depth features. Then the feature fusion is carried out to obtain the fusion features, and finally the three-dimensional human posture is estimated based on the fusion features.

Benefits of technology

By accurately extracting and explicitly characterizing depth features, and using uncertainty to adaptively enhance the depth features, effectively utilize depth information to mitigate the impact of depth blur, and improve the accuracy of three-dimensional human posture estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108044A_ABST
    Figure CN120108044A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, in particular to a depth-enhanced three-dimensional human body posture estimation method and device, equipment and a medium. The method comprises the steps of extracting a depth feature, a two-dimensional human body posture and a plane feature of a to-be-estimated image; regressing the uncertainty of the human body joint points based on the depth features, and acting the uncertainty of the human body joint points on the depth features to obtain enhanced depth features; performing feature fusion on the enhanced depth feature, the two-dimensional human body posture and the plane feature to obtain a fusion feature; and estimating a three-dimensional human body posture of the to-be-estimated image based on the fused features. According to the method, firstly, the depth feature in the human body posture is captured by accurately extracting and explicitly representing the depth feature, and then the depth feature is adaptively enhanced by using uncertainty, so that effective depth information is ensured to be reasonably utilized, interference of noise in the depth feature on a posture estimation task is avoided, the influence of depth blur is effectively relieved, and the accuracy of posture estimation is improved. And the accuracy of three-dimensional human body posture estimation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a depth-enhanced three-dimensional human posture estimation method, device, equipment and medium. Background Art

[0002] Three-dimensional human pose estimation not only provides support for complex application scenarios, but also shows great potential in many fields. For example, in the fields of intelligent monitoring, virtual reality, augmented reality, and medical health, accurate human pose estimation can help improve the intelligence of the system. Three-dimensional human pose estimation aims to infer the position of each joint of the human body in three-dimensional space from images or videos. The current mainstream implementation method is a two-stage method based on lifting. In the first stage, the two-dimensional human coordinates are estimated using a ready-made two-dimensional human pose estimator. In the second stage, the two-dimensional human coordinates identified in the first stage are used as input to lift the two-dimensional human coordinates to three-dimensional human coordinates. However, when mapping two-dimensional human coordinates to three-dimensional human coordinates, there is an inherent depth ambiguity problem, which in turn affects the accuracy of three-dimensional human pose estimation.

[0003] Therefore the prior art still needs to be improved and enhanced. Summary of the invention

[0004] The technical problem to be solved by the present application is to provide a depth-enhanced three-dimensional human posture estimation method, device, equipment and medium in view of the deficiencies in the prior art.

[0005] In order to solve the above technical problems, the first aspect of the present application provides a depth-enhanced 3D human body posture estimation method, wherein the depth-enhanced 3D human body posture estimation method specifically includes: Extract the depth features, two-dimensional human posture and plane features of the image to be estimated; Regressing the uncertainty of human joints based on the depth features, and applying the uncertainty of the human joints to the depth features to obtain enhanced depth features; Fusing the enhanced depth features, the two-dimensional human body posture and the plane features to obtain fused features; The three-dimensional human body posture corresponding to the image to be estimated is estimated based on the fusion feature.

[0006] The depth-enhanced three-dimensional human posture estimation method, wherein the extracting of the depth features, two-dimensional human posture and plane features of the image to be estimated specifically includes: Input the image to be estimated into the depth estimator and the two-dimensional pose estimator respectively; The depth feature of the image to be estimated is extracted by the depth estimator, and the two-dimensional human body posture and plane feature of the image to be estimated are extracted by the two-dimensional posture estimator.

[0007] The depth-enhanced three-dimensional human posture estimation method, wherein, after extracting the depth features, two-dimensional human posture and plane features of the image to be estimated, the method further comprises: An alignment operation is performed on the depth feature, the two-dimensional human body posture, and the plane feature, wherein the alignment operation is to sample the depth feature and the plane feature using the two-dimensional human body posture.

[0008] The depth-enhanced three-dimensional human posture estimation method, wherein, after aligning the depth feature, the two-dimensional human posture and the plane feature, the method further comprises: A depth offset and a depth weight are determined based on the aligned depth features, and the aligned depth features are adaptively sampled based on the depth offset and the depth weight.

[0009] The depth-enhanced three-dimensional human posture estimation method, wherein, after aligning the depth feature, the two-dimensional human posture and the plane feature, the method further comprises: A plane offset and a plane weight are determined based on the aligned plane features, and adaptive sampling is performed on the aligned plane features based on the plane offset and the plane weight.

[0010] The depth-enhanced three-dimensional human posture estimation method, wherein the uncertainty of regressing human joint points based on the depth feature specifically includes: Inputting the depth feature into an adaptive feature enhancement network, and outputting the uncertainty of the human joint points corresponding to the image to be estimated and the depth value of the human joint points through the adaptive feature enhancement network, wherein the adaptive feature enhancement network is expressed as a Bayesian neural network; The uncertainty of the human body joint points is applied to the depth feature through the adaptive feature enhancement network to obtain enhanced depth features.

[0011] The depth-enhanced three-dimensional human posture estimation method, wherein the step of fusing the enhanced depth features, the two-dimensional human posture and the plane features to obtain the fused features specifically includes: Splicing and connecting the enhanced depth feature, the two-dimensional human body posture and the plane feature to obtain a spliced ​​feature; Based on the concatenated features, a fusion feature is determined through a self-attention mechanism.

[0012] A second aspect of the present application provides a depth-enhanced 3D human posture estimation device, wherein the depth-enhanced 3D human posture estimation device specifically comprises: A feature extraction module is used to extract the depth features, two-dimensional human body posture and plane features of the image to be estimated; A feature enhancement module, used for regressing the uncertainty of human joint points based on the depth feature, and applying the uncertainty of the human joint points to the depth feature to obtain an enhanced depth feature; A feature fusion module, used for fusing the enhanced depth feature, the two-dimensional human body posture and the plane feature to obtain a fused feature; A human posture estimation module is used to estimate the three-dimensional human posture corresponding to the image to be estimated based on the fusion feature.

[0013] A third aspect of the present application provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in any of the depth-enhanced three-dimensional human posture estimation methods described above.

[0014] A fourth aspect of the present application provides a terminal device, comprising: a processor and a memory; The memory stores a computer-readable program executable by the processor; When the processor executes the computer-readable program, the steps in any of the above-described depth-enhanced three-dimensional human posture estimation methods are implemented.

[0015] Beneficial effects: Compared with the prior art, the present application provides a method, device, equipment and medium for three-dimensional human posture estimation with depth enhancement, the method comprising extracting depth features, two-dimensional human posture and plane features of the image to be estimated; regressing the uncertainty of human joint points based on the depth features, and applying the uncertainty of human joint points to the depth features to obtain enhanced depth features; fusing the enhanced depth features, two-dimensional human posture and plane features to obtain fused features; estimating the three-dimensional human posture of the image to be estimated based on the fused features. The present application firstly captures the depth features in human posture by accurately extracting and explicitly representing the depth features, and then adaptively enhances the depth features using uncertainty, ensuring that effective depth information is reasonably utilized and avoiding the interference of noise in the depth features on the posture estimation task, effectively mitigating the impact of depth blur, and improving the accuracy of three-dimensional human posture estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 A flowchart of a depth-enhanced three-dimensional human posture estimation method provided in an embodiment of the present application.

[0018] Figure 2 An example diagram of a 3D human body pose.

[0019] Figure 3 This is a principle block diagram of a depth-enhanced three-dimensional human posture estimation device provided in an embodiment of the present application.

[0020] Figure 4 A functional block diagram of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] The embodiments of the present application provide a method, device, equipment and medium for depth-enhanced three-dimensional human posture estimation. In order to make the purpose, technical solution and effect of the present application clearer and more specific, the present application is further described in detail with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0022] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0023] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with those in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.

[0024] It should be understood that the sequence numbers and sizes of the steps in this embodiment do not mean the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.

[0025] The application content is further explained below through the description of embodiments in conjunction with the accompanying drawings.

[0026] The present embodiment provides a depth-enhanced three-dimensional human posture estimation method, which can utilize a trained three-dimensional human posture estimation model, that is, the image to be estimated can be input into the three-dimensional human posture estimation model, and the three-dimensional human posture can be input through the three-dimensional human posture estimation model, wherein the three-dimensional human posture estimation model can be a deep learning model, and specifically includes a feature extraction module, a feature enhancement module, a feature fusion module and a human posture estimation module, the feature extraction module is used to extract the depth features, two-dimensional human posture and plane features of the image to be estimated, the feature enhancement module is used to regress the uncertainty of human joint points based on the depth features, and act on the uncertainty of human joint points to obtain enhanced depth features, the feature fusion module is used to fuse the enhanced depth features, two-dimensional human posture and plane features to obtain fused features; the human posture estimation module is used to estimate the three-dimensional human posture corresponding to the image to be estimated based on the fused features.

[0027] In order to specifically illustrate the depth-enhanced three-dimensional human body posture estimation method provided in the embodiment of the present application, the specific implementation process of the depth-enhanced three-dimensional human body posture estimation method is described below in the form of process steps. The specific process of each module in the three-dimensional human body posture estimation model to implement its corresponding function can refer to the specific description of its corresponding step in the process steps, so as to facilitate the explanation of the specific implementation process of the three-dimensional human body posture estimation model to realize three-dimensional human body posture estimation.

[0028] like Figure 1 As shown, the depth-enhanced 3D human posture estimation method provided in the embodiment of the present application specifically includes: S10, extracting depth features, two-dimensional human body posture and plane features of the image to be estimated.

[0029] Specifically, the image to be estimated can be an RGB image acquired by an image acquisition device, or an RGB image transmitted by an external device, etc. For example, the image to be estimated is a monocular RGB image. The depth feature is used to reflect the depth information related to the human body in the image to be estimated, and is used to assist in the mapping of two-dimensional human body posture to three-dimensional human body posture, and to alleviate the depth ambiguity by explicitly providing depth information as a reference. The plane feature is used to reflect the plane information related to the human body in the image to be estimated, and is used to assist in providing background context information. The two-dimensional human body posture is used to reflect the two-dimensional human body posture information extracted from the image to be estimated, wherein the depth feature, the two-dimensional human body posture and the plane feature can be obtained by extracting features from the image to be estimated by a feature extraction module, and the feature extraction module may include a depth estimator and a two-dimensional posture estimator, and the depth feature is extracted by the depth estimator, and the two-dimensional human body posture and plane feature are extracted by the two-dimensional posture estimator.

[0030] Based on this, the extraction of depth features, two-dimensional human body posture and plane features of the image to be estimated specifically includes: Input the image to be estimated into the depth estimator and the two-dimensional pose estimator respectively; The depth feature of the image to be estimated is extracted by the depth estimator, and the two-dimensional human body posture and plane feature of the image to be estimated are extracted by the two-dimensional posture estimator.

[0031] Specifically, the depth estimator is used to extract the depth features of the image to be estimated, and the two-dimensional posture estimator is used to perform two-dimensional posture estimation on the image to be estimated, so as to extract the two-dimensional human body posture and image plane features of the image to be estimated. That is to say, the feature extraction model in the three-dimensional human body posture estimation model may include a depth estimator and a two-dimensional posture estimator. The image to be estimated is respectively input into the depth estimator and the two-dimensional posture estimator in the feature extraction model, and the depth features are output through the depth estimator, and the two-dimensional human body posture and plane features are output through the two-dimensional posture estimator. For example, for a monocular RGB image , Indicates the image height, Represents the image width, 3 represents the number of channels of the image, and the depth estimator is used to extract pixel-level depth features. , extracting 2D human pose through 2D pose estimator and plane features, where represents the number of human joints, 2 represents the two-dimensional coordinates of the joints, and C represents the feature dimension.

[0032] Further, after obtaining the depth features, two-dimensional human body postures and plane features, in order to avoid unnecessary computational overhead caused by global processing of three-dimensional human body postures based on the depth features, two-dimensional human body postures and plane features, the depth features, two-dimensional human body postures and plane features may be aligned. The alignment operation is to sample the depth features and the plane features using the two-dimensional human body postures, and the alignment operation may be expressed as: , , in, represents the aligned plane features, represents the aligned deep features, Represents a sampling operation, Represents a plane feature, represents the deep features, Represents a 2D human pose.

[0033] The embodiment of the present application samples depth features and plane features through two-dimensional human body posture, and can extract rich depth and plane information related to human body joints from the depth features and plane features, and filter out background information irrelevant to human body joints contained in the depth features and plane features, thereby avoiding interference from background noise in the image to be estimated.

[0034] Further, after the depth features, two-dimensional human body postures and plane features are aligned, the depth features, two-dimensional human body postures and plane features after the alignment operation can be directly used to estimate the three-dimensional human body posture. However, in actual scenarios, since the two-dimensional posture estimator used to extract the two-dimensional human body posture inevitably introduces errors, the joint point coordinates extracted in the two-dimensional human body posture are not true values, which will cause the information about the real joint points in the features to be lost. For this reason, the embodiment of the present application can perform an adaptive sampling operation on the aligned depth features and / or aligned plane features after the depth features, two-dimensional human body postures and plane features are aligned, wherein, in addition, the execution process of the adaptive sampling operation of the depth features and the plane features can be the same, and the process of the adaptive sampling operation can first determine the offset and weight based on the original features, then sample the original features based on the offset, and finally weight the sampled features with the original features based on the weights. The difference between the two lies in the different original features. When the adaptive sampling operation is performed on the depth feature, the original feature is the depth feature, and when the adaptive sampling operation is performed on the plane feature, the original feature is the plane feature. That is, the execution process of the adaptive sampling operation corresponding to the depth feature is to determine the depth offset and the depth weight based on the aligned depth feature, and adaptively sample the aligned depth feature based on the depth offset and the depth weight; the execution process of the adaptive sampling operation corresponding to the plane feature is to determine the plane offset and the plane weight based on the aligned plane feature, and adaptively sample the aligned plane feature based on the plane offset and the plane weight.

[0035] Based on this, taking the deep feature as an example, the execution process of the adaptive sampling operation can be expressed as: , , , in, represents the depth weight, represents a linear mapping, Indicates the depth offset, Represents the deep features after adaptive sampling.

[0036] In addition, it should be noted that first It is determined by performing softmax normalization after linear mapping. It is obtained by executing the tanh activation function after linear mapping, which is used here Secondly, when adaptively sampling the depth features and / or plane features, it can be performed once or multiple times. Here, one adaptive sampling is used as an example for explanation. In actual applications, the number of adaptive sampling can be determined according to actual needs, and no specific limitation is made here.

[0037] In the embodiment of the present application, after the depth features, two-dimensional human body posture and plane features are aligned, adaptive sampling is performed from the aligned depth features and / or plane features, so that the depth features and / or plane features after adaptive sampling contain more complete feature information related to human body joints, thereby avoiding the loss of information about real human body joints in the depth features and / or plane features due to sampling using two-dimensional human body posture, thereby improving the accuracy of the three-dimensional human body posture subsequently determined based on the depth features, two-dimensional human body posture and plane features.

[0038] S20, regressing the uncertainty of human joints based on the depth features, and applying the uncertainty of human joints to the depth features to obtain enhanced depth features.

[0039] Specifically, the uncertainty of the human joints is used to reflect the credibility of the depth feature, and the uncertainty of the human joints includes the uncertainty of each human joint in the image to be estimated, that is, the uncertainty of the human joints regressed based on the depth feature is an uncertainty sequence, and the uncertainty sequence includes the uncertainty of each human joint in the image to be estimated. Among them, the uncertainty of the human joints can be determined by regressing the depth feature, that is, the depth feature is used as the input of the regression head, and the uncertainty of the human joints is output through the regression head. Among them, when the depth feature is aligned and adaptively sampled, the depth feature input to the regression head is the depth feature after adaptive sampling, and when only the depth feature is aligned, the depth feature input to the regression head is the depth feature after alignment.

[0040] In one implementation, the feature enhancement module in the 3D human posture estimation model can use an adaptive feature enhancement network to output enhanced depth features. Accordingly, the uncertainty of human joints regressed based on the depth features, and the uncertainty of human joints applied to the depth features to obtain enhanced depth features specifically include: Inputting the depth feature into an adaptive feature enhancement network, and outputting the uncertainty of the human joint points corresponding to the image to be estimated and the depth value of the human joint points through the adaptive feature enhancement network; The uncertainty of the human body joint points is applied to the depth feature through the adaptive feature enhancement network to obtain enhanced depth features.

[0041] Specifically, the adaptive feature enhancement network is used to determine the uncertainty of the human joint points corresponding to the image to be estimated. At the same time, the adaptive feature enhancement network can also output the depth value of the human joint points, and will apply the uncertainty of the human joint points to the depth features to obtain enhanced depth features, and control the participation of the depth features in feature fusion, wherein the depth value of the human joint points is a depth value sequence, which includes the depth value of each human joint point in the image to be estimated.

[0042] Taking the depth feature after adaptive sampling as an example, the process of determining the enhanced depth feature can be expressed as: , , , in, Indicates the depth value of the human joint point, represents the uncertainty of the human joint points, represents a multi-layer perceptron, Represents enhanced deep features.

[0043] Furthermore, in this application, it is necessary to know which parts of the depth features have a higher noise level, and it is necessary to predict the uncertainty of each human joint point. To this end, the adaptive feature enhancement network is expressed as a Bayesian neural network, and the Gaussian distribution of the depth features is obtained through the Bayesian neural network to obtain the depth value (i.e., the mean of the Gaussian distribution) and the uncertainty (i.e., the variance of the Gaussian distribution). In this way, the Gaussian distribution of the human body can be estimated through the adaptive feature enhancement network, not just predicting the absolute coordinates of the human body joint points. Among them, in the training process of the adaptive feature enhancement network, the uncertainty of the human body joint points is used to represent numerical stability, and the Bayesian loss constraint is adopted (reducing the uncertainty S while making μ closer to the true value of the depth) to make the depth distribution of the joint points closer to the true value, wherein the Bayesian loss constraint can be expressed as: , in, represents the loss function during training of the adaptive feature enhancement network, represents the number of training samples included in the training batch, Indicates The labeled depth value of training samples, Indicates The predicted depth value of training samples, Indicates The prediction uncertainty of each training sample.

[0044] S30, fusing the enhanced depth feature, the two-dimensional human body posture and the plane feature to obtain a fused feature.

[0045] Specifically, the fused feature is obtained by multimodal information interaction between the enhanced depth feature, the two-dimensional human body posture and the plane feature, wherein the multimodal information interaction can be realized by an attention mechanism. That is to say, the feature fusion module in the above-mentioned three-dimensional human body posture estimation model can be configured with an attention mechanism, and the fused feature is obtained by inputting the enhanced depth feature, the two-dimensional human body posture and the plane feature into the feature fusion module, and the fused feature is outputted through the feature fusion module.

[0046] Exemplarily, the step of fusing the enhanced depth feature, the two-dimensional human body posture, and the plane feature to obtain the fused feature specifically includes: Splicing and connecting the enhanced depth feature, the two-dimensional human body posture and the plane feature to obtain a spliced ​​feature; Based on the concatenated features, a fusion feature is determined through a self-attention mechanism.

[0047] Specifically, the enhanced depth features, two-dimensional human body posture and plane features are spliced ​​and connected, which means merging the tokens of the enhanced depth features, two-dimensional human body posture and plane features into a transformer encoder, and simulating the interaction between different features through the transformer encoder to realize the interactive information between the enhanced depth features, two-dimensional human body posture and plane features. The determination process of the fusion feature can be expressed as: , , in, Represents the splicing feature, represents the fusion feature, Represents a splicing operation, Represents multi-head self-attention.

[0048] S40: Estimate a 3D human body posture corresponding to the image to be estimated based on the fusion feature.

[0049] Specifically, after obtaining the fusion features, the 3D human posture can be predicted by the fusion features, for example, Figure 2 The 3D human posture shown. The 3D human posture can be estimated by a multi-layer perceptron, that is, the human posture estimation module in the above 3D human posture estimation model can adopt a multi-layer perceptron, and the process of estimating the 3D human posture based on the fusion feature by the human posture estimation module can be expressed as: , in, Represents 3D human posture.

[0050] In summary, this embodiment provides a depth-enhanced three-dimensional human posture estimation method, which includes extracting depth features, two-dimensional human posture and plane features of an image to be estimated; regressing the uncertainty of human joints based on the depth features, and applying the uncertainty of human joints to the depth features to obtain enhanced depth features; fusing the enhanced depth features, two-dimensional human posture and plane features to obtain fused features; estimating the three-dimensional human posture of the image to be estimated based on the fused features. This application first captures the depth features in human posture by accurately extracting and explicitly representing the depth features, and then uses uncertainty to adaptively enhance the depth features, ensuring that effective depth information is reasonably used and avoiding noise in the depth features from interfering with the posture estimation task, effectively mitigating the impact of depth blur, and improving the accuracy of three-dimensional human posture estimation.

[0051] Based on the above-mentioned depth-enhanced 3D human posture estimation method, this embodiment provides a depth-enhanced 3D human posture estimation device, such as Figure 3 As shown, the depth-enhanced 3D human posture estimation device specifically includes: A feature extraction module 100 is used to extract depth features, two-dimensional human body posture and plane features of the image to be estimated; A feature enhancement module 200, configured to regress the uncertainty of human joints based on the depth feature, and to obtain an enhanced depth feature by applying the uncertainty of the human joints to the depth feature; A feature fusion module 300 is used to fuse the enhanced depth feature, the two-dimensional human body posture and the plane feature to obtain a fused feature; The human body posture estimation module 400 is used to estimate the three-dimensional human body posture corresponding to the image to be estimated based on the fusion feature.

[0052] Based on the above-mentioned depth-enhanced three-dimensional human posture estimation method, this embodiment provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the depth-enhanced three-dimensional human posture estimation method as described in the above-mentioned embodiment.

[0053] Based on the above-mentioned depth-enhanced 3D human posture estimation method, the present application also provides a terminal device, such as Figure 4As shown, it includes at least one processor 20; a display screen 21; and a memory 22, and may also include a communications interface 23 and a bus 24. The processor 20, the display screen 21, the memory 22, and the communications interface 23 may communicate with each other through the bus 24. The display screen 21 is configured to display a preset user guide interface in the initial setting mode. The communications interface 23 may transmit information. The processor 20 may call the logic instructions in the memory 22 to execute the method in the above embodiment.

[0054] In addition, the logic instructions in the memory 22 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0055] The memory 22 is a computer-readable storage medium that can be configured to store software programs, computer executable programs, such as program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions or modules stored in the memory 22, that is, implementing the methods in the above embodiments.

[0056] The memory 22 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, a variety of media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, may also be a transient storage medium.

[0057] In addition, the specific process of loading and executing the multiple instruction processors in the above-mentioned storage medium and the terminal device has been described in detail in the above-mentioned method, and will not be described one by one here.

[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A depth-enhanced three-dimensional human posture estimation method, characterized in that: The depth-enhanced 3D human posture estimation method specifically includes: Extract the depth features, two-dimensional human posture and plane features of the image to be estimated; Regressing the uncertainty of human joints based on the depth features, and applying the uncertainty of the human joints to the depth features to obtain enhanced depth features; Fusing the enhanced depth features, the two-dimensional human body posture and the plane features to obtain fused features; The three-dimensional human body posture corresponding to the image to be estimated is estimated based on the fusion feature.

2. The depth-enhanced 3D human pose estimation method according to claim 1, characterized in that: The extracting of the depth feature, the two-dimensional human body posture and the plane feature of the image to be estimated specifically includes: Input the image to be estimated into the depth estimator and the two-dimensional pose estimator respectively; The depth feature of the image to be estimated is extracted by the depth estimator, and the two-dimensional human body posture and plane feature of the image to be estimated are extracted by the two-dimensional posture estimator.

3. The depth-enhanced 3D human pose estimation method according to claim 1, characterized in that: After extracting the depth features, two-dimensional human body posture and plane features of the image to be estimated, the method further includes: An alignment operation is performed on the depth feature, the two-dimensional human body posture, and the plane feature, wherein the alignment operation is to sample the depth feature and the plane feature using the two-dimensional human body posture.

4. The depth-enhanced 3D human posture estimation method according to claim 3, characterized in that: After aligning the depth feature, the two-dimensional human body posture and the plane feature, the method further includes: A depth offset and a depth weight are determined based on the aligned depth features, and the aligned depth features are adaptively sampled based on the depth offset and the depth weight.

5. The depth-enhanced 3D human posture estimation method according to claim 3 or 4, characterized in that: After aligning the depth feature, the two-dimensional human body posture and the plane feature, the method further includes: A plane offset and a plane weight are determined based on the aligned plane features, and adaptive sampling is performed on the aligned plane features based on the plane offset and the plane weight.

6. The depth-enhanced 3D human posture estimation method according to claim 1, characterized in that: The uncertainty of regressing human joint points based on the depth feature specifically includes: Inputting the depth feature into an adaptive feature enhancement network, and outputting the uncertainty of the human joint points corresponding to the image to be estimated and the depth value of the human joint points through the adaptive feature enhancement network, wherein the adaptive feature enhancement network is expressed as a Bayesian neural network; The uncertainty of the human body joint points is applied to the depth feature through the adaptive feature enhancement network to obtain enhanced depth features.

7. The depth-enhanced 3D human posture estimation method according to claim 1, characterized in that: The step of fusing the enhanced depth features, the two-dimensional human body posture and the plane features to obtain the fused features specifically includes: Splicing and connecting the enhanced depth feature, the two-dimensional human body posture and the plane feature to obtain a spliced ​​feature; Based on the concatenated features, a fusion feature is determined through a self-attention mechanism.

8. A depth-enhanced 3D human posture estimation device, characterized in that: The depth-enhanced 3D human posture estimation device specifically comprises: A feature extraction module is used to extract the depth features, two-dimensional human body posture and plane features of the image to be estimated; A feature enhancement module, used for regressing the uncertainty of human joint points based on the depth feature, and applying the uncertainty of the human joint points to the depth feature to obtain an enhanced depth feature; A feature fusion module, used for fusing the enhanced depth feature, the two-dimensional human body posture and the plane feature to obtain a fused feature; A human posture estimation module is used to estimate the three-dimensional human posture corresponding to the image to be estimated based on the fusion feature.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the depth-enhanced three-dimensional human posture estimation method as described in any one of claims 1-7.

10. A terminal device, characterized in that: include: Processor and memory; The memory stores a computer-readable program executable by the processor; When the processor executes the computer-readable program, the steps in the depth-enhanced three-dimensional human posture estimation method as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • A target attitude estimation method based on deep learning

    CN109903332A

  • Multi-person three-dimensional attitude estimation method and device and electronic equipment

    CN114550282A

  • Monocular three-dimensional human body posture estimation method and device based on diffusion model and feature fusion

    CN119810917A