Three-dimensional human body posture estimation method, medium and system

Through a three-dimensional human posture estimation method, using technologies such as joint node detection, high-dimensional feature embedding and feature fusion, the accuracy problem of existing methods when dealing with occlusion and depth uncertainty is solved, and more efficient posture estimation is achieved.

CN120014709APending Publication Date: 2025-05-16XIAMEN UNIV OF TECH +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510140254.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing three-dimensional human posture estimation method is difficult to efficiently capture the relationship between bones and joints when dealing with the obstruction of the human body or depth uncertainty, resulting in a decrease in the accuracy of posture estimation.

Method used

A three-dimensional human posture estimation method is proposed. By obtaining the human body image to be estimated, it conducts node detection, embeds it in high-dimensional space, expands the receptive field of feature information, learns the relationship between bone joint nodes, performs feature fusion, and inputs the regression head module to output predicted three-dimensional human body coordinates.

Benefits of technology

Effectively capture the relationship between bones and joints and improve the accuracy of posture estimation when the human body is blocked.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014709A_ABST
    Figure CN120014709A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional human body posture estimation method, medium and system, and the method comprises the steps: obtaining a to-be-estimated human body image, and carrying out the joint point detection, so as to obtain a two-dimensional joint point coordinate; embedding the two-dimensional joint point coordinates into a high-dimensional space to obtain high-dimensional feature information corresponding to the two-dimensional key point coordinates; enlarging a receptive field of the high-dimensional feature information to generate a high-dimensional feature vector; based on the high-dimensional feature vector, learning of a relation between skeleton joint points is carried out, and multi-element features are obtained; performing feature fusion on the multi-element features to obtain combined features corresponding to the multi-element features; inputting the combined features into a regression head module, and outputting predicted three-dimensional human body coordinates corresponding to the to-be-estimated human body image through the regression head module; the relation between skeleton joints can be effectively captured, and the estimation accuracy of the human body posture under the condition that the human body is shielded is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision, and in particular to a three-dimensional human posture estimation method, medium and system. Background Art

[0002] The task of 3D human pose estimation is to infer the pose of a human body in 3D space using image and video data.

[0003] At present, relevant methods have made some progress in improving the accuracy of posture estimation. However, there are still a series of challenges such as occlusion problems and depth uncertainty. Specifically, in real scenes, the human body may be occluded by other objects, or the limitation of the camera's viewing angle may lead to uncertainty in posture estimation; in addition, traditional methods are limited in capturing long-distance relationships and local dependencies, and cannot efficiently capture effective information, which seriously affects the accuracy of human posture prediction. Summary of the invention

[0004] The present invention aims to solve one of the technical problems in the related art to at least some extent. To this end, one object of the present invention is to propose a three-dimensional human body posture estimation method that can effectively capture the relationship between bone joints and improve the estimation accuracy of human body posture when the human body is occluded.

[0005] In the first aspect, the present invention proposes a three-dimensional human posture estimation method, comprising: obtaining a human body image to be estimated, and performing joint point detection on the human body image to be estimated to obtain two-dimensional joint point coordinates corresponding to the human body image to be estimated; embedding the two-dimensional joint point coordinates into a high-dimensional space to obtain high-dimensional feature information corresponding to the two-dimensional key point coordinates; expanding the receptive field of the high-dimensional feature information to generate a high-dimensional feature vector; learning the relationship between skeletal joint points based on the high-dimensional feature vector to obtain multi-element features; performing feature fusion on the multi-element features to obtain combined features corresponding to the multi-element features; inputting the combined features into a regression head module to output the predicted three-dimensional human body coordinates corresponding to the human body image to be estimated through the regression head module.

[0006] According to the three-dimensional human posture estimation method of the embodiment of the present invention, firstly, a human body image to be estimated is obtained, and joint point detection is performed on the human body image to be estimated to obtain the two-dimensional joint point coordinates corresponding to the human body image to be estimated; then, the two-dimensional joint point coordinates are embedded into a high-dimensional space to obtain high-dimensional feature information corresponding to the two-dimensional key point coordinates; then, the receptive field of the high-dimensional feature information is expanded to generate a high-dimensional feature vector; then, the relationship between the skeletal joint points is learned based on the high-dimensional feature vector to obtain a multi-element feature; then, the multi-element feature is feature fused to obtain a combined feature corresponding to the multi-element feature; then, the combined feature is input into a regression head module to output the predicted three-dimensional human body coordinates corresponding to the human body image to be estimated through the regression head module. Thereby, the relationship between the skeletal joints is effectively captured, and the estimation accuracy of the human body posture is improved when the human body is occluded.

[0007] In some embodiments, the high-dimensional feature vector is calculated by the following formula: in, represents a high-dimensional feature vector, Represents high-dimensional feature information, represents the symmetric normalized Laplacian matrix, represents the activation function, Indicates The coefficient matrix of the Chebyshev polynomials, represents the symmetric normalized Laplacian matrix Chebyshev polynomial of order, represents the order of Chebyshev polynomial, represents the Laplacian matrix.

[0008] In some embodiments, the Laplacian matrix is ​​expressed by the following formula: in, represents the Laplacian matrix, represents the augmented adjacency matrix of the graph, represents the adjacency matrix of the graph, represents the identity matrix, represents the symmetric normalized degree matrix.

[0009] In some embodiments, the relationship between skeletal joints is learned based on the high-dimensional feature vector to obtain multi-element features, including: inputting the high-dimensional feature vector into a multi-head self-attention block to output initial multi-element features corresponding to the high-dimensional feature vector through the multi-head self-attention block; and learning graph structured data based on the initial multi-element features using a dynamic adjacency matrix based on graph convolution to obtain multi-element features.

[0010] In some embodiments, the combined feature is calculated by the following formula: in, represents the combined features, Represents multi-element features, represents the activation function, Indicates The adjacency matrix of order, represents the identity matrix, represents the symmetric normalized degree matrix, Indicates The weight matrix corresponding to the adjacency matrix of order .

[0011] In some embodiments, the method further includes: calculating the error between the predicted three-dimensional human body coordinates and the actual three-dimensional human body coordinates based on a mean square error loss function to determine the model performance according to the error.

[0012] In a second aspect, an embodiment of the present invention proposes a computer-readable storage medium on which a three-dimensional human body posture estimation program is stored. When the three-dimensional human body posture estimation program is executed by a processor, the three-dimensional human body posture estimation method as described above is implemented.

[0013] In the third aspect, an embodiment of the present invention proposes a three-dimensional human posture estimation system, including: a joint point detection module, the joint point detection module is used to obtain a human body image to be estimated, and perform joint point detection on the human body image to be estimated to obtain two-dimensional joint point coordinates corresponding to the human body image to be estimated; a node embedding module, the node embedding module is used to embed the two-dimensional joint point coordinates into a high-dimensional space to obtain high-dimensional feature information corresponding to the two-dimensional key point coordinates; a Chebyshev graph convolution module, the Chebyshev graph convolution module is used to expand the receptive field of the high-dimensional feature information to generate a high-dimensional feature vector; a dynamic adjacency matrix module, the dynamic adjacency matrix module is used to learn the relationship between skeletal joints based on the high-dimensional feature vector to obtain multi-element features; a high-order graph convolution module, the high-order graph convolution module is used to perform feature fusion on the multi-element features to obtain combined features corresponding to the multi-element features; a regression head module, the regression head module is used to input the combined features into the regression head module, so as to output the predicted three-dimensional human body coordinates corresponding to the human body image to be estimated through the regression head module.

[0014] In some embodiments, the high-dimensional feature vector is calculated by the following formula: in, represents a high-dimensional feature vector, Represents high-dimensional feature information, represents the symmetric normalized Laplacian matrix, represents the activation function, Indicates The coefficient matrix of the Chebyshev polynomials, represents the symmetric normalized Laplacian matrix Chebyshev polynomial of order, represents the order of Chebyshev polynomial, represents the Laplacian matrix.

[0015] In some embodiments, the Laplacian matrix is ​​expressed by the following formula: in, represents the Laplacian matrix, represents the augmented adjacency matrix of the graph, represents the adjacency matrix of the graph, represents the identity matrix, represents the symmetric normalized degree matrix.

[0016] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic flow chart of a method for estimating a three-dimensional human body posture according to an embodiment of the present invention; Figure 2 is a schematic diagram of the prediction result effect according to an embodiment of the present invention; Figure 3 is another schematic diagram of prediction results according to an embodiment of the present invention; Figure 4 is a block diagram of a 3D human posture estimation system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0018] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.

[0019] The following describes a three-dimensional human body posture estimation method according to an embodiment of the present invention with reference to the accompanying drawings.

[0020] See also Figure 1 , Figure 1 FIG. 4 is a flow chart of a method for estimating a 3D human body posture according to an embodiment of the present invention. Figure 1 As shown, the 3D human body posture estimation method comprises the following steps: S101, obtaining a human body image to be estimated, and performing joint point detection on the human body image to be estimated to obtain two-dimensional joint point coordinates corresponding to the human body image to be estimated.

[0021] S102, embedding the two-dimensional joint point coordinates into a high-dimensional space to obtain high-dimensional feature information corresponding to the two-dimensional key point coordinates.

[0022] As an example, the two-dimensional joint point coordinates can be input into the node embedding module to embed the two-dimensional joint point coordinates into a high-dimensional space through a linear layer with position embedding to obtain high-dimensional feature information.

[0023] S103, expanding the receptive field of the high-dimensional feature information to generate a high-dimensional feature vector.

[0024] In some embodiments, the high-dimensional feature vector is calculated by the following formula: in, represents a high-dimensional feature vector, Represents high-dimensional feature information, represents the symmetric normalized Laplacian matrix, represents the activation function, Indicates The coefficient matrix of the Chebyshev polynomials, represents the symmetric normalized Laplacian matrix Chebyshev polynomial of order, represents the order of Chebyshev polynomial, represents the Laplacian matrix.

[0025] In some embodiments, the Laplacian matrix is ​​expressed by the following formula: in, represents the Laplacian matrix, represents the augmented adjacency matrix of the graph, represents the adjacency matrix of the graph, represents the identity matrix, represents the symmetric normalized degree matrix.

[0026] As an example, first, after embedding the two-dimensional joint point coordinates into the high-dimensional space, the corresponding high-dimensional feature information is obtained; then, the Chebyshev graph convolution is used to enhance the network's modeling ability for different graph structures; specifically, the Chebyshev graph convolution is expressed by the following formula: in, represents a high-dimensional feature vector, Represents high-dimensional feature information, represents the symmetric normalized Laplacian matrix, represents the activation function, Indicates The coefficient matrix of the Chebyshev polynomials, represents the symmetric normalized Laplacian matrix Chebyshev polynomial of order, represents the order of Chebyshev polynomial, represents the Laplacian matrix.

[0027] Specifically, in Chebyshev graph convolution, the convolution kernel represents The order polynomial graph Laplacian is called the symmetric normalized Laplacian matrix. Therefore, the Chebyshev graph convolution can integrate the The information of adjacent joint points is used to expand the receptive field, where the Laplacian matrix is ​​expressed by the following formula: in, represents the Laplacian matrix, represents the augmented adjacency matrix of the graph, represents the adjacency matrix of the graph, represents the identity matrix, represents the symmetric normalized degree matrix.

[0028] It can be understood that Chebyshev graph convolution is more suitable for graph structured data.

[0029] S104, learning the relationship between skeletal joints based on the high-dimensional feature vector to obtain multi-element features.

[0030] In some embodiments, the relationship between skeletal joints is learned based on high-dimensional feature vectors to obtain multi-element features, including: inputting the high-dimensional feature vector into a multi-head self-attention block to output initial multi-element features corresponding to the high-dimensional feature vector through the multi-head self-attention block; and learning graph structured data based on the initial multi-element features using a dynamic adjacency matrix based on graph convolution to obtain multi-element features.

[0031] As an example, the system includes a graph-based dynamic adjacency matrix module, which includes a first normalization layer, a multi-head self-attention mechanism layer, a first regularization layer, a second normalization layer, a dynamic adjacency matrix, and a second regularization layer; first, the input high-dimensional feature vector is normalized by the first normalization layer; then, the normalized data is input into the multi-head self-attention mechanism layer to output multiple elements (i.e., initial multi-element features) through the multi-head self-attention mechanism layer; wherein each element of the output contains all two-dimensional joint point information. Then, the multiple elements are regularized and normalized for the second time by the first regularization layer and the second normalization layer; then, the processed data is input into the dynamic adjacency matrix, so that the graph structured data can be learned more flexibly by the dynamic adjacency matrix. It can be understood that by using the dynamic adjacency matrix to represent the dynamic relationship between joints, the accuracy of human posture prediction under human occlusion and depth ambiguity can be improved.

[0032] S105, performing feature fusion on the multi-element features to obtain combined features corresponding to the multi-element features.

[0033] In some embodiments, the combined feature is calculated by the following formula: in, represents the combined features, Represents multi-element features, represents the activation function, Indicates The adjacency matrix of order, represents the identity matrix, represents the symmetric normalized degree matrix, Indicates The weight matrix corresponding to the adjacency matrix of order .

[0034] As an example, the system includes a high-order graph convolutional network to assign different weight matrices to neighbors with different hop counts through the high-order graph convolutional network; and to more accurately and precisely model multi-order neighbors by summing up. The high-order graph convolutional network is expressed by the following formula: in, represents the combined features, Represents multi-element features, represents the activation function, Indicates The adjacency matrix of order, represents the identity matrix, represents the symmetric normalized degree matrix, Indicates The weight matrix corresponding to the adjacency matrix of order .

[0035] It can be understood that the relationship between multi-order neighbors can be better captured through high-order graph convolutional networks, which is achieved by assigning different weight matrices to neighbors of different orders and combining their features by summing.

[0036] S106, inputting the combined features into a regression head module, so as to output predicted three-dimensional human coordinates corresponding to the human body image to be estimated through the regression head module.

[0037] In some embodiments, the method further includes: calculating the error between the predicted three-dimensional human body coordinates and the actual three-dimensional human body coordinates based on a mean square error loss function to determine the model performance according to the error.

[0038] As an example, the mean square error loss function is used to minimize the error between the three-dimensional predicted value and the actual coordinate. The mean square error loss function is used to measure the difference between the model predicted value and the true value. Specifically, the mean square error loss function is expressed as: in, represents the number of samples, Indicates The true value of the samples, Represents the model for The predicted value of a sample.

[0039] It can be understood that the smaller the mean square error loss function value is, the closer the model's predicted value is to the true value, that is, the better the model performance is.

[0040] In order to better illustrate the effect of the three-dimensional human body posture estimation method proposed in the embodiment of the present invention on the estimation effect of the human body posture, Figure 2 and Figure 3 For example, Figure 2 and Figure 3 The two-dimensional picture of the original input, the predicted value and the true value obtained by the three-dimensional human posture estimation method proposed in the present invention based on the original input are shown in FIG. Figure 2 and Figure 3 It can be seen that the 3D human body posture estimation method proposed in the present invention is very close to the true value when estimating human body posture, and the effect is good.

[0041] In summary, according to the three-dimensional human posture estimation method of the embodiment of the present invention, firstly, a human body image to be estimated is obtained, and joint point detection is performed on the human body image to be estimated to obtain the two-dimensional joint point coordinates corresponding to the human body image to be estimated; then, the two-dimensional joint point coordinates are embedded into a high-dimensional space to obtain high-dimensional feature information corresponding to the two-dimensional key point coordinates; then, the receptive field of the high-dimensional feature information is expanded to generate a high-dimensional feature vector; then, the relationship between the skeletal joint points is learned based on the high-dimensional feature vector to obtain a multi-element feature; then, the multi-element feature is feature fused to obtain a combined feature corresponding to the multi-element feature; then, the combined feature is input into a regression head module to output the predicted three-dimensional human body coordinates corresponding to the human body image to be estimated through the regression head module. Thereby, the relationship between the skeletal joints is effectively captured, and the estimation accuracy of the human body posture is improved when the human body is occluded.

[0042] In a second aspect, an embodiment of the present invention proposes a computer-readable storage medium on which a three-dimensional human body posture estimation program is stored. When the three-dimensional human body posture estimation program is executed by a processor, the three-dimensional human body posture estimation method as described above is implemented.

[0043] In a third aspect, an embodiment of the present invention provides a 3D human body posture estimation system, such as Figure 3 As shown, the three-dimensional human posture estimation system includes: a joint point detection module 10, a node embedding module 20, a Chebyshev graph convolution module 30, a dynamic adjacency matrix module 40, a high-order graph convolution module 50 and a regression head module 60.

[0044] The joint point detection module 10 is used to obtain a human body image to be estimated, and perform joint point detection on the human body image to be estimated, so as to obtain two-dimensional joint point coordinates corresponding to the human body image to be estimated; The node embedding module 20 is used to embed the two-dimensional joint point coordinates into a high-dimensional space to obtain high-dimensional feature information corresponding to the two-dimensional key point coordinates; The Chebyshev graph convolution module 30 is used to expand the receptive field of high-dimensional feature information to generate a high-dimensional feature vector; The dynamic adjacency matrix module 40 is used to learn the relationship between the skeletal joints based on the high-dimensional feature vector to obtain multi-element features; The high-order graph convolution module 50 is used to perform feature fusion on multi-element features to obtain combined features corresponding to the multi-element features; The regression head module 60 is used to input the combined features into the regression head module so as to output the predicted three-dimensional human coordinates corresponding to the human image to be estimated through the regression head module.

[0045] In some embodiments, the high-dimensional feature vector is calculated by the following formula: in, represents a high-dimensional feature vector, Represents high-dimensional feature information, represents the symmetric normalized Laplacian matrix, represents the activation function, Indicates The coefficient matrix of the Chebyshev polynomials, represents the symmetric normalized Laplacian matrix Chebyshev polynomial of order, represents the order of Chebyshev polynomial, represents the Laplacian matrix.

[0046] In some embodiments, the Laplacian matrix is ​​expressed by the following formula: in, represents the Laplacian matrix, represents the augmented adjacency matrix of the graph, represents the adjacency matrix of the graph, represents the identity matrix, represents the symmetric normalized degree matrix.

[0047] It should be noted that the above description of the 3D human body posture estimation method is also applicable to the 3D human body posture estimation system, which will not be described in detail here.

[0048] In summary, according to the three-dimensional human posture estimation system of the embodiment of the present invention, a joint point detection module is set to obtain a human body image to be estimated, and joint point detection is performed on the human body image to be estimated to obtain the two-dimensional joint point coordinates corresponding to the human body image to be estimated; the node embedding module is used to embed the two-dimensional joint point coordinates into a high-dimensional space to obtain high-dimensional feature information corresponding to the two-dimensional key point coordinates; the Chebyshev graph convolution module is used to expand the receptive field of the high-dimensional feature information to generate a high-dimensional feature vector; the dynamic adjacency matrix module is used to learn the relationship between skeletal joints based on the high-dimensional feature vector to obtain multi-element features; the high-order graph convolution module is used to perform feature fusion on the multi-element features to obtain the combined features corresponding to the multi-element features; the regression head module is used to input the combined features into the regression head module to output the predicted three-dimensional human coordinates corresponding to the human body image to be estimated through the regression head module; thereby effectively capturing the relationship between skeletal joints and improving the estimation accuracy of human posture when the human body is occluded.

[0049] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.

[0050] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0051] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0052] In the description of the present invention, it is to be understood that the terms “center”, “longitudinal”, “lateral”, “length”, “width”, “thickness”, “up”, “down”, “front”, “back”, “left”, “right”, “vertical”, “horizontal”, “top”, “bottom”, “inside”, “outside”, “clockwise”, “counterclockwise”, “axial”, “radial”, “circumferential”, etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.

[0053] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0054] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0055] In the present invention, unless otherwise clearly specified and limited, a first feature being "above" or "below" a second feature may mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, a first feature being "above", "above" or "above" a second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. A first feature being "below", "below" or "below" a second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.

[0056] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A three-dimensional human body posture estimation method, characterized in that: The following steps are involved: Acquire a human body image to be estimated, and perform joint point detection on the human body image to be estimated to obtain two-dimensional joint point coordinates corresponding to the human body image to be estimated; Embedding the two-dimensional joint point coordinates into a high-dimensional space to obtain high-dimensional feature information corresponding to the two-dimensional key point coordinates; Expanding the receptive field of the high-dimensional feature information to generate a high-dimensional feature vector; Based on the high-dimensional feature vector, the relationship between the skeleton joints is learned to obtain multi-element features; Performing feature fusion on the multi-element features to obtain combined features corresponding to the multi-element features; The combined features are input into a regression head module, so as to output predicted three-dimensional human coordinates corresponding to the human image to be estimated through the regression head module.

2. The three-dimensional human body posture estimation method according to claim 1, characterized in that: The high-dimensional feature vector is calculated by the following formula: in, represents a high-dimensional feature vector, Represents high-dimensional feature information, represents the symmetric normalized Laplacian matrix, represents the activation function, Indicates The coefficient matrix of the Chebyshev polynomials, represents the symmetric normalized Laplacian matrix Chebyshev polynomial of order, represents the order of Chebyshev polynomial, represents the Laplacian matrix.

3. The three-dimensional human body posture estimation method according to claim 2, characterized in that: The Laplace matrix is ​​expressed by the following formula: in, represents the Laplacian matrix, represents the augmented adjacency matrix of the graph, represents the adjacency matrix of the graph, represents the identity matrix, represents the symmetric normalized degree matrix.

4. The three-dimensional human body posture estimation method according to claim 1, characterized in that: Based on the high-dimensional feature vector, the relationship between the skeletal joints is learned to obtain multi-element features, including: Inputting the high-dimensional feature vector into a multi-head self-attention block to output initial multi-element features corresponding to the high-dimensional feature vector through the multi-head self-attention block; The dynamic adjacency matrix based on graph convolution learns graph structured data according to the initial multi-element features to obtain multi-element features.

5. The three-dimensional human body posture estimation method according to claim 1, characterized in that: The combined feature is calculated by the following formula: in, represents the combined features, Represents multi-element features, represents the activation function, Indicates The adjacency matrix of order, represents the identity matrix, represents the symmetric normalized degree matrix, Indicates The weight matrix corresponding to the adjacency matrix of order .

6. The three-dimensional human body posture estimation method according to claim 1, characterized in that: Also includes: The error between the predicted three-dimensional human body coordinates and the actual three-dimensional human body coordinates is calculated based on a mean square error loss function to determine the model performance according to the error.

7. A computer-readable storage medium, characterized in that: A three-dimensional human body posture estimation program is stored thereon, and when the three-dimensional human body posture estimation program is executed by a processor, a three-dimensional human body posture estimation method as described in any one of claims 1 to 6 is implemented.

8. A three-dimensional human posture estimation system, characterized in that: include: A joint point detection module, the joint point detection module is used to obtain a human body image to be estimated, and perform joint point detection on the human body image to be estimated to obtain two-dimensional joint point coordinates corresponding to the human body image to be estimated; A node embedding module, the node embedding module is used to embed the two-dimensional joint point coordinates into a high-dimensional space to obtain high-dimensional feature information corresponding to the two-dimensional key point coordinates; A Chebyshev graph convolution module, wherein the Chebyshev graph convolution module is used to expand the receptive field of the high-dimensional feature information to generate a high-dimensional feature vector; A dynamic adjacency matrix module, wherein the dynamic adjacency matrix module is used to learn the relationship between skeletal joint points based on the high-dimensional feature vector to obtain multi-element features; A high-order graph convolution module, wherein the high-order graph convolution module is used to perform feature fusion on the multi-element features to obtain a combined feature corresponding to the multi-element features; A regression head module, wherein the regression head module is used to input the combined features into the regression head module so as to output the predicted three-dimensional human coordinates corresponding to the human image to be estimated through the regression head module.

9. The three-dimensional human posture estimation system according to claim 8, characterized in that: The high-dimensional feature vector is calculated by the following formula: in, represents a high-dimensional feature vector, Represents high-dimensional feature information, represents the symmetric normalized Laplacian matrix, represents the activation function, Indicates The coefficient matrix of the Chebyshev polynomials, represents the symmetric normalized Laplacian matrix Chebyshev polynomial of order, represents the order of Chebyshev polynomial, represents the Laplacian matrix.

10. The three-dimensional human posture estimation system according to claim 9, characterized in that: The Laplace matrix is ​​expressed by the following formula: in, represents the Laplacian matrix, represents the augmented adjacency matrix of the graph, represents the adjacency matrix of the graph, represents the identity matrix, represents the symmetric normalized degree matrix.

Citation Information

Cited By

  • Multi-view three-dimensional human body posture estimation method and system

    CN121544715A