Three-dimensional human body motion posture real-time estimation method and system based on deep learning

Through deep learning methods, the three-dimensional joint position is optimized, and the problems of high computational complexity and poor environmental adaptability in the existing technology are solved, and efficient and accurate estimation of three-dimensional human movement postures are achieved.

CN120259366AInactive Publication Date: 2025-07-04ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510368348.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing three-dimensional human posture estimation methods are highly complex in calculation, poor real-time and not adaptable to complex environments, especially in multi-person scenarios and dynamic environments, which are difficult to achieve accurate and real-time three-dimensional human movement posture estimation.

Method used

Deep learning method is adopted, and two-dimensional joint position information is extracted by combining convolutional neural networks and graph convolutional networks, and three-dimensional joint position is optimized through deep order perception and overlap loss perception, residual networks and knowledge distillation technology are introduced to improve model stability and performance, and iterative optimization is used for SMPL model, combining deep sorting networks and overlap loss avoidance functions to solve joint order and overlap problems.

Benefits of technology

It realizes efficient and accurate three-dimensional human body movement posture estimation, ensures real-time and adaptability to complex environments, reduces calculation complexity and improves the accuracy and robustness of the estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259366A_ABST
    Figure CN120259366A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of three-dimensional human body posture estimation, and provides a three-dimensional human body motion posture real-time estimation method and system based on deep learning, and the method comprises the steps: receiving an image or video sequence containing a human body; utilizing a deep neural network model to extract two-dimensional joint position information of the human body from the image or video sequence; inputting the two-dimensional joint position information into a three-dimensional attitude estimation module, sensing l oss and overlapping l oss through a depth sequence, and predicting and optimizing the three-dimensional joint position of the human body; outputting the estimated three-dimensional human body motion posture; an image or video sequence containing a human body is received, two-dimensional joint position information is extracted through a deep neural network model and input into a three-dimensional posture estimation module to predict and optimize the three-dimensional joint position, and finally the estimated three-dimensional human body motion posture is output, so that real-time estimation is achieved; the method has the remarkable beneficial effects that the three-dimensional human body motion posture can be efficiently and accurately estimated, and meanwhile, the real-time performance and the adaptability to a complex environment are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of three-dimensional human pose estimation, and specifically relates to a method and system for real-time estimation of three-dimensional human motion poses based on deep learning. Background Art

[0002] With the rapid development of computer vision and deep learning technologies, three-dimensional human pose estimation has been widely applied in many fields such as animation, virtual reality, human-computer interaction, video surveillance, and action recognition. However, most of the existing three-dimensional human pose estimation methods have problems such as high computational complexity, poor real-time performance, and weak adaptability to complex environments. Especially in multi-person scenarios and dynamic environments, how to accurately and real-time estimate the three-dimensional human motion poses remains a challenge.

[0003] Traditional three-dimensional human pose estimation methods usually rely on multi-camera systems or depth sensors. These methods not only have high equipment costs, but are also easily affected by occlusion and lighting conditions in complex environments. In recent years, with the rise of deep learning, three-dimensional human pose estimation methods based on deep neural networks have gradually emerged; these methods can directly estimate the three-dimensional poses of humans from single images or video sequences by training neural network models, and have advantages such as high computational efficiency and good real-time performance. However, most of the existing three-dimensional human pose estimation methods based on deep learning focus on improving the accuracy of estimation, while ignoring real-time performance and adaptability to complex environments.

[0004] Therefore, those skilled in the art have proposed a method and system for real-time estimation of three-dimensional human motion poses based on deep learning, aiming to solve the above problems in the prior art and achieve efficient, accurate, and real-time three-dimensional human motion pose estimation. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides a method and system for real-time estimation of three-dimensional human motion poses based on deep learning, so as to solve the problems that most of the existing three-dimensional human pose estimation methods based on deep learning focus on improving the accuracy of estimation, while ignoring real-time performance and adaptability to complex environments.

[0006] The method and system for real-time estimation of three-dimensional human motion poses based on deep learning include:

[0007] S1. Receive an image or video sequence containing a human body;

[0008] S2. Use a deep neural network model to extract two-dimensional joint position information of the human body from the image or video sequence;

[0009] S3. Input the two-dimensional joint position information into the three-dimensional pose estimation module, and predict and optimize the three-dimensional joint positions of the human body through depth order perception loss and overlap loss;

[0010] S4. Output the estimated three-dimensional human motion pose to achieve real-time estimation.

[0011] Preferably, in step S2, the deep neural network model includes a convolutional neural network layer for feature extraction and a graph convolutional network layer for joint position prediction.

[0012] Preferably, a residual network module is introduced into the deep neural network model to solve the problems of gradient vanishing and gradient explosion in the deep network, and improve the training stability and performance of the network.

[0013] Preferably, in step S3, the three-dimensional pose estimation module further includes a regression method based on the SMPL model, and improves the accuracy of three-dimensional pose estimation by iteratively optimizing the model parameters.

[0014] Preferably, in step S3, when optimizing the three-dimensional joint positions of the human body, the formulas of depth order perception loss and overlap loss are introduced.

[0015] Preferably, when performing the depth order perception loss step, a depth sorting network is used to learn the depth order relationship between joints and improve the accuracy of the three-dimensional pose in the depth direction;

[0016] When performing the overlap loss step, an overlap avoidance loss function is used to penalize the overlap situation between joints and ensure that the estimated three-dimensional pose is physically reasonable.

[0017] Preferably, in the optimization process, knowledge distillation technology is used to compress the knowledge of the large model into the small model, so as to reduce the computational amount while maintaining the performance.

[0018] The real-time three-dimensional human motion pose estimation system based on deep learning, using the above real-time three-dimensional human motion pose estimation method based on deep learning, includes:

[0019] An image / video receiving module for receiving an image or video sequence containing a human body;

[0020] A deep neural network model for extracting two-dimensional joint position information of the human body from the image or video sequence;

[0021] A three-dimensional pose estimation module connected to the deep neural network model for receiving the two-dimensional joint position information and predicting and optimizing the three-dimensional joint positions of the human body;

[0022] An output module for outputting the estimated three-dimensional human motion posture to achieve real-time estimation.

[0023] A processor configured to execute a method for real-time estimation of three-dimensional human motion posture based on deep learning as described above.

[0024] A computer-readable storage medium having a computer program stored thereon, where the computer program, when executed by a processor, implements a method for real-time estimation of three-dimensional human motion posture based on deep learning as described above.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] 1. By introducing a deep neural network model, especially combining a convolutional neural network and a graph convolutional network, the present invention realizes the efficient extraction of two-dimensional joint position information from an image or video sequence and accurately predicts the three-dimensional human posture, improving the accuracy and real-time performance of the estimation.

[0027] 2. By introducing a depth order perception loss and an overlap loss in the three-dimensional posture estimation module, the present invention effectively solves the problems of incorrect joint order in the depth direction and joint overlap, further improving the accuracy of three-dimensional posture estimation.

[0028] 3. By introducing a residual network module and a knowledge distillation technique, the present invention improves the training stability and performance of the deep neural network model, while reducing the computational amount, enabling the method to achieve real-time and accurate three-dimensional human motion posture estimation in a complex environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a flowchart of a method for real-time estimation of three-dimensional human motion posture based on deep learning according to the present invention;

[0030] Figure 2 It is a framework diagram of a system for real-time estimation of three-dimensional human motion posture based on deep learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The following further describes in detail the embodiments of the present invention in conjunction with the drawings and examples. The following examples are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.

[0032] Example: The present invention provides a method for real-time estimation of three-dimensional human motion posture based on deep learning, as Figure 1 shown, including:

[0033] S1. Receiving an image or video sequence containing a human body;

[0034] S2. Using a deep neural network model to extract two-dimensional joint position information of the human body from the image or video sequence;

[0035] S3. Input the two-dimensional joint position information into the three-dimensional pose estimation module, and predict and optimize the three-dimensional joint positions of the human body through the depth order perception loss and the overlap loss.

[0036] S4. Output the estimated three-dimensional human motion pose to achieve real-time estimation.

[0037] As can be seen from the above, by receiving an image or video sequence containing a human body, using a deep neural network model to extract two-dimensional joint position information, and inputting it into the three-dimensional pose estimation module to predict and optimize the three-dimensional joint positions, and finally outputting the estimated three-dimensional human motion pose to achieve real-time estimation; this method has significant beneficial effects, can efficiently and accurately estimate the three-dimensional human motion pose, while ensuring real-time performance and adaptability to complex environments, and solves the problems of insufficient estimation accuracy, real-time performance, and environmental adaptability in the prior art.

[0038] Further, in step S2, the deep neural network model includes a convolutional neural network layer for feature extraction and a graph convolutional network layer for joint position prediction; the convolutional neural network formula includes:

[0039] The input image or video frame is represented as X.

[0040] The output of the convolutional layer is represented as Y = f(W·X + b), where W is the convolutional kernel weight, b is the bias term, and f is the activation function.

[0041] The graph convolutional network formula includes:

[0042] The joint position information graph is represented as G = (V, E), where V is the set of nodes (joints) and E is the set of edges (joint connections).

[0043] The output of the graph convolutional layer is represented as where H (l) is the node feature matrix of the l-th layer, N (i) is the set of neighbor nodes of node i, C ij is the normalization constant, W (l) is the weight matrix of the l-th layer, and σ is the activation function.

[0044] As can be seen from the above, the deep neural network model adopted by the present invention realizes effective feature extraction and joint position prediction for human body images or video sequences by combining a convolutional neural network layer and a graph convolutional network layer; the convolutional neural network layer can capture local features in the image or video frame, while the graph convolutional network layer can use the connection relationship between joints for joint position prediction; this combination not only improves the accuracy of feature extraction, but also provides a more reliable basis for subsequent three-dimensional pose estimation, thus improving the overall performance of three-dimensional human motion pose estimation.

[0045] Further, a residual network module is introduced into the deep neural network model to solve the problems of gradient disappearance and gradient explosion in the deep network, and improve the training stability and performance of the network. The formula of the residual network module is as follows:

[0046] y = F(x, {W i}+ x;

[0047] where x is the input feature, F(x, {W i} is the residual function to be learned, and y is the output feature.

[0048] As can be seen from the above, by introducing the residual network module and learning the residual function between the input feature and the output feature, the problems of gradient disappearance and gradient explosion in the deep network are effectively solved; this introduction not only significantly improves the training stability and performance of the network, but also enables the network to learn the characteristics of human motion postures more deeply, thereby further enhancing the accuracy and robustness of 3D human motion posture estimation.

[0049] Further, in step S3, the 3D pose estimation module further includes a regression method based on the SMPL model. By iteratively optimizing the model parameters, the accuracy of 3D pose estimation is improved. The parametric human model formula includes:

[0050] The human pose is represented by θ (joint angles) and the shape is represented by β;

[0051] The human mesh vertices are represented by M(θ, β), where M is a function of the SMPL model;

[0052] The regression target is to minimize the error between the predicted mesh vertices and the true mesh vertices, that is:

[0053] min θ,β ||M(θ, β) - M true || 2 .

[0054] As can be seen from the above, the present invention further introduces a regression method based on the SMPL model in the 3D pose estimation module. By iteratively optimizing the model parameters, the human pose and shape can be represented more precisely, thereby significantly improving the accuracy of 3D pose estimation; this method uses the SMPL model to parametrically represent the human body and optimizes the model parameters by minimizing the error between the predicted mesh vertices and the true mesh vertices, making the estimated 3D human motion posture closer to the real situation and enhancing the reliability and practicality of the estimation results.

[0055] Further, in step S3, when optimizing the three-dimensional joint positions of the human body, the formulas of depth order perception loss and overlap loss are introduced; the depth order perception loss formula is used to ensure the correct order of joints in the depth direction, and its formula is as follows:

[0056] L′ = ∑ i,j∈joints max(0, d i - d j + ε);

[0057] Where, d i and d j are the depth values of joints i and j, and ε is a small positive number threshold;

[0058] The overlap loss formula is used to avoid the overlap between joints, and its formula is as follows:

[0059] L * = ∑ i,j∈joints max(0, -dist(p i , p j ) + r);

[0060] Where, p i and p j are the three-dimensional positions of joints i and j, dist(p i , p j ) is the distance between them, and r is the radius of the joint.

[0061] As can be seen from the above, when optimizing the three-dimensional joint positions of the human body, the present invention introduces the formulas of depth order perception loss and overlap loss. This innovative measure effectively solves the problems of incorrect order of joints in the depth direction and joint overlap; the depth order perception loss improves the accuracy of the three-dimensional pose in the depth direction by ensuring the correct order of joints in the depth direction; while the overlap loss ensures the physical reasonableness of the estimated three-dimensional pose by avoiding the overlap between joints; the introduction of these two loss functions further improves the accuracy and robustness of the three-dimensional human motion pose estimation, making the estimation result more in line with the actual situation of human motion.

[0062] Further, when performing the depth order perception loss step, a depth sorting network is adopted to learn the depth order relationship between joints and improve the accuracy of the three-dimensional pose in the depth direction; the formula of the depth sorting network is as follows:

[0063] D′ = DON(J 2D , F);

[0064] Where, J 2Dis the two-dimensional joint position, F is the additional feature (such as image feature, time series feature, etc.), and D′ is the output depth sorting information;

[0065] When performing the overlap loss step, an overlap avoidance loss function is adopted to punish the overlap situation between joints and ensure that the estimated three-dimensional pose is physically reasonable. The formula of the overlap avoidance loss function is as follows:

[0066]

[0067] where J i and J j are the three-dimensional positions of two joints, and ∈ is the overlap threshold.

[0068] As can be seen from the above, the present invention respectively adopts a depth sorting network and an overlap avoidance loss function; the depth sorting network effectively improves the accuracy of the three-dimensional pose in the depth direction by learning the depth order relationship between joints, making the estimated three-dimensional human motion pose more in line with reality in the depth direction; while the overlap avoidance loss function ensures that the estimated three-dimensional pose is physically reasonable by punishing the overlap situation between joints and avoids unreasonable overlap phenomena between joints; the combination of these two technologies further enhances the accuracy and reasonableness of the three-dimensional human motion pose estimation and improves the quality and credibility of the estimation results.

[0069] Furthermore, in the optimization process, the knowledge distillation technology is used to compress the knowledge of the large model into the small model, thereby reducing the computational amount while maintaining the performance. The formula of the knowledge distillation technology is as follows:

[0070] L d =L CE (Y s ,Y T )+λL KL (Y s ,Y t );

[0071] where L CE is the cross-entropy loss, L KL is the KL divergence loss, λ is the distillation coefficient, and Y s and Y t are the outputs of the student model and the teacher model respectively.

[0072] As can be seen from the above, in the optimization process, the present invention ingeniously uses the knowledge distillation technology to successfully compress the rich knowledge in the large model into the small model. This technology not only significantly reduces the computational amount and improves the operation efficiency, but also can achieve the lightweight of the model while maintaining the performance of three-dimensional human motion pose estimation. By adjusting the distillation coefficient and balancing the cross-entropy loss and KL divergence loss, the present invention ensures that the student model can fully learn the knowledge of the teacher model, thereby realizing more efficient real-time three-dimensional human motion pose estimation on the premise of ensuring the estimation accuracy.

[0073] A real-time three-dimensional human motion pose estimation system based on deep learning, such as Figure 2 shown, uses the above-mentioned real-time three-dimensional human motion pose estimation method based on deep learning, including:

[0074] An image / video receiving module for receiving an image or video sequence containing a human body;

[0075] A deep neural network model for extracting two-dimensional joint position information of a human body from an image or video sequence;

[0076] A three-dimensional pose estimation module connected to the deep neural network model for receiving the two-dimensional joint position information and predicting and optimizing the three-dimensional joint position of the human body;

[0077] An output module for outputting the estimated three-dimensional human motion pose to achieve real-time estimation.

[0078] Working principle: First, receive an image or video sequence containing a human body, then use the deep neural network model to extract the two-dimensional joint position information of the human body, then input this two-dimensional joint position information into the three-dimensional pose estimation module, predict and optimize the three-dimensional joint position of the human body through the depth-order perception loss and the overlap loss, and finally output the estimated three-dimensional human motion pose to achieve real-time estimation.

[0079] Furthermore, the effects of a real-time three-dimensional human motion pose estimation method according to an embodiment of the present invention are compared with a traditional three-dimensional human pose estimation method (comparative example), and the following table is obtained:

[0080] Project Comparative Example Example Computational Complexity High Lower Real-time Performance Poor Good Adaptability to Complex Environments Weak Strong Accuracy Medium High

[0081] As can be seen from the above table, 1) The traditional method usually relies on complex algorithms and multi-camera systems or depth sensors, resulting in a high computational complexity; while the method in the embodiment effectively reduces the computational complexity through a deep learning model, especially by using the combination of a convolutional neural network and a graph convolutional network;

[0082] 2) Traditional methods usually rely on complex algorithms and multi-camera systems or depth sensors, resulting in high computational complexity; while the method in the embodiment effectively reduces the computational complexity through a deep learning model, especially by using the combination of a convolutional neural network and a graph convolutional network;

[0083] 3) Traditional methods are prone to being affected by factors such as occlusion and lighting conditions in complex environments, resulting in poor estimation effects; the method in the embodiment enhances the adaptability of the model to complex environments by introducing innovative technologies such as depth order perception loss and overlap loss;

[0084] 4) Traditional methods are prone to being affected by factors such as occlusion and lighting conditions in complex environments, resulting in poor estimation effects; the method in the embodiment enhances the adaptability of the model to complex environments by introducing innovative technologies such as depth order perception loss and overlap loss.

[0085] An embodiment of the present application provides an electronic device, which is applicable to the above-mentioned real-time three-dimensional human motion posture estimation method based on deep learning, and includes:

[0086] A memory for storing computer programs and data;

[0087] A processor for running system programs.

[0088] An embodiment of the present application provides a computer storage medium, which is applicable to the above-mentioned real-time three-dimensional human motion posture estimation method based on deep learning, and performs hierarchical confidentiality management on the above system and data according to confidentiality management requirements.

[0089] Those skilled in the art should understand that the embodiments of the present application can be provided as a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0090] The present application is described with reference to the flowcharts and / or block diagrams of devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate for implementing the processes Figure 1one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks.

[0091] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks.

[0092] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks.

[0093] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0094] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0095] Computer-readable media includes both permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0096] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, article or apparatus. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, article or apparatus comprising the element.

[0097] The embodiments of the present invention are given for purposes of illustration and description. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A real-time estimation method for three-dimensional human motion postures based on deep learning, characterized in that: Including: S1. Receive an image or video sequence including a human body; S2. Use a deep neural network model to extract two-dimensional joint position information of the human body from the image or video sequence; S3. Input the two-dimensional joint position information into a three-dimensional pose estimation module, and predict and optimize the three-dimensional joint position of the human body through depth order perception loss and overlap loss; S4. Output the estimated three-dimensional human motion pose.

2. The real-time three-dimensional human motion posture estimation method based on deep learning according to claim 1, characterized in that: In step S2, the deep neural network model includes a convolutional neural network layer for feature extraction and a graph convolutional network layer for joint position prediction.

3. The real-time estimation method of three-dimensional human motion posture based on deep learning according to claim 2, characterized in that: A residual network module is introduced into the deep neural network model to solve the problems of gradient vanishing and gradient explosion in deep networks.

4. The real-time estimation method of three-dimensional human motion posture based on deep learning according to claim 1, characterized in that: In step S3, the three-dimensional pose estimation module further includes a regression method based on the SMPL model, and improves the accuracy of three-dimensional pose estimation by iteratively optimizing the model parameters.

5. The real-time estimation method of three-dimensional human motion posture based on deep learning according to claim 1, characterized in that: In step S3, when optimizing the three-dimensional joint position of the human body, the formulas of depth order perception loss and overlap loss are introduced.

6. The real-time three-dimensional human body motion posture estimation method based on deep learning according to claim 5, characterized in that: When performing the depth order perception loss step, a depth sorting network is adopted to learn the depth order relationship between joints; When performing the overlap loss step, an overlap avoidance loss function is adopted to punish the overlap situation between joints.

7. The real-time estimation method of three-dimensional human motion postures based on deep learning according to claim 5, characterized in that: During the optimization process, knowledge distillation technology is used to compress the knowledge of the large model into the small model, so as to reduce the computational amount while maintaining the performance.

8. A real-time three-dimensional human motion posture estimation system based on deep learning, characterized in that: Using the real-time estimation method for three-dimensional human motion pose based on deep learning according to any one of claims 1-7, including: An image / video receiving module for receiving an image or video sequence including a human body; A deep neural network model for extracting two-dimensional joint position information of the human body from the image or video sequence; A three-dimensional pose estimation module connected to the deep neural network model for receiving the two-dimensional joint position information and predicting and optimizing the three-dimensional joint position of the human body; An output module for outputting the estimated three-dimensional human motion pose.

9. A processor, characterized in that: Configured to execute a real-time estimation method for three-dimensional human motion pose based on deep learning according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Stored thereon is a computer program, and when the computer program is executed by a processor, it implements a real-time estimation method for three-dimensional human motion pose based on deep learning according to any one of claims 1 to 7.