A 3D Human Pose Spatial Feature Modeling Method
The local and global spatial features are captured through the Origin-based Part Transformer Block and Spatial Transformer Block in parallel structures, and combined with the timing characteristics of the Temporal Transformer Block, the accuracy problem of the 3D human pose estimation model under monocular data is solved, and a higher precision 3D pose estimation is achieved.
Patent Information
- Application Number
- CN202411718693.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-11-28
AI Technical Summary
Existing 3D human pose estimation models are difficult to accurately restore accurate 3D poses under monocular data, and existing methods fail to effectively capture spatial correlations between joints when designing Spatial Blocks.
The parallel structure of Origin-based Part Transformer Block and Spatial Transformer Block is adopted to capture local and global spatial features respectively, and time series features are captured through adaptive fusion and Temporal Transformer Block to generate more accurate 3D poses.
The accuracy and continuity of 3D pose estimation are improved, the modeling ability of spatial features in different ranges is enhanced, and more accurate 3D human pose estimation results are generated.
Smart Images

Figure CN119229027B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of 3D human pose estimation in computer vision, and more particularly, relates to a method for modeling 3D human pose spatial features. Background Art
[0002] Human pose estimation (HPE) is an important and challenging task in the field of computer vision, and is of great significance to many application fields such as action recognition, virtual reality, and human-robot interaction. The purpose of HPE is to predict the positions of each human joint from the input image or video. According to whether the predicted joints contain depth information, human pose estimation can be divided into 2D human pose estimation (2D HPE) and 3D human pose estimation (3D HPE). With the development of deep learning technology, the 2D HPE field has currently developed relatively maturely. The accuracy and generalization of 2D detectors have reached an advanced level, but the model output only contains limited 2D information. In contrast, although 3D HPE faces more challenges, the addition of depth information enables it to provide richer 3D spatial information for human poses and has a better understanding of human actions and interactions. Therefore, the 2D-to-3D method that applies the developed 2D detector to the 3D HPE task has become a typical monocular solution.
[0003] Due to the depth ambiguity problem in monocular data, multiple potential 3D poses may be mapped from the same 2D pose. Therefore, it is difficult to recover the accurate 3D pose only based on the single-frame 2D joint position information. Recently, driven by the ability of Transformer to capture long-range dependencies, the 2D-to-3D solution using video frame sequences containing motion temporal features has made significant progress. Among them, MixSTE further divides the entire human pose into multiple joints to model more refined temporal features. However, most of the existing solutions roughly use the entire 2D skeleton as the source of the model spatial features or simply divide it into several independent parts when designing the Spatial Block.
[0004] How to overcome the above-mentioned defects and deficiencies faced by 3D human pose estimation is a technical problem that needs to be solved currently. Summary of the Invention
[0005] The modeling performance of the Spatial Block in MixSTE for the spatial correlation between joints was observed, and the average dependence degree among 17 joints in the Human3.6M dataset was studied. An interpretable scheme was applied to average all the attention weights calculated in the middle layer of the model and then plot them as a confusion matrix. It can be observed that not all human joints have a high dependence in spatial position. The spatial correlation between joints that are far apart is generally low, such as the elbow joint and the knee joint. Such joint pairs are not directly related in the human geometric structure and are also independent of each other during movement. While joints belonging to the same part are closer in distance and generally have a higher spatial correlation, such as the hip joint, the knee joint, and the ankle joint. Such joint pairs are closely connected in the human geometric structure and also move synchronously during movement. In addition, it was observed that all joints have a high degree of association with the hip joint, which is called the Origin of the human skeleton. The position of the Origin directly determines the spatial position of the human body, which in turn affects the spatial positions of all other joints. Therefore, the Origin is used as a bridge connecting different parts to protect the integrity of the human pose. Specifically, the Origin is combined with each human part to form multiple Origin-based Parts. The spatial features of each Origin-based Part are modeled separately, and finally concatenated together to form the complete local spatial features.
[0006] In addition, the spatial relationships outside the parts were not ignored, such as the left ankle and right ankle joints, and the left knee and right knee joints. Although such joint pairs do not belong to the same part, they still have a certain degree of dependence. The relationships between all joints are modeled to generate global spatial features that are complementary to the local spatial features. Therefore, the proposed invention has a parallel structure design, and two parallel channels are responsible for capturing global and local spatial features respectively. Finally, the outputs of the two channels are fused and then passed through the Temporal Block to obtain the temporal features of each joint at multiple time steps to generate a more accurate 3D pose result.
[0007] The content of the present invention can be summarized into three aspects: 1. A new Origin-based PartTransformer (OPFormer) block is designed to model the spatial relationships of joints within different parts of the model, and at the same time, the Origin of the skeleton is used as a bridge connecting different parts to protect the integrity of the human pose.
[0008] 2. A new alternating network structure is designed, which has a dual-channel parallel structure to capture spatial features in different ranges and temporal blocks to capture the temporal features of different joints, so as to improve the results of 3D pose estimation.
[0009] 3. The overall network structure mainly consists of five components. The Origin-based Part Transformer Block and the Spatial Transformer Block form a two-channel parallel structure for capturing spatial features in different size ranges of a single-frame 2D pose. Among them, the proposed OPFormer Block is responsible for capturing the correlations between joints within a part and generating local spatial features. Also, the Spatial Transformer Block is used to capture the dependencies of all human joints and generate global spatial features. The results of the two channels are fused in an adaptive manner, and the combination of the two is used to enhance the network's ability to model spatial features in different ranges. The fused spatial features are then input into the Temporal Transformer Block to capture the temporal features of each joint. The modeling processes of spatial and temporal features are executed alternately. After several rounds of iteration of the structure, the final 3D pose sequence is obtained through the Regression Head.
[0010] A method for modeling spatial features of 3D human poses, the method comprising the following steps: Step 1: Use a pre-trained 2D human pose detector to process the input image or video frame, detect and generate the two-dimensional coordinates of each human joint, and generate a 2D skeleton. This process ensures accurate input data for subsequent 3D pose estimation.
[0011] Step 2: Divide the 2D skeleton into several main parts, including the right leg, left leg, head, right arm, and left arm. This division is based on the spatial and motion correlations between joints. For example, joints in the same part such as the knee and ankle, and the shoulder and elbow usually have a high correlation during movement. By taking these parts as objects to capture local spatial features, the movement characteristics of local joints can be better understood while reducing the complexity of global processing.
[0012] Step 3: Combine the joint points of each part with the skeleton origin to form multiple Origin-based Parts. This design, by associating local joints with a global reference point (Origin), can not only capture the movement relationships between joints within a part but also maintain the continuity and integrity of the overall human pose. The connection between the Origin and each part ensures that the global consistency of the human pose is not lost when estimating local poses.
[0013] Step 4: Add Spatial Positional Encoding to the 2D skeleton that has formed multiple Origin-based Parts. This encoding can represent the relative positions of each joint point in space. Then, input the processed data into the Spatial Transformer Encoder (STE) to capture the global spatial features of the entire skeleton. Through the self-attention mechanism, STE can effectively capture the long-range dependencies between various joints of the human body. This process helps the model understand the interactions between non-locally related joints, such as the hands and feet, the shoulders and knees, etc., thereby improving the prediction accuracy of the overall pose.
[0014] Step 5: In addition to capturing global spatial features, local spatial feature modeling is also performed on each Origin-based Part. Using STE, with each Origin-based Part as a unit, extract the spatial dependency features inside the joints through the self-attention mechanism to achieve local spatial feature modeling of each Origin-based Part. These local features can well capture the motion correlations inside the parts. Next, fuse the global spatial features and the local spatial features to generate a spatial feature containing comprehensive information, thereby improving the accuracy of 3D pose prediction.
[0015] The present invention adopts a dual-channel parallel design in the model structure. One channel is used to capture local spatial features, and the other channel is used to capture global spatial features, and these two types of features are fused to generate more accurate spatial features.
[0016] Step 6: After completing the modeling of the spatial features, next, add Temporal Positional Encoding to the fused spatial features to represent the motion trajectories of each joint point between multiple time frames. Through the self-attention mechanism, capture the dependencies of each joint point in the time dimension. Finally, through the regression module, convert the spatio-temporal features into the three-dimensional coordinates of the joint points to generate the 3D human pose estimation result.
[0017] Furthermore, the generated 2D joint points form the skeleton structure of the human body, including the head, shoulders, hips, knees, and ankles. The generated 2D joint point data usually contains the x and y coordinates of each joint point and the detection confidence. Define the Hip joint as the origin (Origin) of the human body because it has a high spatial correlation in the human body structure and is the reference point for the spatial positions of other joints.
[0018] Further, the fusion process is carried out in an adaptive fusion manner, including performing element-wise multiplication on local spatial features and global spatial features, and weighted averaging the two features with weights generated by linear transformation, so as to generate a spatial feature containing comprehensive information, ensuring that the information of both can be fully integrated.
[0019] Further, the fused spatial feature is further passed through a Temporal Transformer Encoder (TTE) to capture the temporal features of each joint in multiple time frames, that is, to capture the dependencies of each joint point in the time dimension, so as to generate a more accurate 3D pose estimation result; the TTE can effectively process video sequences, analyze the motion patterns of joints over time, and thus improve the model's understanding of dynamic poses.
[0020] Further, the image or video frame is based on a video captured by a monocular camera.
[0021] Further, the optimization objectives for the generated 3D human pose estimation result include a position loss function and a velocity loss function ; wherein, the position loss function is used to measure the accuracy of the estimated position, and the velocity loss function is used to measure the smoothness of joint movement; . Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0023] Figure 1 is a schematic flowchart of a 3D human pose spatial feature modeling method disclosed in an embodiment of the present invention. Detailed Embodiments
[0024] The following specific embodiments illustrate the implementation manners of the present application. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.
[0025] In addition, the technical features involved in different embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.
[0026] The overall network structure mainly consists of five components. The Origin-based Part TransformerBlock and the Spatial Transformer Block form a two-channel parallel structure for capturing spatial features of single-frame 2D poses in different size ranges. Among them, the proposed OPFormer Block is responsible for capturing the correlation between joints within a part and generating local spatial features. Also, the Spatial Transformer Block is used to capture the dependencies of all joints of the human body and generate global spatial features. An adaptive method is used to fuse the results of the two channels, and the combination of the two is used to enhance the network's ability to model spatial features in different ranges. The fused spatial features are then input into the Temporal TransformerBlock to capture the temporal features of each joint. The modeling processes of spatial and temporal features are executed alternately. After several rounds of iteration of the structure, the final 3D pose sequence is obtained through processing by the Regression Head.
[0027] As Figure 1 shown, a method for modeling 3D human pose spatial features is provided, including the following steps: Step 1: The model inputs a 2D skeleton sequence , where is the number of frames, is the number of joints, is the number of input channels. First, the input is mapped into a high-dimensional feature using the Linear Embedding layer, and then the learnable spatial position encoding is added. After the prepared features are input into the parallel channels, the Origin-basedPart Transformer Block is used to calculate the local spatial features , where is the number of channels of the embedded feature, , is the network depth.
[0028] Step 2: The Origin-based Part Transformer Block is responsible for capturing the local spatial features of single-frame 2D poses. First, the complete human skeleton is divided into 5 parts: right leg, left leg, head, right arm, leftarm. Specifically, the complete human body feature is split into 5 part features , where is the number of joints in the th part, . The above process is defined as: , based on the fact that all joints are closely related to the hip joint, the hip joint is defined as the Origin of the human skeleton. For each part, the Origin and the joints inside it are combined into an Origin-based Part. Specifically, the features corresponding to the Origin and are concatenated through Concat to obtain a new feature , . The above process is defined as: .
[0029] Step 3: For the features corresponding to each Origin-based Part , first use a Linear layer to map it to the corresponding query , key and value , and pass them into for calculation. This generates the spatial features of a specific part, . not only contains the spatial dependencies of the joints inside the part, but also contains the spatial position information of the part relative to the Origin. The above process is defined as: ; where represents the matrix of the linear mapping, represents the calculation process of the Multi-Head Self-Attention Layer for spatial modeling.
[0030] Step 4: Then, the obtained features are re-divided into Origin features and part features , . Finally, the calculated from different Origin-based Parts are re-concatenated into a feature , and then the calculated different part features together with the Origin features are re-combined into the complete human body spatial feature . The above process is defined as: ; where is a learnable linear transformation.
[0031] Step 5; Use the Spatial Transformer Block to calculate the global spatial features , where is the number of channels of the embedded feature, , is the network depth. After two parallel channels model spatial features in different ranges respectively, the outputs of the two channels are fused in an adaptive fusion manner to generate complete spatial features , . This process is defined as: ; where is element-wise production, and the calculation process of the adaptive fusion weight and is defined as: ; where is a learnable linear transformation.
[0032] Step Six: Finally, add the fused feature to the temporal position encoding , and input it into the TemporalTransformer Block to calculate the temporal features of joint motion . . Finally, apply a linear transformation Regression Head to to estimate the final 3D pose , where is the number of channels of the output feature.
[0033] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0034] Although the embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of this disclosure is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in this disclosure. Further, various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many elements described herein can be replaced by equivalent elements that appear after this disclosure.
Claims
1. A 3D human body pose spatial feature modeling method, characterized in that, The method includes the following steps: Step 1: Use a pre-trained 2D human pose detector to process the input image or video frame, detect and generate the 2D coordinates of each human joint, and generate a 2D skeleton; The 2D joint point data included in the 2D skeleton contains the x and y coordinates of each joint point and the detection confidence; and, define the Hip joint as the origin of the human body; Step 2: Divide the 2D skeleton into several main parts, including the right leg, left leg, head, right arm, and left arm; Complete human body features are split into 5 part features , where is the number of joints, is the number of joints of the th part, , and the above process is defined as: ; Step 3: Combine the joint points of each part with the skeleton origin to form multiple Origin-based Parts; The features corresponding to Origin and obtain new features through Concat , , and the above process is defined as: ; Step 4: Add spatial position encoding to the 2D skeleton that has formed multiple Origin-based Parts. This encoding can represent the relative positions of each joint point in space; input the processed data into the Spatial Transformer Encoder to capture the global spatial features of the entire 2D skeleton; Step 5: Use the STE to extract the spatial dependence features inside the joints for each Origin-based Part through the self-attention mechanism, realize the local spatial feature modeling of each Origin-based Part, and fuse the global spatial features with the local spatial features to generate a spatial feature containing comprehensive information; For each feature corresponding to the Origin-based Part , first use the Linear layer to map it to the corresponding query , key and value , and pass them into for calculation; generate the spatial features of a specific part , ; not only contains the spatial dependencies of the internal joints of the part, but also contains the spatial position information of the part relative to the Origin; the above process is defined as: ; Among them, represents the matrix of the linear mapping; Use the Spatial Transformer Block to calculate global spatial features , where is the number of channels of the embedded feature, , is the network depth; after two parallel channels model spatial features in different ranges respectively, the outputs of the two channels are fused in an adaptive fusion manner to generate complete spatial features , ; This process is defined as: ; wherein, is element-wise multiplication, and the calculation process of the adaptive fusion weight and is defined as: ; wherein, is a learnable linear transformation; Step 6: Add temporal position encoding to the spatial feature to represent the movement trajectories of each joint point between multiple time frames; capture the dependence relationships of each joint point in the time dimension through the self-attention mechanism; through the regression module, convert the spatio-temporal features into the 3D coordinates of the joint points to generate the 3D human pose estimation result; The fused features are added with temporal position encoding and fed into the Temporal Transformer Block to calculate the temporal features of joint motion , Finally, a linear transformation Regression Head is applied to estimate the final 3D pose , where is the number of channels of the output features; The image or video frame is a video captured based on a monocular camera; The optimization objectives for the generated 3D human pose estimation results include a position loss function and a velocity loss function ; among them, the position loss function is used to measure the accuracy of the estimated position, and the velocity loss function is used to measure the smoothness of joint movement; .
Citation Information
Patent Citations
Monocular three-dimensional human body posture estimation method and system fusing spatial-temporal characteristics
CN114581945A
Human posture estimation method and device fused with human intelligence manufacturing
CN118918614A