A 3D Human Pose Estimation Method, System, Electronic Device and Storage Medium
By adopting a global-local feature fusion network and pose optimization network with a two-branch structure in 3D human pose estimation, combined with 2D pose comparison learning and multimodal constraint optimization, the problems of local details loss and model uncertainty are solved, and the accuracy of 3D pose estimation is significantly improved.
Patent Information
- Application Number
- CN202510488584.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing 3D human pose estimation method based on monocular 2D videos has problems with local details and model uncertainty, resulting in inconsistent timing of the estimated 3D pose and insufficient local details.
The global-local feature fusion network (GLFF-Net) with a two-branch structure is adopted to extract global and local features through the pose-level timing interaction module and the joint-level spatial interaction module, and generate fusion features through an adaptive fusion strategy. At the same time, 2D pose contrast learning and multimodal constraint optimization network are used to generate 3D human poses, and a sparse-intensive training strategy is used to optimize the model.
The accuracy of 3D human posture estimation is significantly improved, the ability to restore local details is enhanced, and the model uncertainty caused by depth blur is reduced.
Smart Images

Figure CN120014713B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a 3D human pose estimation method, system, electronic device, and storage medium based on global-local feature fusion and pose optimization. Background Art
[0002] Estimating 3D human poses (3D-HPE) from monocular 2D videos is a classic and challenging task in the field of computer vision, and can be widely applied to various human motion-related scenarios such as virtual reality and human-computer interaction. The 3D-HPE task can be performed under multi-view or single-view settings. Due to the high deployment cost of multi-view methods and the difficulty in widely applying them in real scenarios, estimating 3D human poses from monocular videos has become the focus of research. Currently, the 3D human pose estimation methods based on monocular 2D videos are mainly divided into direct regression methods and lifting methods from 2D to 3D. The latter is superior in performance to the former because it uses advanced 2D pose estimators. To alleviate the model uncertainty caused by depth ambiguity, probability distribution-based methods use generative models to predict multiple 3D pose hypotheses; uncertainty methods use Transformers to learn temporal relationships to alleviate the depth ambiguity problem.
[0003] Although probability distribution-based methods can predict multiple 3D pose hypotheses, they need to aggregate multiple hypotheses into a single 3D pose, which is inefficient, and pre-determining the number of generated hypotheses limits the flexibility of designing hypothesis regression models. The deterministic methods based on Transformers have many deficiencies. Some models can only effectively capture long-range global temporal dependencies, ignoring the subtle temporal changes in local contexts, resulting in the loss of local details; some models only rely on one branch to learn global and local spatio-temporal relationships, causing local context information to be masked by long-term global information, and ultimately resulting in inconsistent 3D poses in terms of time series and insufficient local details. In addition, deterministic methods do not fully consider the uncertainty of the model, resulting in unsatisfactory evaluation results, such as inaccurate alignment of local joints. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a 3D human pose estimation method, system, electronic device, and storage medium, aiming to decouple global and local features through a dual-branch structure, combine 2D pose contrast learning and multi-modal constraint optimization, improve the ability to restore local details, and reduce the model uncertainty caused by depth ambiguity, thereby significantly improving the accuracy of 3D pose estimation.
[0005] To solve the above technical problems, the present application is implemented as follows:
[0006] In a first aspect, the embodiments of the present application provide a 3D human pose estimation method, and the method includes:
[0007] S1. Extract the 2D key point sequence from the monocular 2D video for pose space embedding to generate a high-dimensional feature representation for each frame;
[0008] S2. Apply a random masking operation and positional encoding to the high-dimensional feature representation to obtain a sparse temporal Token sequence;
[0009] S3. Use a pose-level temporal interaction module based on Transformer to extract global temporal features from the temporal Token sequence;
[0010] S4. For the current frame and its neighboring frames, obtain the joint-level Token sequence using joint-level spatial embedding, and adopt a joint-level spatial interaction module and hierarchical convolution to capture local temporal details to obtain local features;
[0011] S5. Fuse the global temporal features and the local features through an adaptive fusion strategy to generate fused features;
[0012] S6. Input the fused features into a pose optimization network, which includes a 2D pose contrastive learning module, a global 3D pose contrastive learning module, and a central frame 3D pose contrastive learning module, to generate 3D human poses;
[0013] S7. Calculate the overall loss based on the generated 3D human poses and the corresponding prior information, and optimize and train the pose optimization network using a sparse-dense training strategy;
[0014] S8. Output the 3D human poses corresponding to the monocular 2D video based on the optimized and trained pose optimization network.
[0015] As an alternative implementation of the first aspect of the present application, the pose-level temporal interaction module in step S3 uses Transformer as the backbone network to implement the multi-head self-attention mechanism for self-attention interaction of the temporal Token sequence.
[0016] As an alternative implementation of the first aspect of the present application, after step S3, it further includes: randomly initializing the temporal Token sequence to generate masked temporal Tokens, and splicing the masked temporal Tokens with the unmasked Tokens to obtain globally aligned temporal features after dimension alignment.
[0017] As an alternative implementation of the first aspect of the present application, the step S4 includes: performing joint-level spatial embedding on the current frame and local neighboring frames to construct a joint-level token sequence including a query matrix, a key matrix, and a value matrix; capturing the spatial relationship between joints through a joint-level spatial interaction module based on Transformer; using two-layer hierarchical convolution to capture local fine temporal changes in the spatial relationship between joints; the kernel size of the first-layer convolution is 3 to capture the correlation of local context; the kernel size of the second-layer convolution is 1 to enhance context excitation; the two-layer hierarchical convolution uses residual connection to fully learn detailed features.
[0018] As an alternative implementation of the first aspect of the present application, in the step S6: through a 2D pose contrast learning module, performing masked attention processing on the fused features to generate a predicted 2D pose, and calculating the contrast loss with the 2D prior information; through a global 3D pose contrast learning module, performing a 3D pose elevation operation to generate a global 3D pose sequence, and calculating the contrast loss between the estimated 3D pose and the true 3D pose; through a central frame 3D pose contrast learning module, using a Transformer with a stride to reduce the global information to the target frame, and using the global information to constrain the features of the target frame, and calculating the contrast loss between the predicted 3D pose of the central frame and the true 3D pose.
[0019] As an alternative implementation of the first aspect of the present application, in the step S7, the total loss function is defined as: , where , and are hyperparameters representing the weights of each type of loss function, represents the contrast loss between the predicted 2D pose and the 2D prior information, represents the contrast loss between the estimated 3D pose and the true 3D pose, represents the contrast loss between the predicted 3D pose of the central frame and the true 3D pose.
[0020] As an alternative implementation of the first aspect of the present application, in the step S7, the sparse-dense training strategy includes: in the sparse training stage, setting a non-zero random masking rate for sparse training and randomly mining the inter-frame consistency; in the dense training stage, setting a random masking rate of 0 for dense training to capture the global spatio-temporal representation.
[0021] In a second aspect, an embodiment of the present application provides a 3D human pose estimation system, and the system includes:
[0022] A data preprocessing module, which is used to extract a 2D key point sequence from a monocular 2D video to perform pose space embedding and generate a high-dimensional feature representation for each frame;
[0023] A global-local feature fusion module, which is used to perform a random masking operation and position encoding on the high-dimensional feature representation to obtain a sparse temporal Token sequence; use a pose-level temporal interaction module based on Transformer to extract global temporal features from the temporal Token sequence; for the current frame and its neighboring frames, obtain joint-level Token sequences by using joint-level spatial embedding, and use a joint-level spatial interaction module and hierarchical convolution to capture local temporal details to obtain local features; fuse the global temporal features and the local features through an adaptive fusion strategy to generate fused features;
[0024] A pose optimization module, which is used to input the fused features into a pose optimization network. The pose optimization network includes a 2D pose contrast learning module, a global 3D pose contrast learning module, and a central frame 3D pose contrast learning module to generate 3D human poses; calculate an overall loss based on the generated 3D human poses and corresponding prior information, and use a sparse-dense training strategy to optimize and train the pose optimization network;
[0025] An output module, which is used to output 3D human poses corresponding to the monocular 2D video based on the optimized and trained pose optimization network.
[0026] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0027] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0028] Compared with the prior art, the present invention proposes a 3D human pose estimation method: First, by extracting key points from a monocular 2D video and performing pose space embedding, a high-dimensional feature representation of each frame is obtained, which provides rich and accurate pose information for subsequent modeling. Then, random masks and position encoding are applied to the high-dimensional features to construct a sparse temporal Token sequence, effectively reducing redundancy and enhancing the model's ability to capture key information. Subsequently, a pose-level temporal interaction module based on Transformer is used to extract global temporal features, enabling the modeling of long-range dependencies and overall motion trends in the video. At the same time, local Tokens are obtained through joint-level spatial embedding of the current frame and its neighboring frames, and fine-grained temporal details are captured through a joint-level interaction module and hierarchical convolution, making up for the deficiency of global features in describing local motion details. The subsequent adaptive fusion strategy integrates the global temporal features and local features to generate fusion features that have both macroscopic semantics and microscopic details. Next, the fusion features are input into a pose optimization network containing 2D and multi-level 3D pose contrast learning modules, and the pose information is further refined and corrected through multi-angle supervision. Finally, the overall loss is calculated based on the generated 3D human pose and prior information, and a sparse-dense training strategy is used for optimization training, enabling the network to fully exploit and integrate sparse features and dense features during training, and finally output a 3D human pose that is highly consistent and accurate with the input 2D video. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a flowchart of a 3D human pose estimation method provided by the first embodiment of the present invention;
[0030] Figure 2 is a structural diagram of a global-local feature fusion and pose optimization network (GLFFPO-Net) proposed in the first embodiment of the present application;
[0031] Figure 3 is a schematic structural diagram of a 3D human pose estimation system provided by the second embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0033] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.
[0034] In order to illustrate the technical solutions described in this application, the following will be described through specific embodiments.
[0035] Embodiment 1
[0036] Please refer to Figure 1 , which is a flowchart of a 3D human pose estimation method proposed in the first embodiment of this application. Please refer to Figure 2 , which is a structural diagram of the Global-Local Feature Fusion and Pose Optimization Network (GLFFPO-Net) proposed in the first embodiment of this application. The steps of the proposed method are as follows.
[0037] S1. Extract the 2D key point sequence from the monocular 2D video to perform pose space embedding and generate a high-dimensional feature representation for each frame.
[0038] S2. Apply a random masking operation and positional encoding to the high-dimensional feature representation to obtain a sparse temporal Token sequence.
[0039] S3. Use a Transformer-based pose-level temporal interaction module to extract global temporal features from the temporal Token sequence.
[0040] S4. For the current frame and its neighboring frames, use joint-level space embedding to obtain a joint-level Token sequence, and adopt a joint-level space interaction module and hierarchical convolution to capture local temporal details and obtain local features.
[0041] It should be noted that existing 3D-HPE methods tend to learn long-range global temporal dependencies, but are not sensitive enough to subtle changes in local motion. This can lead to problems such as temporal inconsistency and inaccurate prediction of joint positions in 3D pose results. To solve this problem, this application proposes a Global-Local Feature Fusion Network (GLFF-Net). This network uses a dual-branch structure to decouple global and local spatio-temporal dependencies, and applies an adaptive fusion strategy to enhance the globality and locality of the fused features.
[0042] Regarding the global branch, the specific details of GLFF-Net are as follows:
[0043] 1. Temporal Token Sparse Modeling. First, use an existing 2D pose estimator to extract continuous 2D human key points from a monocular video, obtaining the coordinate information of each frame containing the main human joints (such as joints, head, limbs, etc.). Subsequently, in order to endow these 2D coordinates with stronger expressive power and discriminability, the original 2D key point data needs to be transformed into high-dimensional features through a series of preprocessing and embedding steps.
[0044] Then, use the random masked frame operation with a masking rate to discard a part of the frames, obtaining a sparse 2D pose representation , where denotes the feature dimension; denote the indices of the masked frames as the set q , , T , and is the number of input frames. Then, add the positional encoding to the sparse 2D pose representation to obtain the input sequence of the temporal interaction module, and then obtain the query matrix , key matrix and value matrix
[0045] through a separation operation. The sparsification process of temporal features can further explore the correlation between frames.
[0046] 2. Pose-level Temporal Interaction. In some embodiments, the pose-level temporal interaction module uses a Transformer as the backbone network to implement the multi-head self-attention mechanism for self-attention interaction on the temporal Token sequence. Optionally, randomly initialize the temporal Token sequence to generate masked temporal Tokens, and concatenate the masked temporal Tokens with the unmasked Tokens to obtain the globally aligned temporal features after dimension alignment. Specifically, design a pose-level temporal interaction module (Pose-level Temporal Interaction, PTI) with a Transformer as the backbone network, and the input temporal Token sequence captures the temporal relationship between frames , where
[0047]
[0048]
[0049] where represents the intermediate output, represents the output of the previous layer, represents a normalization layer. The above formula can use a function to represent:
[0050]
[0051] Multi-head Self-attention ( ) and Multi-layer Perception ( ) are important components of Transformer. In , scaled dot-product attention is single-head attention, and its calculation formula is as follows:
[0052]
[0053] Among them, represents the attention mechanism, represents the activation function, divides the query matrix , the key matrix and the value matrix into h heads, and then the formula (4) is executed in parallel to obtain :
[0054]
[0055]
[0056] Among them represents the concatenation operation, represents the h -th head in the multi-head attention, is the parameter matrix of the linear transformation.
[0057] consists of two linear transformations and an activation function. Taking the intermediate output of formula (1) as an example, is defined as:
[0058]
[0059] Among them represents the activation function, , , , are the parameter matrices of the linear transformation, is the dimension of the intermediate layer.
[0060] To ensure that the output dimension is consistent with the input dimension, a masked temporal token sequence is generated using random initialization and concatenated with the unmasked tokens to obtain a global representation .
[0061] For the local branch, the specific details of GLFF-Net are as follows:
[0062] In some embodiments, joint-level spatial embedding is performed on the current frame and local neighboring frames to construct a joint-level token sequence including a query matrix, a key matrix, and a value matrix; through a Transformer-based joint-level spatial interaction module, the joint-level token sequence is used to capture the spatial relationships between joints; two-layer hierarchical convolution is used to capture local fine temporal variations in the spatial relationships between joints; the kernel size of the first layer of convolution is 3 to capture the relevance of local context; the kernel size of the second layer of convolution is 1 to enhance context excitation; the two-layer hierarchical convolution uses residual connections to fully learn detailed features.
[0063] 1. Context neighborhood spatial token acquisition. First, focus on the current frame and its local neighboring frames in time. For these selected frame sets, joint-level spatial embedding is performed to generate high-dimensional feature representations for each joint point. Based on these high-dimensional joint features, the query matrix, key matrix, and value matrix are obtained by separating the features, preparing for the subsequent interaction module.
[0064] 2. Joint-level spatial interaction. These joint-level token sequences are input into the designed joint-level spatial interaction module (Joint-level Spatial Interaction, JSI). This module uses Transformer as the backbone network and aims to effectively capture the spatial relationships between joint points through components such as its built-in multi-head self-attention mechanism , represents the number of selected frames, represents the number of human joints, represents the joint feature dimension, where is the layer index of the JSI module. Similar to the PTI module, the JSI module can also use a function to represent:
[0065]
[0066] where represents the output of the previous layer.
[0067] 3. Local Context Temporal Feature Capture. After capturing the main joint space structure through the JSI module, in order to further model the local subtle temporal changes contained in these spatial relationships, a Hierarchical Convolution (HC) module is introduced. The HC module is designed to consist of two convolutional layers and is connected by an activation function. Specifically, the first convolutional layer uses a convolutional kernel of size 3 to capture the relevance of local context; subsequently, through the activation function (in this embodiment, the Sigmoid activation function, denoted as ), and then connected to the second convolutional layer, which uses a convolutional kernel of size 1 to enhance context excitation. In order to be able to fully learn detailed features and improve the gradient flow, residual connections are applied before and after the entire hierarchical convolution operation. After the serial processing of the JSI module and the HC module, the final local temporal representation , representing the local feature dimension, can be expressed as formula (9):
[0068]
[0069] where represents the convolutional operation, represents the GELU activation function, represents 's linear transformation.
[0070] S5. Fuse the global temporal feature and the local feature through an adaptive fusion strategy to generate a fused feature.
[0071] It should be noted that the global temporal feature has two characteristics: temporal consistency and globality, while the local temporal feature incorporates the local subtle changes in human joint movements. In order to simultaneously leverage the advantages of these two types of features, an Adaptive Fusion (AF) strategy is designed. Specifically, this strategy first integrates the global and local features using a function, and then adaptively adjusts the weights of the local components in the fused feature using a normalization layer and a linear layer to balance the global and local features, and finally obtains the fused feature , representing the fused feature dimension.
[0072] S6. Input the fused feature into the pose optimization network, which includes a 2D pose contrast learning module, a global 3D pose contrast learning module, and a central frame 3D pose contrast learning module, to generate a 3D human pose.
[0073] It should be noted that the 3D human pose estimation task predicts the corresponding 3D pose from 2D key points. However, due to the lack of depth information, this task becomes an inverse problem with multiple possible solutions, and thus the learned model has a considerable degree of uncertainty. Although 2D priors only contain 2D information, they effectively restrict the reasonable positions of human joints in 3D space (i.e., each joint lies on the ray extending from the camera optical center to the corresponding 2D key point). Therefore, the most direct way to solve this problem is to use the data of 2D and 3D structures to constrain the learning process of the model to ensure that the learned features can be mapped to the determined 2D or 3D priors. For this purpose, this application proposes a Pose Optimization Network (PO-Net).
[0074] In some embodiments, through the 2D pose contrast learning module, masked attention processing is performed on the fused features to generate the predicted 2D pose, and the contrast loss between the predicted 2D pose and the 2D prior information is calculated; through the global 3D pose contrast learning module, 3D pose elevation operations are performed to generate the global 3D pose sequence, and the contrast loss between the estimated 3D pose and the true 3D pose is calculated; through the central frame 3D pose contrast learning module, the global information is reduced to the target frame using a Transformer with a stride, the features of the target frame are constrained using the global information, and the contrast loss between the predicted 3D pose of the central frame and the true 3D pose is calculated.
[0075] It can be understood that the Pose Optimization Network (PO-Net) contains three branches, and these branches cooperate to constrain the learning process of the network by using 2D prior information and 3D pose structures. First, based on the fused features, a masked attention mechanism is designed to ensure that the features of each frame can communicate with the unmasked features, thereby enhancing the inter-frame correlation. Subsequently, these enhanced features are decoded to generate new 2D poses, and they are aligned with the 2D prior information through contrast learning. This method ensures that the features enhanced by masked attention can recover poses consistent with the 2D prior information. Based on the 3D pose structure, 3D poses are generated for the entire sequence and the central frame respectively, and contrast learning is performed with the real sample data to constrain the learning of the network from the perspective of 3D space. This method of using multiple structure data to constrain the network learning can, to a certain extent, reduce the uncertainty of the network to optimize the generation of 3D human poses.
[0076] 1. 2D Pose Contrastive Learning Module: Through observation, it is found that even minor changes in 2D joint positions and appearances can provide useful information in the absence of depth information. Therefore, this application attempts to utilize 2D prior information to constrain the learning process of the network. Meanwhile, in the research of GLFF-Net, in order to further explore the inter-frame correlation in human motion, this application randomly masks some frames and attempts to recover them through random initialization. However, this method cannot guarantee the accuracy of mask frame recovery. For this reason, this application proposes a 2D Pose Contrastive Learning (2D-PCL) method, which directly constrains the learning process of mask frame features while ensuring that the learned features can be aligned with 2D prior information, thereby reducing the uncertainty of the network.
[0077] In the previous feature fusion stage, random frames were appended after the unmasked frames. These random frames contained a large amount of noise, making the process of recovering the masked frames similar to the language translation process, with a focus on the impact of prior information on the current result. Inspired by the great progress achieved by Transformer in the field of natural language processing, this application then proposes to apply masked attention to recover the masked frames. This method can further explore the correlation between masked frames and unmasked frames. Specifically, for the fused features , first, positional encodings are added to the features of unmasked frames and masked frames respectively. Then, a Transformer based on masked attention (MT) is used to obtain enhanced features. The above process can be expressed as:
[0078]
[0079]
[0080] where represents the concatenation operation, represents the layer index of MT, represents the set of unmasked frames, represents the positional encoding of unmasked features, represents the positional encoding of masked features, , and the optimized features are obtained through linear transformation.
[0081] To be consistent with the data structure of prior information, a 2D pose decoder is used to generate 2D poses . The 2D pose decoder consists of a batch normalization layer (BN layer) and a 1D temporal convolutional layer:
[0082]
[0083] Finally, calculate the contrast loss between the predicted 2D pose and the 2D prior information. . During the previous process of aligning the global information dimension, the masked frames were directly appended after the unmasked frames. To ensure the one-to-one correspondence between the output result and the 2D prior information, the present application first rearranges the order of the 2D prior information. After that, the contrast loss is calculated. The above process is defined as follows:
[0084] (13)
[0085]
[0086] where represents the 2D prior information after adjusting the order, represents the 2D pose of the j-th frame, and respectively represent the position of the -th joint of the predicted -th frame and the position of the -th joint of the existing -th frame, represents the number of human joints.
[0087] 2. Global 3D Pose Contrast Learning Module: Since the ultimate goal is to predict 3D poses from 2D video frames, to ensure the accuracy of the generated 3D poses and the smoothness of their time series, the global 3D pose structure is used to further constrain the learning process of the network. Specifically, use 1D convolution to process to obtain , and then adopt a 3D pose elevation operation to generate a global 3D pose sequence . Calculate the loss between the estimated 3D pose and the true 3D pose:
[0088]
[0089] where and respectively represent the position of the -th joint of the estimated -th frame and the position of the -th joint of the true -th frame.
[0090] 3. Central Frame 3D Pose Contrast Learning Module: To ensure that the target frame has a specific representation, the present application introduces central frame 3D pose contrast learning. Specifically, the present application uses a Transformer with a stride to reduce the global information to the target frame, and uses the global information to constrain the features of the target frame , thus ensuring the accuracy of the 3D pose of the target frame. Estimate the 3D pose of the central frame through the pose enhancement module . Calculate the loss between the predicted 3D pose of the central frame and the true 3D pose :
[0091]
[0092] where and are the positions of the -th joint of the predicted central frame and the -th joint of the true central frame respectively.
[0093] S7. Based on the generated 3D human pose and the corresponding prior information, calculate the overall loss, and optimize and train the pose optimization network using the sparse-dense training strategy.
[0094] Specifically, train the proposed network model in an end-to-end manner, and the total loss function is defined as:
[0095]
[0096] where , and are hyperparameters representing the weights of each type of loss function, represents the contrast loss between the predicted 2D pose and the 2D prior information, represents the contrast loss between the estimated 3D pose and the true 3D pose, represents the contrast loss between the predicted 3D pose of the central frame and the true 3D pose.
[0097] Furthermore, in the sparse training stage, set the random mask rate to a non-zero value for sparse training to randomly mine the inter-frame consistency; in the dense training stage, set the random mask rate to 0 for dense training to capture the global spatio-temporal representation.
[0098] S8. Output the 3D human pose corresponding to the monocular 2D video based on the optimized and trained pose optimization network.
[0099] In summary, a 3D human pose estimation method based on global-local feature fusion and pose optimization network (GLFFPO-Net) proposed in this application includes:
[0100] First, the Global-Local Feature Fusion Network (GLFF-Net) adopts a dual-branch structure to learn global and local spatio-temporal representations respectively, so as to solve the problem of missing local details. To achieve this goal, this application designs a global pose-level temporal interaction attention mechanism to capture the global temporal dependencies between different poses. At the same time, in order to further explore the coherence and consistency in human motion, this application applies a random masking strategy to the entire input sequence. In addition, considering that both past and future frame information are important guiding factors for predicting the content of the current frame, this application designs a local joint-level spatial interaction attention mechanism specifically for local regions, enabling it to effectively capture the spatial dependencies between joints. In addition, hierarchical convolution is proposed to learn the local temporal dependencies in the context, thus highlighting the subtle temporal changes in local motion. Finally, an adaptive fusion strategy is adopted to dynamically adjust the weight of the local component in the global temporal features, effectively enhancing the globality and locality of the fused features.
[0101] Second, existing 2D key points are introduced as prior information to constrain the uncertainty of the model. Although the 2D prior information cannot directly represent depth information, they effectively limit the reasonable positions of human joints in 3D space (i.e., each joint lies on the ray extending from the camera optical center to the corresponding 2D key point). Therefore, the 2D prior information plays a role in reducing the uncertainty of the network model. Based on this, this application proposes a Pose Optimization Network (PO-Net), aiming to use 2D pose priors and 3D pose structures to constrain the learning process of the model, thus alleviating the uncertainty problem generated when learning 3D structures from 2D data. This application designs a masked attention mechanism to optimize the fused features, ensuring that the optimized features contain the information transmitted by each unmasked feature representation. Subsequently, contrastive learning is performed using 2D pose prior information to ensure that the learned features can be mapped to the corresponding determined 2D poses. The structural information of 3D poses is also utilized to further guide the learning process of the model.
[0102] Therefore, in the embodiments of this application, the main contributions include the following three aspects:
[0103] 1. A new method for 3D human pose estimation based on Global-Local Feature Fusion and Pose Optimization Network (GLFFPO-Net) is proposed, which is used to estimate 3D human poses from monocular videos and can effectively learn global and local spatio-temporal representations.
[0104] 2. A global-local feature fusion network (GLFF-Net) is proposed to adaptively fuse global and local temporal representations, thereby enhancing the global structural information and local details of pose features. In addition, a pose optimization network (PO-Net) is designed, which combines two-dimensional prior information to reduce the uncertainty of the model. This application considers the problem of model uncertainty in deterministic 3D-HPE methods for the first time.
[0105] 3. A sparse-dense training (SDT) method is proposed. In the sparse training stage, a random masking rate is set to a non-zero value for sparse training to randomly mine the inter-frame consistency. In the dense training stage, the random masking rate is set to 0 for dense training to capture the global spatio-temporal representation. This training method ensures the smoothness of the estimated 3D pose.
[0106] Embodiment 2
[0107] Please refer to Figure 3 , which shows a schematic structural diagram of a 3D human pose estimation system proposed in the second embodiment of this application. The system includes:
[0108] A data preprocessing module 100, which is used to extract a 2D key point sequence from a monocular 2D video for pose space embedding to generate a high-dimensional feature representation for each frame; a global-local feature fusion module 200, which is used to perform a random masking operation and position encoding on the high-dimensional feature representation to obtain a sparse temporal Token sequence; use a Transformer-based pose-level temporal interaction module to extract global temporal features from the temporal Token sequence; for the current frame and its neighboring frames, use joint-level spatial embedding to obtain a joint-level Token sequence, and use a joint-level spatial interaction module and hierarchical convolution to capture local temporal details to obtain local features; fuse the global temporal features and the local features through an adaptive fusion strategy to generate a fusion feature; a pose optimization module 300, which is used to input the fusion feature into a pose optimization network. The pose optimization network includes a 2D pose contrast learning module, a global 3D pose contrast learning module, and a central frame 3D pose contrast learning module to generate a 3D human pose; calculate the overall loss based on the generated 3D human pose and the corresponding prior information, and use a sparse-dense training strategy to optimize and train the pose optimization network;
[0109] An output module 400, which is used to output a 3D human pose corresponding to the monocular 2D video based on the optimized and trained pose optimization network.
[0110] A 3D human pose estimation system in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc. The embodiments of the present application do not make specific limitations.
[0111] A device of a 3D human pose estimation system in an embodiment of the present application. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0112] A 3D human pose estimation system provided by an embodiment of the present application can implement Figure 1 each process implemented by a 3D human pose estimation method in the method embodiment. To avoid repetition, it will not be elaborated here.
[0113] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned 3D human pose estimation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0114] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-mentioned 3D human pose estimation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0115] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc.
[0116] It should be noted that, in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0117] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0118] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
Claims
1. A 3D human body posture estimation method, characterized in that: The method comprises the following steps: S1, extracting 2D key point sequences from monocular 2D videos to embed posture space and generate high-dimensional feature representations for each frame; S2. Apply random masking operation and position encoding to the high-dimensional feature representation to obtain a sparse time-series Token sequence; S3, using a Transformer-based posture-level timing interaction module to extract global timing features from the timing Token sequence; S4. For the current frame and its neighboring frames, joint-level spatial embedding is used to obtain joint-level Token sequences, and joint-level spatial interaction modules and hierarchical convolutions are used to capture local temporal details and obtain local features. S5, fusing the global temporal features and the local features through an adaptive fusion strategy to generate a fusion feature; S6, inputting the fusion features into a posture optimization network, wherein the posture optimization network includes a 2D posture contrast learning module, a global 3D posture contrast learning module and a center frame 3D posture contrast learning module, so as to generate a 3D human body posture; S7, based on the generated 3D human body posture and the corresponding prior information, calculating the overall loss, and optimizing the posture optimization network using a sparse-dense training strategy; S8. Outputting a 3D human body posture corresponding to the monocular 2D video based on the posture optimization network after optimization training.
2. A 3D human body posture estimation method according to claim 1, characterized in that: The posture-level timing interaction module in step S3 uses Transformer as the backbone network to implement a multi-head self-attention mechanism to perform self-attention interaction on the timing Token sequence.
3. A 3D human body posture estimation method according to claim 1, characterized in that: After step S3, the following steps are also included: The timing Token sequence is randomly initialized to generate a masked timing Token, and the masked timing Token is concatenated with an unmasked Token to obtain a dimensionally aligned global timing feature.
4. A 3D human body posture estimation method according to claim 1, characterized in that: The step S4 comprises: Perform joint-level spatial embedding on the current frame and local neighboring frames, and construct a joint-level Token sequence including a query matrix, a key matrix, and a value matrix; The joint-level Token sequence is used to capture the spatial relationship between joints through a Transformer-based joint-level spatial interaction module. Two layers of layered convolution are used to capture the local subtle temporal changes in the spatial relationship between the joints; the kernel size of the first layer of convolution is 3 to capture the relevance of the local context; the kernel size of the second layer of convolution is 1 to enhance contextual excitation; the two layers of layered convolution use residual connections to fully learn detailed features.
5. A 3D human body posture estimation method according to claim 1, characterized in that: In step S6: Through the 2D pose contrast learning module, the fused features are subjected to mask attention processing to generate the predicted 2D pose and calculate the contrast loss with the 2D prior information; Through the global 3D pose contrast learning module, a 3D pose lifting operation is performed to generate a global 3D pose sequence, and the contrast loss between the estimated 3D pose and the true 3D pose is calculated; Through the center frame 3D pose contrast learning module, a Transformer with a step size is used to reduce the global information to the target frame, the global information is used to constrain the features of the target frame, and the contrast loss between the predicted 3D pose of the center frame and the true 3D pose is calculated.
6. A 3D human body posture estimation method according to claim 5, characterized in that: In step S7, the function of the overall loss is defined as: in, , and is a hyperparameter, representing the weight of each type of loss function, represents the contrast loss between the predicted 2D pose and the 2D prior information, represents the contrast loss between the estimated 3D pose and the true 3D pose, represents the contrast loss between the predicted 3D pose of the center frame and the ground-truth 3D pose.
7. A 3D human body posture estimation method according to claim 1, characterized in that: In step S7, the sparse-dense training strategy includes: In the sparse training stage, the random mask rate is set to a non-zero value for sparse training to randomly mine the consistency between frames; In the dense training stage, the random mask rate is set to 0 for dense training to capture the global spatiotemporal representation.
8. A 3D human body posture estimation system, characterized in that: The system comprises: The data preprocessing module is used to extract 2D key point sequences from monocular 2D videos for posture space embedding and generate high-dimensional feature representations for each frame; A global-local feature fusion module is used to apply random mask operations and position encoding to the high-dimensional feature representation to obtain a sparse temporal Token sequence; a Transformer-based posture-level temporal interaction module is used to extract global temporal features from the temporal Token sequence; for the current frame and its neighboring frames, a joint-level spatial embedding is used to obtain a joint-level Token sequence, and a joint-level spatial interaction module and hierarchical convolution are used to capture local temporal details to obtain local features; the global temporal features and the local features are fused through an adaptive fusion strategy to generate fused features; A posture optimization module is used to input the fusion features into a posture optimization network, wherein the posture optimization network includes a 2D posture contrast learning module, a global 3D posture contrast learning module and a center frame 3D posture contrast learning module, so as to generate a 3D human body posture; based on the generated 3D human body posture and the corresponding prior information, the overall loss is calculated, and the posture optimization network is optimized and trained using a sparse-dense training strategy; The output module is used to output a 3D human body posture corresponding to the monocular 2D video based on the posture optimization network after optimization training.
9. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a 3D human body posture estimation method as described in any one of claims 1 to 7 are implemented.
10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of a 3D human body posture estimation method as described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Method and device for constructing attitude estimation model and attitude estimation method
CN117114083A
Feature fusion 3D human body posture estimation method based on GCN and Transform
CN118212689A