3D human body posture estimation method and system, electronic equipment and storage medium
By adopting a global-local feature fusion network with a two-branch structure in 3D human pose estimation, local details loss and model uncertainty problems are solved, and the accuracy and consistency of 3D pose estimation are significantly improved.
Patent Information
- Application Number
- CN202510488584.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing 3D human pose estimation method based on monocular 2D videos has problems with local details and model uncertainty, resulting in inconsistent timing of the estimated 3D pose and insufficient local details.
The global-local feature fusion network (GLFF-Net) with a dual-branch structure is adopted to extract global and local timing features through the posture-level timing interaction module and the joint-level spatial interaction module, and generate fusion features through an adaptive fusion strategy. At the same time, 2D pose comparison learning and multimodal constraint optimization are introduced to reduce model uncertainty.
The accuracy of 3D human posture estimation is significantly improved, the ability to restore local details is enhanced, and model uncertainty caused by depth blur is reduced.
Smart Images

Figure CN120014713A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a 3D human body posture estimation method, system, electronic device and storage medium based on global-local feature fusion and posture optimization. Background Art
[0002] Estimating 3D human pose from monocular 2D video (3D-HPE) is a classic and challenging task in the field of computer vision. It can be widely used in various scenarios related to human motion, such as virtual reality and human-computer interaction. The 3D-HPE task can be performed in multi-view or single-view settings. Since the multi-view method has high deployment cost and is difficult to be widely used in real scenes, estimating 3D human pose from monocular video has become a research focus. At present, the 3D human pose estimation methods based on monocular 2D video are mainly divided into direct regression methods and 2D to 3D lifting methods. The latter has better performance than the former due to the use of advanced 2D pose estimators. To alleviate the model uncertainty caused by depth ambiguity, the probability distribution-based method uses a generative model to predict multiple 3D pose hypotheses; the uncertainty method uses Transformer to learn temporal relationships to alleviate the depth ambiguity problem.
[0003] Although the probability distribution-based method can predict multiple 3D pose hypotheses, it needs to aggregate multiple hypotheses into a unique 3D pose, which is inefficient. In addition, the number of generated hypotheses is predetermined, which limits the flexibility of designing hypothesis regression models. Transformer-based deterministic methods have many shortcomings. Some models can only effectively capture long-distance global temporal dependencies, ignoring subtle temporal changes in local contexts, resulting in loss of local details; some models rely on only one branch to learn the global and local spatiotemporal relationship, so that local context information is masked by long-term global information, ultimately resulting in the estimated 3D pose being inconsistent in time and lacking local details. In addition, deterministic methods fail to fully consider the uncertainty of the model, resulting in unsatisfactory evaluation results, such as inaccurate alignment of local joints. Summary of the invention
[0004] The purpose of the embodiments of the present application is to provide a 3D human posture estimation method, system, electronic device and storage medium, which aims to decouple global and local features through a dual-branch structure, combine 2D posture contrast learning and multimodal constraint optimization, improve the ability to restore local details, and reduce the model uncertainty caused by depth blur, thereby significantly improving the accuracy of 3D posture estimation.
[0005] In order to solve the above technical problems, this application is implemented as follows: In a first aspect, an embodiment of the present application provides a 3D human body posture estimation method, the method comprising: S1, extracting 2D key point sequences from monocular 2D videos to embed posture space and generate high-dimensional feature representations for each frame; S2. Applying random masking operation and position encoding to the high-dimensional feature representation to obtain a sparse time-series Token sequence; S3, using a Transformer-based posture-level timing interaction module to extract global timing features from the timing Token sequence; S4. For the current frame and its neighboring frames, joint-level spatial embedding is used to obtain joint-level Token sequences, and joint-level spatial interaction modules and hierarchical convolutions are used to capture local temporal details and obtain local features. S5, fusing the global temporal features and the local features through an adaptive fusion strategy to generate a fusion feature; S6, inputting the fusion features into a posture optimization network, wherein the posture optimization network includes a 2D posture contrast learning module, a global 3D posture contrast learning module and a center frame 3D posture contrast learning module, so as to generate a 3D human body posture; S7, based on the generated 3D human body posture and the corresponding prior information, calculating the overall loss, and optimizing the posture optimization network using a sparse-dense training strategy; S8. Outputting a 3D human body posture corresponding to the monocular 2D video based on the posture optimization network after optimization training.
[0006] As an optional implementation of the first aspect of the present application, the posture-level timing interaction module in step S3 uses Transformer as the backbone network to implement a multi-head self-attention mechanism to perform self-attention interaction on the timing Token sequence.
[0007] As an optional implementation of the first aspect of the present application, step S3 also includes: randomly initializing the timing Token sequence, generating a masked timing Token, splicing the masked timing Token with an unmasked Token, and obtaining a global timing feature after dimension alignment.
[0008] As an optional implementation of the first aspect of the present application, step S4 includes: performing joint-level spatial embedding on the current frame and local neighboring frames, constructing a joint-level Token sequence including a query matrix, a key matrix and a value matrix; capturing the spatial relationship between joints for the joint-level Token sequence through a Transformer-based joint-level spatial interaction module; using two layers of hierarchical convolution to capture local subtle temporal changes in the spatial relationship between the joints; the kernel size of the first layer of convolution is 3 to capture the relevance of local context; the kernel size of the second layer of convolution is 1 to enhance contextual excitation; the two layers of hierarchical convolution use residual connections to fully learn detailed features.
[0009] As an optional implementation of the first aspect of the present application, in step S6: through the 2D pose contrast learning module, mask attention processing is performed on the fused features to generate a predicted 2D pose, and the contrast loss between the 2D prior information is calculated; through the global 3D pose contrast learning module, a 3D pose lifting operation is performed to generate a global 3D pose sequence, and the contrast loss between the estimated 3D pose and the true 3D pose is calculated; through the center frame 3D pose contrast learning module, a Transformer with a step size is used to reduce the global information to the target frame, the global information is used to constrain the features of the target frame, and the contrast loss between the predicted 3D pose of the center frame and the true 3D pose is calculated.
[0010] As an optional implementation of the first aspect of the present application, in step S7, the total loss function is defined as: ,in, , and is a hyperparameter, representing the weight of each type of loss function, represents the contrast loss between the predicted 2D pose and the 2D prior information, represents the contrast loss between the estimated 3D pose and the true 3D pose, represents the contrast loss between the predicted 3D pose of the center frame and the ground-truth 3D pose.
[0011] As an optional implementation of the first aspect of the present application, in step S7, the sparse-dense training strategy includes: in the sparse training stage, setting the random mask rate to a non-zero value for sparse training and randomly mining the consistency between frames; in the dense training stage, setting the random mask rate to 0 for dense training to capture the global spatiotemporal representation.
[0012] In a second aspect, an embodiment of the present application provides a 3D human body posture estimation system, the system comprising: The data preprocessing module is used to extract 2D key point sequences from monocular 2D videos for posture space embedding and generate high-dimensional feature representations for each frame; A global-local feature fusion module is used to apply random mask operations and position encoding to the high-dimensional feature representation to obtain a sparse temporal Token sequence; a Transformer-based posture-level temporal interaction module is used to extract global temporal features from the temporal Token sequence; for the current frame and its neighboring frames, a joint-level spatial embedding is used to obtain a joint-level Token sequence, and a joint-level spatial interaction module and hierarchical convolution are used to capture local temporal details to obtain local features; the global temporal features and the local features are fused through an adaptive fusion strategy to generate fused features; A posture optimization module is used to input the fusion features into a posture optimization network, wherein the posture optimization network includes a 2D posture contrast learning module, a global 3D posture contrast learning module and a center frame 3D posture contrast learning module, so as to generate a 3D human body posture; based on the generated 3D human body posture and the corresponding prior information, the overall loss is calculated, and the posture optimization network is optimized and trained using a sparse-dense training strategy; The output module is used to output a 3D human body posture corresponding to the monocular 2D video based on the posture optimization network after optimization training.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.
[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0015] Compared with the prior art, the present invention proposes a 3D human posture estimation method: first, by extracting key points from monocular 2D videos and embedding them in posture space, a high-dimensional feature representation of each frame is obtained, which provides rich and accurate posture information for subsequent modeling. Then, random masks and position encoding are applied to high-dimensional features to construct a sparse temporal Token sequence, which effectively reduces redundancy and enhances the model's ability to capture key information. Subsequently, a Transformer-based posture-level temporal interaction module is used to extract global temporal features, so that long-range dependencies and overall motion trends in the video can be modeled. At the same time, local tokens are obtained for the current frame and its neighboring frames through joint-level spatial embedding, and fine-grained temporal details are captured through joint-level interaction modules and hierarchical convolutions, making up for the shortcomings of global features in describing local motion details. The subsequent adaptive fusion strategy integrates global temporal features with local features to generate fused features that have both macro semantics and micro details. Next, the fused features are input into a posture optimization network containing 2D and multi-level 3D posture contrast learning modules, and the posture information is further refined and corrected through multi-angle supervision. Finally, the overall loss is calculated based on the generated 3D human posture and prior information, and the sparse-dense training strategy is used for optimization training, so that the network can fully mine and fuse sparse features and dense features during the training process, and finally output a 3D human posture that is highly consistent and accurate with the input 2D video. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flow chart of a 3D human body posture estimation method provided by the first embodiment of the present invention; Figure 2 It is a structural diagram of the global-local feature fusion and posture optimization network (GLFFPO-Net) proposed in the first embodiment of the present application; Figure 3 It is a structural schematic diagram of a 3D human body posture estimation system provided by the second embodiment of the present invention. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0018] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, the "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.
[0019] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.
[0020] Example 1 See also Figure 1 , which is a flow chart of a 3D human body posture estimation method proposed in the first embodiment of the present application, please refer to Figure 2 , is a structural diagram of the global-local feature fusion and posture optimization network (GLFFPO-Net) proposed in the first embodiment of the present application. The steps of the proposed method are as follows.
[0021] S1. Extract 2D key point sequences from monocular 2D videos to embed them into the pose space and generate high-dimensional feature representations for each frame.
[0022] S2. Apply random masking operations and position encoding to the high-dimensional feature representation to obtain a sparse temporal Token sequence.
[0023] S3. Use the Transformer-based posture-level temporal interaction module to extract global temporal features from the temporal token sequence.
[0024] S4. For the current frame and its neighboring frames, joint-level spatial embedding is used to obtain the joint-level Token sequence, and the joint-level spatial interaction module and hierarchical convolution are used to capture local temporal details and obtain local features.
[0025] It should be noted that existing 3D-HPE methods tend to learn long-distance global temporal dependencies, but are not sensitive enough to subtle changes in local motion. This can lead to problems such as temporal inconsistency and inaccurate prediction of joint positions in 3D pose results. To solve this problem, this application proposes a global-local feature fusion network (GLFF-Net). The network adopts a dual-branch structure to decouple global and local spatial-temporal dependencies, and uses an adaptive fusion strategy to enhance the globality and locality of the fused features.
[0026] For the global branch, the specific details of GLFF-Net are as follows: 1. Time-series Token Sparse Modeling. First, use the existing 2D posture estimator to extract continuous 2D human key points from the monocular video to obtain the coordinate information of the main joints of the human body (such as joints, head, limbs, etc.) in each frame. Then, in order to make these two-dimensional coordinates more expressive and discriminative, the original 2D key point data needs to be converted into high-dimensional features through a series of preprocessing and embedding steps.
[0027] Then use the random mask frame operation to mask the rate Discard some frames to get a sparse 2D pose representation , Represents the feature dimension; the index of the mask frame is recorded as a set , q , , T is the number of input frames. Then, we give the sparse 2D pose representation Add positional encoding Get the input sequence of the timing interaction module , and then obtain the query matrix through separation operation , key matrix Sum Matrix The sparse processing of temporal features can further mine the correlation between frames.
[0028] 2. Posture-level temporal interaction. In some embodiments, the posture-level temporal interaction module uses Transformer as the backbone network to implement a multi-head self-attention mechanism and perform self-attention interaction on the temporal token sequence. Optionally, the temporal token sequence is randomly initialized to generate a masked temporal token, and the masked temporal token is concatenated with the unmasked token to obtain the global temporal features after dimension alignment.
[0029] Specifically, we design a pose-level temporal interaction module (PTI), using Transformer as the backbone network and inputting a temporal token sequence to capture the temporal relationship between frames. ,in The process can be expressed by the following formula: in Represents the intermediate output, represents the output of the previous layer, represents the normalization layer. The above formula can be used with a function express: Multi-head Self-attention ) and Multi-layer Perception ) is an important component of Transformer. In , the scaled dot product attention is a single-head attention, and its calculation formula is as follows: in, represents the attention mechanism, represents the activation function, The query matrix , key matrix Sum Matrix Divide h Then, executing formula (4) in parallel, we can get : in Indicates a connection operation. Indicates the first h Head, is the parameter matrix of the linear transformation.
[0030] It consists of two linear transformations and an activation function. For example, is defined as: in express Activation function, , , , is the parameter matrix of the linear transformation, It is the dimension of the middle layer.
[0031] To ensure that the output dimension is consistent with the input dimension, a masked time sequence Token sequence is generated using random initialization , and concatenate it with the unmasked token to obtain a global representation .
[0032] For local branches, the specific details of GLFF-Net are as follows: In some embodiments, joint-level spatial embedding is performed on the current frame and local neighboring frames to construct a joint-level Token sequence including a query matrix, a key matrix, and a value matrix; the spatial relationship between joints is captured in the joint-level Token sequence through a Transformer-based joint-level spatial interaction module; two layers of hierarchical convolution are used to capture local subtle temporal changes in the spatial relationship between joints; the kernel size of the first layer of convolution is 3 to capture the relevance of local context; the kernel size of the second layer of convolution is 1 to enhance contextual excitation; the two layers of hierarchical convolution use residual connections to fully learn detailed features.
[0033] 1. Obtaining the contextual neighborhood space token. First, focus on the current frame and its local neighboring frames in time. Perform joint-level spatial embedding on these selected frame sets to generate high-dimensional feature representations for each joint point. Based on these high-dimensional joint features, query matrix, key matrix and value matrix are obtained by feature separation to prepare for the subsequent interaction module.
[0034] 2. Joint-level spatial interaction. These joint-level token sequences are input into the designed joint-level spatial interaction module (JSI). This module uses Transformer as the backbone network, aiming to effectively capture the spatial relationship between joint points through its built-in multi-head self-attention mechanism and other components. , Indicates the number of selected frames, Represents the number of human joints, represents the joint feature dimension, where is the layer index of the JSI module. Similar to the PTI module, the JSI module can also use a function express: in Represents the output of the previous layer.
[0035] 3. Capture of local contextual temporal features. After capturing the main joint spatial structure through the JSI module, a hierarchical convolution (HC) module is introduced to further model the local subtle temporal changes contained in these spatial relationships. The HC module is designed to contain two convolution layers connected by an activation function. Specifically, the first convolution layer uses a convolution kernel of size 3 to capture the relevance of the local context; then the activation function (in this embodiment, the Sigmoid activation function, expressed as ), and then connected to the second layer of convolution, which uses a convolution kernel of size 1 to enhance contextual excitation. In order to fully learn detailed features and improve gradient flow, residual connections are applied before and after the entire layered convolution operation. After the cascade processing of the JSI module and the HC module, the final local temporal representation is obtained , represents the local feature dimension, which can be expressed as formula (9): in represents the convolution operation, represents the GELU activation function, express Linear conversion.
[0036] S5. The global temporal features and local features are fused through an adaptive fusion strategy to generate fused features.
[0037] It should be noted that global temporal features have two characteristics: temporal consistency and globality, while local temporal features integrate local subtle changes in human joint movements. In order to give full play to the advantages of these two types of features, an adaptive fusion (AF) strategy is designed. This strategy first uses a function to integrate global and local features, and then uses a normalization layer and a linear layer to adaptively adjust the weights of local components in the fusion feature to balance the global and local features, and finally obtain the fusion feature. , Represents the fusion feature dimension.
[0038] S6. Input the fused features into a posture optimization network, which includes a 2D posture contrast learning module, a global 3D posture contrast learning module and a center frame 3D posture contrast learning module to generate a 3D human body posture.
[0039] It should be noted that the task of 3D human pose estimation is to predict the corresponding 3D pose from 2D key points. However, due to the lack of depth information, this task becomes an inverse problem with multiple possible solutions, so that the learned model has considerable uncertainty. Although the 2D priors only contain 2D information, they effectively restrict the reasonable positions of human joints in 3D space (i.e., each joint is located on a ray extending from the optical center of the camera to the corresponding 2D key point). Therefore, the most direct way to solve this problem is to use the data of 2D and 3D structures to constrain the learning process of the model to ensure that the learned features can be mapped to a certain 2D or 3D prior. To this end, this application proposes a pose optimization network (PO-Net).
[0040] In some embodiments, a 2D pose contrast learning module is used to perform masked attention processing on the fused features to generate a predicted 2D pose, and the contrast loss with the 2D prior information is calculated; a global 3D pose contrast learning module is used to perform a 3D pose lifting operation to generate a global 3D pose sequence, and the contrast loss between the estimated 3D pose and the true 3D pose is calculated; a central frame 3D pose contrast learning module is used to reduce the global information to the target frame using a Transformer with a step size, and the features of the target frame are constrained using the global information, and the contrast loss between the predicted 3D pose of the central frame and the true 3D pose is calculated.
[0041] Understandably, the pose optimization network (PO-Net) contains three branches that synergistically constrain the network's learning process by leveraging 2D prior information and 3D pose structure. First, based on the fused features, a masked attention mechanism is designed to ensure that the features of each frame can communicate with the unmasked features, thereby enhancing the inter-frame correlation. Subsequently, these enhanced features are decoded to generate new 2D poses, which are aligned with the 2D prior information through contrastive learning. This approach ensures that the features enhanced by masked attention can recover poses consistent with the 2D prior information. Based on the 3D pose structure, 3D poses are generated for the entire sequence and the center frame respectively, and contrastive learning is performed with the real sample data to constrain the network's learning from the perspective of three-dimensional space. This method of using multiple structural data to constrain network learning can alleviate the uncertainty of the network to a certain extent to optimize the generation of 3D human poses.
[0042] 1. 2D Pose Contrastive Learning Module: It has been observed that even small changes in 2D joint position and appearance can provide useful information in the absence of depth information. Therefore, this application attempts to use 2D prior information to constrain the learning process of the network. At the same time, in the study of GLFF-Net, in order to further explore the inter-frame correlation in human motion, this application randomly masked some frames and tried to restore them through random initialization. However, this method cannot guarantee the accuracy of mask frame recovery. To this end, this application proposes a 2D Pose Contrastive Learning (2D-PCL) method, which directly constrains the learning process of mask frame features while ensuring that the learned features can be aligned with the 2D prior information, thereby reducing the uncertainty of the network.
[0043] In the previous feature fusion stage, random frames are appended after the unmasked frames. These random frames contain a lot of noise, making the process of recovering the masked frames similar to the language translation process, focusing on the impact of previous information on the current result. Inspired by the great progress made by Transformer in the field of natural language processing, this application proposes to apply mask attention to recover the masked frames. This method can further explore the correlation between masked frames and unmasked frames. Specifically, for the fused feature , first add position encoding to the features of the unmasked frame and the masked frame respectively. Then, the mask attention-based Transformer (MT) is used to obtain enhanced features. The above process can be expressed as: in Indicates a connection operation. Indicates the layer index of MT, represents the set of unmasked frames, represents the positional encoding of the unmasked features, represents the positional encoding of the mask feature, , After linear transformation, the optimized features are obtained .
[0044] In order to keep consistent with the data structure of the prior information, the 2D pose decoder is used to generate the 2D pose The 2D pose decoder consists of a batch normalization layer (BN layer) and a 1D temporal convolution layer: Finally, the contrastive loss between the predicted 2D pose and the 2D prior information is calculated . Previously, in the process of global information dimension alignment, the mask frame was directly attached to the unmasked frame. In order to ensure a one-to-one correspondence between the output result and the 2D prior information, this application first rearranges the order of the two-dimensional prior information. After that, the contrast loss is calculated. The above process is defined as follows: (13) in, Represents the 2D prior information after adjusting the order, represents the 2D pose of the jth frame, and Respectively represent the predicted Frame No. The position of the joint and the existing Frame No. The position of the joints, Represents the number of joints in the human body.
[0045] 2. Global 3D pose contrast learning module: Since the ultimate goal is to predict 3D poses from 2D video frames, in order to ensure the accuracy of the generated 3D poses and the smoothness of their timing, the global 3D pose structure is used to further constrain the network learning process. Specifically, 1D convolution is used to process get , and then use the 3D posture lifting operation to generate a global 3D posture sequence . Calculate the loss between the estimated 3D pose and the true 3D pose : in and Respectively represent the estimated Frame No. The position of the joint and the true Frame No. The position of the joints.
[0046] 3. Center frame 3D pose contrast learning module: In order to ensure that the target frame has a specific representation, this application introduces center frame 3D pose contrast learning. Specifically, this application uses a Transformer with a step size to reduce the global information to the target frame, and uses the global information to constrain the features of the target frame. , thereby ensuring the accuracy of the 3D pose of the target frame. The 3D pose of the center frame is estimated through the pose lifting module . Calculate the loss between the predicted 3D pose of the center frame and the true 3D pose : in and They are the predicted center frame The position of the joint and the true center frame The position of the joints.
[0047] S7. Based on the generated 3D human posture and the corresponding prior information, the overall loss is calculated, and the sparse-dense training strategy is used to optimize the posture optimization network.
[0048] Specifically, the proposed network model is trained in an end-to-end manner, and the total loss function is defined as: in, , and is a hyperparameter, representing the weight of each type of loss function, represents the contrast loss between the predicted 2D pose and the 2D prior information, represents the contrast loss between the estimated 3D pose and the true 3D pose, represents the contrast loss between the predicted 3D pose of the center frame and the ground-truth 3D pose.
[0049] Furthermore, in the sparse training stage, the random mask rate is set to a non-zero value for sparse training to randomly mine the consistency between frames; in the dense training stage, the random mask rate is set to 0 for dense training to capture the global spatiotemporal representation.
[0050] S8. Based on the posture optimization network after optimization training, the network outputs a 3D human posture corresponding to the monocular 2D video.
[0051] In summary, the present application proposes a 3D human pose estimation method based on a global-local feature fusion and pose optimization network (GLFFPO-Net), which includes: First, the global-local feature fusion network (GLFF-Net) adopts a dual-branch structure to learn global and local spatiotemporal representations respectively to solve the problem of local detail loss. To achieve this goal, this application designs a global pose-level temporal interaction attention mechanism to capture the global temporal dependencies between different poses. At the same time, in order to further explore the coherence in human motion, this application applies a random masking strategy to the entire input sequence. In addition, given that both past and future frame information are important guiding factors for predicting the content of the current frame, this application designs a local joint-level spatial interaction attention mechanism specifically for local regions, so that it can effectively capture the spatial dependencies between joints. In addition, hierarchical convolution is proposed to learn local temporal dependencies in context, thereby highlighting subtle temporal changes in local motion. Finally, an adaptive fusion strategy is adopted to dynamically adjust the weights of local components in the global temporal features, effectively enhancing the globality and locality of the fused features.
[0052] Secondly, the existing two-dimensional key points are introduced as prior information to constrain the uncertainty of the model. Although the two-dimensional prior information cannot directly represent the depth information, they effectively restrict the reasonable positions of the human body joints in the three-dimensional space (that is, each joint is located on the ray extending from the optical center of the camera to the corresponding two-dimensional key point). Therefore, the two-dimensional prior information plays a role in reducing the uncertainty of the network model. Based on this, the present application proposes a posture optimization network (PO-Net), which aims to use the two-dimensional posture prior and the three-dimensional posture structure to constrain the learning process of the model, thereby alleviating the uncertainty problem generated when learning the three-dimensional structure from the two-dimensional data. The present application designs a masked attention mechanism to optimize the fused features to ensure that the optimized features contain the information conveyed by each unmasked feature representation. Subsequently, the two-dimensional posture prior information is used for comparative learning to ensure that the learned features can be mapped to the corresponding determined two-dimensional posture. The structural information of the three-dimensional posture is also used to further guide the learning process of the model.
[0053] Therefore, in the embodiments of the present application, the main contributions include the following three aspects: 1. A new method for 3D human pose estimation based on global-local feature fusion and pose optimization network (GLFFPO-Net) is proposed for estimating 3D human pose from monocular video, which can effectively learn global and local spatiotemporal representations.
[0054] 2. A global-local feature fusion network (GLFF-Net) is proposed to adaptively fuse global and local temporal representations to enhance the global structural information and local details of pose features. In addition, a pose optimization network (PO-Net) is designed, which combines two-dimensional prior information to reduce model uncertainty. This application considers the model uncertainty problem in the deterministic 3D-HPE method for the first time.
[0055] 3. A sparse-dense training (SDT) method is proposed. In the sparse training stage, the random mask rate is set to a non-zero value for sparse training, and the consistency between frames is randomly mined. In the dense training stage, the random mask rate is set to 0, and dense training is performed to capture the global spatiotemporal representation. This training method ensures the smoothness of the estimated 3D pose.
[0056] Example 2 See also Figure 3 , which is a schematic diagram of the structure of a 3D human body posture estimation system proposed in the second embodiment of the present application, wherein the system comprises: The data preprocessing module 100 is used to extract a 2D key point sequence from a monocular 2D video to perform posture space embedding and generate a high-dimensional feature representation for each frame; the global-local feature fusion module 200 is used to apply random mask operations and position encoding to the high-dimensional feature representation to obtain a sparse temporal Token sequence; the Transformer-based posture-level temporal interaction module is used to extract global temporal features from the temporal Token sequence; for the current frame and its neighboring frames, the joint-level spatial embedding is used to obtain a joint-level Token sequence, and the joint-level spatial interaction module and the layered convolution are used to capture local temporal details to obtain local features; the global temporal features and the local features are fused through an adaptive fusion strategy to generate fused features; the posture optimization module 300 is used to input the fused features into a posture optimization network, the posture optimization network includes a 2D posture contrast learning module, a global 3D posture contrast learning module and a center frame 3D posture contrast learning module to generate a 3D human posture; based on the generated 3D human posture and the corresponding prior information, the overall loss is calculated, and the posture optimization network is optimized and trained using a sparse-dense training strategy; The output module 400 is used to output the 3D human body posture corresponding to the monocular 2D video based on the posture optimization network after optimization training.
[0057] A 3D human posture estimation system in an embodiment of the present application may be a device, or a component, integrated circuit, or chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a network attached storage (NAS), a personal computer (PC), etc., which is not specifically limited in the embodiment of the present application.
[0058] A device of a 3D human posture estimation system in an embodiment of the present application. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0059] The 3D human body posture estimation system provided in the embodiment of the present application can achieve Figure 1In order to avoid repetition, each process of implementing a 3D human body posture estimation method in a method embodiment will not be described here.
[0060] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, each process of the above-mentioned 3D human body posture estimation method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0061] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the embodiment of the above-mentioned 3D human body posture estimation method are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0062] The processor is a processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0063] It should be noted that, in this article, the term "comprises", "includes" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0064] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0065] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A 3D human body posture estimation method, characterized in that: The method comprises the following steps: S1, extracting 2D key point sequences from monocular 2D videos to embed posture space and generate high-dimensional feature representations for each frame; S2. Applying random masking operation and position encoding to the high-dimensional feature representation to obtain a sparse time-series Token sequence; S3, using a Transformer-based posture-level timing interaction module to extract global timing features from the timing Token sequence; S4. For the current frame and its neighboring frames, joint-level spatial embedding is used to obtain joint-level Token sequences, and joint-level spatial interaction modules and hierarchical convolutions are used to capture local temporal details and obtain local features. S5, fusing the global temporal features and the local features through an adaptive fusion strategy to generate a fusion feature; S6, inputting the fusion features into a posture optimization network, wherein the posture optimization network includes a 2D posture contrast learning module, a global 3D posture contrast learning module and a center frame 3D posture contrast learning module, so as to generate a 3D human body posture; S7, based on the generated 3D human body posture and the corresponding prior information, calculating the overall loss, and optimizing the posture optimization network using a sparse-dense training strategy; S8. Outputting a 3D human body posture corresponding to the monocular 2D video based on the posture optimization network after optimization training.
2. A 3D human body posture estimation method according to claim 1, characterized in that: The posture-level timing interaction module in step S3 uses Transformer as the backbone network to implement a multi-head self-attention mechanism to perform self-attention interaction on the timing Token sequence.
3. A 3D human body posture estimation method according to claim 1, characterized in that: After step S3, the following steps are also included: The timing Token sequence is randomly initialized to generate a masked timing Token, and the masked timing Token is concatenated with an unmasked Token to obtain a dimensionally aligned global timing feature.
4. A 3D human body posture estimation method according to claim 1, characterized in that: The step S4 comprises: Perform joint-level spatial embedding on the current frame and local neighboring frames, and construct a joint-level Token sequence including a query matrix, a key matrix, and a value matrix; The joint-level Token sequence is used to capture the spatial relationship between joints through a Transformer-based joint-level spatial interaction module. Two layers of layered convolution are used to capture the local subtle temporal changes in the spatial relationship between the joints; the kernel size of the first layer of convolution is 3 to capture the relevance of the local context; the kernel size of the second layer of convolution is 1 to enhance contextual excitation; the two layers of layered convolution use residual connections to fully learn detailed features.
5. A 3D human body posture estimation method according to claim 1, characterized in that: In step S6: Through the 2D pose contrast learning module, the fused features are subjected to mask attention processing to generate the predicted 2D pose and calculate the contrast loss with the 2D prior information; Through the global 3D pose contrast learning module, a 3D pose lifting operation is performed to generate a global 3D pose sequence, and the contrast loss between the estimated 3D pose and the true 3D pose is calculated; Through the center frame 3D pose contrast learning module, a Transformer with a step size is used to reduce the global information to the target frame, the global information is used to constrain the features of the target frame, and the contrast loss between the predicted 3D pose of the center frame and the true 3D pose is calculated.
6. A 3D human body posture estimation method according to claim 5, characterized in that: In step S7, the function of the overall loss is defined as: in, , and is a hyperparameter, representing the weight of each type of loss function, represents the contrast loss between the predicted 2D pose and the 2D prior information, represents the contrast loss between the estimated 3D pose and the true 3D pose, represents the contrast loss between the predicted 3D pose of the center frame and the ground-truth 3D pose.
7. A 3D human body posture estimation method according to claim 1, characterized in that: In step S7, the sparse-dense training strategy includes: In the sparse training stage, the random mask rate is set to a non-zero value for sparse training to randomly mine the consistency between frames; In the dense training stage, the random mask rate is set to 0 for dense training to capture the global spatiotemporal representation.
8. A 3D human body posture estimation system, characterized in that: The system comprises: The data preprocessing module is used to extract 2D key point sequences from monocular 2D videos for posture space embedding and generate high-dimensional feature representations for each frame; A global-local feature fusion module is used to apply random mask operations and position encoding to the high-dimensional feature representation to obtain a sparse temporal Token sequence; a Transformer-based posture-level temporal interaction module is used to extract global temporal features from the temporal Token sequence; for the current frame and its neighboring frames, a joint-level spatial embedding is used to obtain a joint-level Token sequence, and a joint-level spatial interaction module and hierarchical convolution are used to capture local temporal details to obtain local features; the global temporal features and the local features are fused through an adaptive fusion strategy to generate fused features; A posture optimization module is used to input the fusion features into a posture optimization network, wherein the posture optimization network includes a 2D posture contrast learning module, a global 3D posture contrast learning module and a center frame 3D posture contrast learning module, so as to generate a 3D human body posture; based on the generated 3D human body posture and the corresponding prior information, the overall loss is calculated, and the posture optimization network is optimized and trained using a sparse-dense training strategy; The output module is used to output a 3D human body posture corresponding to the monocular 2D video based on the posture optimization network after optimization training.
9. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a 3D human body posture estimation method as described in any one of claims 1 to 7 are implemented.
10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of a 3D human body posture estimation method as described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Method and device for constructing attitude estimation model and attitude estimation method
CN117114083A
Feature fusion 3D human body posture estimation method based on GCN and Transform
CN118212689A
Human body posture estimation method based on conditional double-branch diffusion model
CN118968552A
Three-dimensional human body posture estimation method based on space-time Transform
CN119418399A
Method for generating virtual object, method for processing three-dimensional pose, and electronic device
WO2024228663A1
Cited By
Training method and estimation method of single-view multi-person three-dimensional attitude estimation model
CN120356037A