3D human body posture estimation method in dynamic shielding scene

By fusing structure-enhanced position embedding and occlusion-aware attention embedding in the Transformer network, and combining them with temporal and spatial attention modules, the problem of inaccurate keypoint localization in 3D pose estimation under occlusion scenarios is solved, achieving efficient and accurate 3D human pose reconstruction.

CN120976978APending Publication Date: 2025-11-18NANCHANG INST OF TECH

Patent Information

Application Number
CN202511495917.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing two-stage 3D pose estimation methods suffer from inaccurate joint localization in occluded scenes, and have high model complexity, consume a lot of training resources, and are difficult to adapt to diverse human movement patterns.

Method used

We employ a Transformer-based deep learning network, combining structure-enhanced position embedding and occlusion-aware attention embedding to construct a feature selection module. We capture joint temporal correlations through the temporal attention module of the mobile vision Transformer, and enhance spatial context information through strided convolution and cross-stage feature fusion to achieve efficient 3D pose estimation.

Benefits of technology

It improves the accuracy of 3D human pose estimation in occluded scenes, reduces model complexity, enhances the network's generalization ability and pose estimation accuracy, and significantly reduces MPJPE and P-MPJPE values.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976978A_ABST
    Figure CN120976978A_ABST
Patent Text Reader

Abstract

The invention records a 3D human body posture estimation method in a dynamic shielding scene. The method comprises the following steps: constructing a feature screening module; performing coordinate mapping and feature fusion on the human body 2D posture sequence, and outputting original scale features and high-dimensional features; constructing a time attention MVT2vs module, connecting the attitude embedding and attention mechanism of the time encoder, capturing the time correlation between human joints through the attention mechanism, performing output calculation on different layers of the time encoder, and outputting a feature vector; carrying out space down-sampling on the feature tensor through step convolution receiving, aggregating local context information of step convolution, and outputting down-sampling scale features and the processed feature tensor; and constructing cross-stage feature fusion, applying space-time attention and channel double attention in parallel to sequentially carry out same-scale splicing, cross-scale splicing and 1 * 1 convolution fusion on multi-scale features, generating a final output fusion feature tensor, inputting the final output fusion feature tensor into a three-dimensional attitude regression head, and outputting an optimized three-dimensional human body attitude coordinate sequence.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of computer vision and artificial intelligence, and particularly relates to a 3D human pose estimation method in a dynamic occlusion scene. BACKGROUND

[0002] Two-stage 3D pose estimation is a commonly used 3D human pose estimation method at present. In the first stage, a mature 2D pose detector is used to obtain 2D coordinates of the joint nodes. In the second stage, a 2D-to-3D lifting network is used to map the 2D coordinate sequence to a three-dimensional space. This method integrates temporal information in the second stage to mine the spatial dependence and temporal continuity of human joints, so as to alleviate the positioning error of the joint nodes caused by occlusion. For example, the SPGformer network combines the modeling ability of the graph convolution network (GCN) for the topology structure of the human skeleton and the global context awareness ability of the Transformer, which can effectively reduce the interference of occlusion. However, the use of the serial-parallel multi-branch coupled structure significantly increases the model complexity, consumes a large amount of computing resources in the training process, and the preset static graph structure is difficult to flexibly adapt to diversified human motion patterns.

[0003] Based on the Transformer architecture, the multi-head self-attention mechanism can dynamically allocate weights according to the global context of the input data, adaptively process the occluded area, and reasonably infer the missing information according to the visible information even if part of the joint is occluded, such as the Strided Transformer network, which uses a stride convolution to replace the fully connected layer in the standard Transformer encoder, gradually compresses the sequence length and aggregates local context information, and realizes efficient mapping from a 2D sequence to a 3D pose. However, its dependence on a single linear layer to extract spatial features limits the ability to model spatial information. SUMMARY

[0004] In order to solve the problems existing in the prior art to some extent, based on the Transformer network with attention mechanism which can effectively capture long sequence dependence, the present application proposes a 3D human pose estimation method in a dynamic occlusion scene, and the specific technical solutions are as follows:

[0005] A 3D human pose estimation method in a dynamic occlusion scene, comprising the following steps:

[0006] S1, in the Transformer-based deep learning network, the structure-enhanced position embedding and the occlusion-aware attention embedding are fused to construct a feature screening module; the 2D pose sequence of the human body is mapped and the features are fused through the feature screening module to output original scale features and high-dimensional features;

[0007] S2, constructing a time attention MVT2vs module based on a mobile vision Transformer, connecting a pose embedding of a time encoder and an attention mechanism, and capturing a time correlation between human body joints through the attention mechanism, outputting a feature tensor by output calculation on different layers of the time encoder;

[0008] S3, receiving the feature tensor through a stride convolution, spatially down-sampling the feature tensor, aggregating local context information of the stride convolution, and outputting a down-sampled scale feature and a processed feature tensor;

[0009] S4, outputting multi-scale features based on steps S1-S3, constructing cross-stage feature fusion, and sequentially performing same-scale splicing, cross-scale splicing and 1×1 convolution fusion on the multi-scale features through parallel application of space-time attention and channel double attention, to output a fusion feature tensor;

[0010] S5, inputting the fusion feature tensor into a three-dimensional pose regression head, adjusting a 3D human body pose coordinate sequence, predicting three-dimensional coordinates of each joint of the human body, and outputting an optimized three-dimensional human body pose coordinate sequence.

[0011] Further,

[0012] mapping the coordinates of the human body 2D pose sequence into original scale features through one-dimensional convolution, wherein the human body 2D pose sequence is obtained from an image or a video frame, and the expression is:

[0013] ,

[0014] In the formula, is a human body 2D pose sequence; is a frame number of the 2D human body joint sequence, is a 2D joint position of the Tth frame, , is a real matrix, is a number of joints.

[0015] Further,

[0016] The structure-enhanced position embedding is obtained by the following method: constructing a relative position matrix between human body joints, and mapping the relative position matrix to a space with the same feature dimension through a learnable linear layer, and the expression is:

[0017] ,

[0018] In the formula, is a structure-enhanced position embedding; denotes a linear transformation from dimension to dimension; is a vector of dimensions is a vector of dimensions is a vector of dimensions

[0019] The occlusion-aware attention embedding is obtained by the following method: predicting the occlusion confidence of each joint of the human body, mapping the occlusion confidence to the same space as its feature dimension through a learnable linear layer, and the expression is:

[0020] ,

[0021] In the formula, is an occlusion-aware attention embedding; represents a linear transformation from a J-dimensional space to a dimensional space; is a J-dimensional vector; is a vector of dimensions

[0022] Further,

[0023] Add the structure-enhanced position embedding and the occlusion-aware attention embedding to generate a guide signal;

[0024] Nonlinearly transform the guide signal through the SwiGLU activation function to output the enhanced structure-aware position feature;

[0025] Add the enhanced structure-aware position feature to the original feature to output a one-dimensional convolution-processed high-dimensional feature;

[0026] Add the one-dimensional convolution-processed high-dimensional feature to the enhanced structure-aware position feature to output a high-dimensional feature .

[0027] Further,

[0028] The time attention MVT2vs module is constructed by the following steps:

[0029] S201, defining the high-dimensional feature as , the high-dimensional feature is reshaped into an input feature ;

[0030] S202, the time encoder uses a learnable position embedding to perform residual processing with the input feature , and outputs an embedded feature ;

[0031] S203, setting Branch 1 to sequentially perform Linear and Softmax operations on the embedded feature , and output attention weights , the expression is:

[0032] ,

[0033] In the formula, Attention weights; For embedded features; The weights of the linear layer in branch 1; The bias for the linear layer of branch 1;

[0034] S204, Set branch 2 parallel to branch 1 for the embedded feature. Perform Linear and Maxout activation function operations sequentially to output nonlinear transformation characteristics. The expression is:

[0035] ,

[0036] In the formula, It exhibits nonlinear transformation characteristics; For embedded features; The weights of the linear layer in branch 2; The bias for the linear layer of branch 2;

[0037] S205, the attention weights Characteristics of nonlinear transformation Perform element-wise multiplication and accumulation operations to output the fused features. ;

[0038] S206, Set branch 3 to integrate the features Characteristics of nonlinear transformation Perform element-wise multiplication, then perform Linear and Maxout operations to output the enhanced features. The expression is:

[0039] ,

[0040] In the formula, To enhance features; It exhibits nonlinear transformation characteristics; Features of fusion; The weights of the linear layer in branch 3; This is the bias for the linear layer of branch 3.

[0041] Furthermore,

[0042] The output of the time encoder is calculated through the Transformer encoder layers at different levels, and the output feature tensor is as follows:

[0043] When the number of time encoder layers At that time, the feature tensor output by the time encoder is The expression is:

[0044] ,

[0045] When the number of time encoder layers At that time, the feature tensor output by the time encoder is The expression is:

[0046] ,

[0047] In the formula, This is the feature tensor output when the time encoder has 1 layer. To enhance features; Number of layers for time encoder The feature tensor output at that time; For the time encoder The output of the layer preceding the previous layer; For the attention of multiple parties; For layer normalization; It is a feedforward network.

[0048] Furthermore,

[0049] In step S3, the specific steps of spatial downsampling are as follows:

[0050] S301, Construct a hierarchical downsampling unit. The hierarchical downsampling unit consists of cascaded processing layers. Each processing layer performs feature transformation through a unified calculation formula, as shown below:

[0051] ,

[0052] In the formula, The features are those processed at the k-th layer; These are the features before processing at the (k-1)th layer; For max pooling; For the attention of multiple parties; For layer normalization; It is a convolutional feedforward network;

[0053] S302, performs a recursive calculation mechanism on the hierarchical downsampling unit: when k=1, the input is When k≥2, the input is the output of the previous layer. .

[0054] Furthermore,

[0055] Branch 1 is configured to perform layer normalization, linear layer and Sigmoid operation sequentially on the feature vector and / or the processed feature tensor, and output a weight vector.

[0056] Branch 2, which runs parallel to branch 1, performs adaptive average pooling and Softmax normalization operations sequentially on the feature vector and / or the processed feature tensor, and outputs the weights.

[0057] The feature vector and / or the processed feature tensor, weight vector, and weights are weighted and fused to output the fused feature vector, expressed as:

[0058] ,

[0059] In the formula, The feature tensor after fusion; For feature vectors and / or processed feature tensors; This is a weighted fusion function; This is the weight vector; As weight.

[0060] Furthermore,

[0061] Step S4 includes the following specific steps:

[0062] a. Extract the original scale features of step S1 at Scale 1:1, and extract the 1 / 2 downsampled scale features and 1 / 4 downsampled scale features of step S3 at Scale 1:2 and Scale 1:4 to construct multi-scale spatiotemporal features;

[0063] b. Input the multi-scale spatiotemporal features into the parallel spatiotemporal attention module and channel attention module, and output the multi-scale spatiotemporal attention enhanced feature map and the multi-scale channel attention enhanced feature map;

[0064] c. For the multi-scale spatiotemporal features, their corresponding spatiotemporal attention enhancement feature maps and channel attention enhancement feature maps are concatenated along the channel dimension to output a multi-scale fusion feature map.

[0065] d. The multi-scale fused feature map is spliced ​​along the channel dimension to form a cross-scale aggregated feature tensor; a 1×1 convolution operation is performed on the cross-scale aggregated feature tensor to realize cross-channel information fusion and feature dimension compression, and output the fused feature map.

[0066] e. The fused feature map is added element-wise to the original scale feature map to generate the final output fused feature tensor.

[0067] Furthermore,

[0068] The expression for the fused feature tensor is:

[0069] ,

[0070] In the formula, To fuse feature tensors; For cross-scale splicing operations; Original scale features; Features at a 1 / 2 downsampling scale; Features at a 1 / 4 downsampling scale; The input features are the original scale.

[0071] Based on the above technical solution, the present invention has the following beneficial effects:

[0072] 1. The method described in this invention adopts a three-stage training paradigm, such as... Figure 1 As shown in (a): the pre-training stage (steps S1-S4), the cross-stage feature fusion stage (step S5), and the fine-tuning stage (step S6) are based on Transformer. The spatiotemporal attention mechanism can capture and fuse the spatiotemporal features of human pose more accurately and adaptively adjust the feature weights to achieve higher accuracy in 3D human pose reconstruction.

[0073] 2. The cross-stage feature fusion described in this invention: ① By deploying spatiotemporal attention modules and channel attention modules in parallel and applying them to three input scales respectively, it achieves synchronous enhancement of key spatial regions and important channel features, which can significantly improve feature discrimination in occluded scenes; ② It adopts a three-level processing flow of same-scale stitching → cross-scale stitching → 1×1 convolution fusion to ensure efficient integration of multi-granularity information, while realizing feature space adaptation and dimensionality reduction through convolution operations; the final output retains the detailed information of the original scale features through residual connections, avoiding information loss during the fusion process and improving gradient propagation efficiency; ③ It solves the problem of inconsistent feature distribution between the pre-training and fine-tuning stages, and significantly improves the pose estimation accuracy in the subsequent fine-tuning stage through multi-scale context fusion and attention weighting mechanism.

[0074] 3. Based on the Human3.6M dataset and Protocol#1 evaluation protocol, the evaluation conclusions of the algorithm performance of the method described in this invention through experiments show that the method exhibits excellent network generalization ability and overall 3D human pose estimation ability. Specifically: ① Under noisy input and unreliable data conditions, the method described in this invention performs best or second best in most actions. Compared with the benchmark algorithm Strided Transformer, when the number of frames T is 351, the MPJPE value of the algorithm in this paper is reduced by 2.5% (1.1mm) and the P-MPJPE value is reduced by 2.3% (0.8mm) when the 2D joints extracted by CPN are used as input. Overall, the average error is the lowest and the performance is the best; ② When the data is relatively clean, the method described in this invention can achieve the best pose estimation results in most behavioral scenarios, indicating that when the input data has less noise interference, the algorithm can effectively reduce the pose estimation error and obtain better performance than the benchmark algorithm Strided Transformer. Attached Figure Description

[0075] Figure 1 This is a schematic diagram of the overall network architecture of the method of the present invention, including a flowchart (a), a spatiotemporal attention module (b), and a channel attention module (c).

[0076] Figure 2 This is a schematic diagram of the improved feature filtering module structure;

[0077] Figure 3 This is a schematic diagram of a time encoder network;

[0078] Figure 4 The diagrams show a comparison of the MVT2vs module structure before and after the improvement, including a structural diagram (a) and a basic module diagram (b).

[0079] Figure 5 This is a schematic diagram of the space encoder structure;

[0080] Figure 6 A schematic diagram of the improved channel attention module;

[0081] Figure 7 This is a schematic diagram of cross-stage feature fusion. Detailed Implementation

[0082] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0083] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning as understood by one of ordinary skill in the art to which this disclosure pertains.

[0084] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0085] In all examples shown and discussed herein, any specific value should be interpreted as merely exemplary and not as a limitation; therefore, other examples of exemplary embodiments may have different values.

[0086] For ease of description, in the specification and claims of this disclosure, "Transformer-based deep learning network" will be simply referred to as "deep learning network".

[0087] This embodiment describes a 3D human pose estimation method in a dynamic occlusion scenario. Figure 1 The overall process of the method is shown, which includes the following steps.

[0088] S1, in a Transformer-based deep learning network, structurally enhanced positional embeddings and occlusion-aware attention embeddings are fused to construct a Feature Selection Module (FFM), such as... Figure 2 As shown, the feature selection module performs coordinate mapping and feature fusion on the 2D pose sequence of the human body, thereby enhancing the deep learning network's ability to perceive the spatial position of human joints.

[0089] Specifically, the following steps are included:

[0090] S101, continuously extract the coordinate sequence of two-dimensional human body key points from image or video frames to obtain the 2D human body pose sequence, as shown in the following expression:

[0091] ,

[0092] In the formula, This is a 2D human pose sequence; The number of frames in a 2D human joint sequence. Let T be the 2D joint position of the T-th frame. , It is a real matrix. This represents the number of joints.

[0093] S102 maps the coordinates of a 2D human pose sequence to the original scale features through one-dimensional convolution.

[0094] S103, construct the relative position matrix between human joints, and map the relative position matrix to the feature dimension d through a learnable linear layer. m In the same space, the location embedding of structural enhancement is obtained. (SPE) The expression is:

[0095] ,

[0096] ,

[0097] In the formula, For structural enhancement, the positional embedding has a dimension of ; Indicates a from Dimension (number of joint pairs) mapped to Linear transformation of the feature dimension; It is A 3D vector containing the relative position information between all joint pairs; is the dimension of the feature, i.e., the number of feature channels after one-dimensional convolution; T is the number of frames in the input sequence.

[0098] Structurally enhanced positional embeddings can map the relative positional information between joints to a space with the same feature dimension through a learnable linear layer, so as to integrate it with the original features obtained through one-dimensional convolution. By integrating these technologies, deep learning networks can be enhanced to perceive the spatial position of human joints, especially when dealing with occlusion problems.

[0099] S104 predicts the occlusion confidence for each joint of the human body (using the joint confidence output by the pose detector or through an additional lightweight CNN network), mapping the occlusion confidence to a feature dimension d through a learnable linear layer. m In the same space, occlusion perception attention embedding E (OAA) The expression is:

[0100] ,

[0101] ,

[0102] In the formula, To embed attentional information for occlusion perception, the dimension is... ; For a mapping from J-dimensional (number of joint pairs) to Linear transformation of the feature dimension; It is a J-dimensional vector; is the dimension of the feature, i.e., the number of feature channels after one-dimensional convolution; T is the number of frames in the input sequence.

[0103] Occlusion-aware attentional embedding maps the occlusion confidence of each joint to a space with the same feature dimension through a learnable linear layer, so as to be correlated with the original features obtained through one-dimensional convolution. By fusing the data, the model's ability to perceive occluded areas is enhanced, thereby improving the accuracy of pose estimation under occlusion conditions.

[0104] S105 fuses the structure-enhanced location embedding and occlusion-aware attention embedding with the original feature Z0 to output a high-dimensional feature processed by one-dimensional convolution. The specific steps are as follows:

[0105] S105-1, the structure-enhanced position embedding and occlusion-aware attention embedding are added together to generate a joint structure-occlusion guidance signal. The expression is , G ∈ ;

[0106] S105-2, because the SwiGLU activation function (Swish-Gated Linear Unit) can effectively handle gradient flow and enhance the nonlinear expressive power of the model, especially beneficial for preserving weak feature signals of occluded joints, this embodiment uses the SwiGLU activation function to perform a nonlinear transformation on the guiding signal G, outputting enhanced structure-aware position features. The expression is as follows:

[0107] ,

[0108] In the formula, σ(G) is the enhanced structure-aware location feature, that is, the final generated location feature that integrates structural prior and occlusion information; For activation functions; For guiding signals;

[0109] S105-3 adds the enhanced structure-aware location feature σ(G) to the original feature Z0, outputting a high-dimensional feature processed by one-dimensional convolution. ,Right now:

[0110] ,

[0111] S105-4, High-dimensional features processed by one-dimensional convolution Compared with enhanced structure-aware location features Add them together to output high-dimensional features. ,Right now:

[0112] ,

[0113] Based on the above description, the introduction of structure-enhanced position embedding can solve the problem of insufficient modeling of the relative position of joints in traditional position embedding; the introduction of occlusion-aware attention embedding can identify the weak feature signals of occluded joints and reconstruct the features of occluded joints based on the spatial relationship of adjacent joints; the SwiGLU function can solve the gradient vanishing problem in joint feature learning and retain the weak feature signals of occluded joints.

[0114] S2. In order to enhance the ability of deep learning networks to capture the temporal nonlinear relationship in human pose and improve the accuracy and generalization ability of human pose estimation, this step constructs a temporal attention MobileViTv2vs module (MVT2vs module for short) based on mobile vision Transformer. It connects the pose embedding and attention mechanism of the temporal encoder, and captures the temporal correlation between all human joints through the attention mechanism. It performs output calculation on different layers of the temporal encoder and outputs feature tensors.

[0115] The specific steps are as follows:

[0116] S201, regarding the high-dimensional features recorded in step S1 The input features are reshaped to obtain the reshaped features. , ;

[0117] S202, the time encoder uses learnable position embeddings , and the reshaped input features Perform residual processing and output embedded features. The expression is:

[0118] ,

[0119] S203, such as Figure 4 As shown in (a), a temporal attention MVT2vs module of the mobile vision Transformer is constructed to process the embedded features output in step S202. Enhanced features are output after computation and processing. The specific steps are as follows:

[0120] S203-1, Set branch 1 to embed features Perform Linear and Softmax operations sequentially to output attention weights. The expression is:

[0121] ,

[0122] In the formula, Attention weights; For embedded features; The weights of the linear layer in branch 1; This is the bias for the linear layer of branch 1.

[0123] S203-2, Set up branch 2 parallel to branch 1 to embed features The Linear (linear layer) and Maxout activation functions are applied sequentially to output nonlinear transformation features. The expression is:

[0124] ,

[0125] In the formula, It exhibits nonlinear transformation characteristics; For embedded features; The weights of the linear layer in branch 2; This is the bias for the linear layer of branch 2.

[0126] The Maxout activation function can better capture the nonlinear relationships of local features, enabling deep learning networks to have excellent nonlinear representation capabilities and map the pose feature signals of 3D human bodies to a larger range, so as to better preserve the details and differences of local features.

[0127] S203-3, Attention Weights Characteristics of nonlinear transformation Perform element-wise multiplication and accumulation operations to output the fused features. The expression is:

[0128] ,

[0129] In the formula, Features of fusion; Attention weights; It exhibits nonlinear transformation characteristics.

[0130] S203-4, Setting branch 3 will merge features Characteristics of nonlinear transformation Perform element-wise multiplication, then perform Linear and Maxout operations to output the enhanced features. The expression is:

[0131] ,

[0132] In the formula, To enhance features; It exhibits nonlinear transformation characteristics; Features of fusion; The weights of the linear layer in branch 3; This is the bias for the linear layer of branch 3.

[0133] By introducing the non-linear activation function Maxout at key locations, the MVT2vs module can better handle and learn complex human pose distributions. Key locations refer to the feature locations in a multi-head attention mechanism that the model needs to pay special attention to for capturing the temporal correlation of human joints. These locations are learned through previous network layers (such as the MVTTA module), such as... Figure 4 As shown in (b) above, or determined through the data preprocessing stage.

[0134] S204, Enhanced features based on the output of step S203 By using standard Transformer encoder layers that include Layer Normalization (LN), Multi-Head Attention (MSA), and Feedforward Network (FFN), the output of the temporal encoder at different layers is calculated, and the output feature vector is generated, such as... Figure 3 As shown, the details are as follows:

[0135] When the number of time encoder layers At that time, the feature tensor output by the time encoder is The expression is:

[0136] ,

[0137] In the formula, This is the feature tensor output when the time encoder has 1 layer. To enhance features; For the attention of multiple parties; For layer normalization; It is a feedforward network.

[0138] When the number of time encoder layers At that time, the feature tensor output by the time encoder is :

[0139] ,

[0140] In the formula, Number of layers for time encoder The feature tensor output at that time; For the time encoder The output of the layer preceding the previous layer; For the attention of multiple parties; For layer normalization; It is a feedforward network.

[0141] In this step:

[0142] It is the first layer time encoder ( The output of (=1) represents the state of the input features after preliminary time modeling. At this point, the model has captured the most basic and shallow time dependencies between joints.

[0143] It is the last layer of time encoder (the The output of each layer represents the final state of the input features after deep temporal modeling. As the number of layers increases, each layer integrates more complex and abstract temporal information based on the previous layer.

[0144] This progressive structure enables the model to capture long-range temporal dependencies in human pose sequences from a shallow to a deep level. (Shallow output) focuses more on simple motion patterns of joints between adjacent frames (e.g., the movement vector of the wrist over two or three consecutive frames). (Deep output) enables the understanding of more complex global temporal contexts (e.g., the coordinated movements of the shoulder, elbow, and wrist joints throughout a complete "waving" motion cycle). This is crucial for understanding periodic motions and inferring frame information missing due to occlusion.

[0145] S3, as Figure 5 As shown, in order to effectively aggregate local spatial context information and optimize feature dimensions, a strided convolution (Strided Transformer Encoder) receives the feature tensor output by the temporal encoder in step S2, performs spatial downsampling on it, outputs a downsampled scale feature map, and aggregates the local context information from the strided convolution to output a processed feature map. Feature tensors after layer processing .

[0146] The specific steps of spatial downsampling are as follows:

[0147] S301, Construct a hierarchical downsampling unit, the hierarchical downsampling unit consists of... The system consists of cascaded processing layers, with each layer performing feature transformation through a unified computational formula, as shown below:

[0148] ,

[0149] In the formula, The features are those processed at the k-th layer; For the first Features before layer processing; For max pooling; For the attention of multiple parties; For layer normalization; It is a convolutional feedforward network.

[0150] S302, performs a recursive calculation mechanism on the hierarchical downsampling unit: when k=1, the input is When k≥2, the input is the output of the previous layer. .

[0151] S4. To improve the feature discrimination power and training convergence of deep learning networks, a channel attention module (CAM) is constructed, such as... Figure 1 (c) and Figure 6 As shown, the feature tensor output in step S2 and / or the processed feature tensor output in step S3 are input to highlight human pose features and suppress background interference.

[0152] The specific steps are as follows:

[0153] S401, set branch 1 to perform layer normalization, linear layer and sigmoid operation on the feature vector and / or the processed feature tensor in sequence, and output the weight vector;

[0154] S402, set branch 2, which runs parallel to branch 1, to perform adaptive average pooling and Softmax normalization operations on the feature vector and / or the processed feature tensor in sequence, and output the weights;

[0155] S403 performs a weighted fusion of the feature tensor, weight vector, and weights, outputting the fused feature vector, expressed as follows:

[0156] ,

[0157] In the formula, The feature tensor after fusion; For feature vectors and / or processed feature tensors; This is a weighted fusion function; This is the weight vector; As weight.

[0158] S5, constructing cross-stage feature fusion, such as Figure 7 As shown, by applying spatiotemporal attention and channel dual attention in parallel, the multi-scale features output in steps S1-S4 are sequentially subjected to same-scale concatenation, cross-scale concatenation, and 1×1 convolution fusion to output fused features, and the original features are preserved through residual connections. The original details are obtained to resolve the issue of inconsistent feature distributions in steps S1-S4.

[0159] The specific steps are as follows:

[0160] S501, extract the original scale feature map of step S1 at Scale 1:1, and extract the 1 / 2 downsampled scale feature map and 1 / 4 downsampled scale feature map of step S3 at Scale 1:2 and Scale 1:4 respectively, to construct multi-scale spatiotemporal features, and input the multi-scale spatiotemporal features into the parallel spatiotemporal attention module and channel attention module, and output the multi-scale spatiotemporal attention enhanced feature map and the multi-scale channel attention enhanced feature map.

[0161] The specific steps are as follows:

[0162] S501-1, as follows Figure 1 As shown in (b), the spatiotemporal attention module is computed through the following steps:

[0163] ① The time encoder first uses the learnable position embedding and the reconstructed input features to perform residual processing to obtain the embedding features;

[0164] ② Introduce the mobile vision Transformer temporal attention MVT2vs module to bridge the gap between the pose embedding and attention mechanism of the temporal encoder. Utilize its attention mechanism to capture the temporal correlation between all joints, enhance the network's ability to capture local pose feature temporal information, and thus improve the network's overall ability to capture pose features.

[0165] ③ The MVT2vs module performs computations through a parallel branching design, including two branches: one for generating attention weights and the other for generating features. The attention weight branch performs Linear and Softmax operations on the embedded features to obtain the attention weights; the feature branch performs Linear and Maxout activation function operations on the embedded features to obtain the features.

[0166] ④ Perform element-wise multiplication and summation operations on the attention weights and features to obtain the fused features;

[0167] ⑤ Perform element-wise multiplication between the fused features and the original features, and then perform Linear and Maxout operations to obtain the output of the MVT2vs module.

[0168] The spatiotemporal attention module is applied to the original scale feature map, the 1 / 2 downsampled scale feature map, and the 1 / 4 downsampled scale feature map, respectively. It can calculate and enhance the feature responses that are important to spatial location, and output the spatiotemporal attention enhanced feature map at the original scale, the spatiotemporal attention enhanced feature map at the 1 / 2 downsampled scale, and the spatiotemporal attention enhanced feature map at the 1 / 4 downsampled scale.

[0169] S501-2, such as Figure 1 (c) and Figure 6 As shown, the Channel Attention Module (CAM) operation is performed through the following steps:

[0170] ① Input feature tensor:

[0171] CAM receives feature tensors from the output of the spatiotemporal attention module (temporal / spatial encoder), and the feature vectors contain information encoded in both time and space.

[0172] ② Two-branch processing:

[0173] CAM consists of two parallel branches, each of which processes the feature tensor to generate a different weight vector.

[0174] Branch 1: Process the feature tensor using Layer Normalization (LN), LinearLayer transformation, and Sigmoid activation function to generate a weight vector X. v This indicates the importance of each channel.

[0175] Branch 2: Perform adaptive average pooling and softmax normalization on the feature tensor to generate another weight vector X. p This also indicates the importance of the channel.

[0176] ③ Weighted fusion:

[0177] The weight vector X v and weight vector X p The fusion process is performed to highlight important feature channels and suppress unimportant channels. This is typically achieved through element-wise multiplication, where corresponding elements of the two weight vectors are multiplied for each channel to obtain the final weighted feature X′.

[0178] ④ Output feature tensor:

[0179] The weighted fused feature tensor X′ contains enhanced human pose features and suppressed background interference, which will be used in subsequent network layers or as the final output.

[0180] The channel attention module is applied to the original scale feature map, the 1 / 2 downsampled scale feature map, and the 1 / 4 downsampled scale feature map, respectively. It can calculate and enhance the feature responses that are important to the feature channels, and output the channel attention enhanced feature map at the original scale, the channel attention enhanced feature map at the 1 / 2 downsampled scale, and the channel attention enhanced feature map at the 1 / 2 downsampled scale.

[0181] S502: For multi-scale spatiotemporal features, the corresponding spatiotemporal attention-enhanced feature maps and channel attention-enhanced feature maps are concatenated along the channel dimension to output a multi-scale fused feature map.

[0182] S503 concatenates the multi-scale fused feature maps along the channel dimension to form a cross-scale aggregated feature tensor; performs a 1×1 convolution operation on the cross-scale aggregated feature tensor to achieve cross-channel information fusion and feature dimension compression, and outputs a fused feature map.

[0183] S504, the fused feature map is added element-wise to the original scale feature map to generate the final output fused feature tensor. The expression is:

[0184] ,

[0185] In the formula, The final output is the fused feature tensor; For cross-scale splicing operations; This is the original scale feature map; Feature map at a 1 / 2 downsampling scale; This is a feature map at a 1 / 4 downsampling scale; The input feature map is the original scale.

[0186] S6, fuse the final output obtained in step S5 with the feature tensor. The input is a 3D pose regression head, which is the preferred one. The model is trained on the 3D pose regression head to make fine adjustments to the 3D human pose coordinate sequence, directly predict the 3D coordinates (x, y, z) of each joint of the human body, and output the final optimized 3D human pose coordinate sequence.

[0187] Based on the Human3.6M dataset and Protocol#1 evaluation protocol, the algorithm performance of the method described in this embodiment is evaluated through experimental design. I. Brief Description of Terms

[0188] Before presenting the experimental design, the Human 3.6M dataset and Protocol #1 evaluation protocol are briefly described below: 1. Human 3.6M dataset

[0189] The Human3.6M dataset is primarily used for research in areas such as human pose estimation, motion analysis, and action recognition. It contains approximately 3.6 million human poses and their corresponding images. The dataset was filmed indoors from four different perspectives by 11 professional actors (A1-A11) depicting 15 everyday behaviors: “Dir.” (directions), “Disc.” (discussion), “Eat” (eating), “Greet” (greeting), “Phone” (talking on the phone), “Photo” (taking a photo), “Pose” (posing), “Purch.” (making purchases), “Sit” (sitting), “SitD.” (sitting down), “Smoke” (smoking), “Wait” (waiting), “WalkD.” (walking the dog), “Walk” (walking), and “WalkT.” (walking together).

[0190] The Human3.6M dataset provides 2D and corresponding 3D pose annotation data. Among them, only A1, A5, A6, A7, A8, A9, and A11 have 3D human pose labels. In this experiment, the action samples of these five people, A1, A5, A6, A7, and A8, are used as the training set, and the action samples of A9 and A11 are used as the test set to test the algorithm performance of the method described in this embodiment. 2. Protocol #1 Evaluation Agreement

[0191] Protocol #1 is a commonly used evaluation protocol in the field of 3D human pose estimation, used to measure the performance of models in different task scenarios. Protocol #1 is a relatively comprehensive evaluation method, taking into account factors such as character rotation and scaling, scale bias, etc., and better reflects the overall capability of 3D human pose estimation networks. It uses the mean per joint position error (MPJPE) as the evaluation metric. The smaller the MPJPE and P-MPJPE values, the smaller the gap between the estimated 3D human pose and the true value, and the higher the accuracy of the model's 3D human pose estimation. II. Setting up the experimental plan

[0192] The algorithm described in this embodiment is compared with other classic 3D human pose estimation algorithms (including StridedTransformer, FTCM_S, Real-Time Pose3D, ARGP-Pose, SRNet, etc.) on the Human3.6M dataset. Option 1:

[0193] Using 2D keypoints extracted from images by a cascaded pyramid network (CPN) for 2D pose detectors as input, the generalization ability and overall 3D human pose estimation ability of the network were evaluated under conditions of noisy input and unreliable data. The Protocol#1 evaluation metric was used, and the results are shown in Table 1.

[0194] Table 1. Performance comparison of algorithms using CPN as input under Protocol #1

[0195] Note: The optimal value is indicated in bold, and the suboptimal value is indicated in underline.

[0196] As shown in Table 1: Because of significant self-occlusion in the "Sit" and "SitD." actions, the input 2D human joints contain large errors, resulting in relatively high MPJPE values ​​for the algorithms. For the "Phone" and "Photo" actions, the self-occlusion and complexity of the actions lead to significant errors in the pose estimation of the algorithms. "Walk" and "WalkT." are two relatively simple periodic actions, so all algorithms can achieve good pose estimation results.

[0197] The algorithm described in this embodiment performs best or second best in most actions. Compared with the benchmark algorithm StridedTransformer, when the number of frames T is 351, the MPJPE value of the algorithm in this paper is reduced by 2.5% (1.1mm) and the P-MPJPE value is reduced by 2.3% (0.8mm) when the 2D joints extracted by CPN are used as input. Overall, the average error is the lowest and the performance is the best. Option 2:

[0198] Using the 2D ground truth (GT), i.e., the 2D labeled joints in the Human3.6M dataset, as input, we can evaluate the network's ability to estimate 3D human pose when the data is relatively clean. The Protocol#1 evaluation metric is used, and the results are shown in Table 2.

[0199] Table 2 Comparison of algorithm performance using ground truth (GT) as input under Protocol #1

[0200] Note: The optimal value is indicated in bold, and the suboptimal value is indicated in underline.

[0201] As shown in Table 2: Compared to the baseline algorithm Strided Transformer, the algorithm described in this embodiment reduces the MPJPE value by 1.9 mm (6.7%) when the frame number T is 351 and the 2D keypoints (GT) labeled on the Human3.6M dataset are used as input. The improvement is most significant for "Sit," with an MPJPE value reduction of 3.8 mm, indicating that the algorithm effectively alleviates the self-occlusion problem. For regular actions like "Eat," the MPJPE value decreases by 0.7 mm, meaning that even with regularly labeled 2D keypoints, the algorithm still improves accuracy and obtains more precise 3D human pose. Experimental results show that the algorithm achieves the best pose estimation results in most behavioral scenarios, indicating that when inputting data with low noise interference, it effectively reduces pose estimation errors and achieves better performance than the baseline algorithm Strided Transformer, further verifying the superior performance of the algorithm.

[0202] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. A method for 3D human pose estimation in dynamic occlusion scenarios, characterized in that, Includes the following steps: S1, In the Transformer-based deep learning network, the structure-enhanced position embedding and the occlusion-aware attention embedding are fused to construct a feature selection module; the feature selection module performs coordinate mapping and feature fusion on the human 2D pose sequence to output the original scale features and high-dimensional features; S2, construct a temporal attention MVT2vs module based on mobile vision Transformer, connect the pose embedding and attention mechanism of the temporal encoder, capture the temporal correlation between human joints through the attention mechanism, perform output calculation on different layers of the temporal encoder, and output feature tensors; S3, receive the feature tensor through strided convolution, perform spatial downsampling on it, aggregate the local context information of strided convolution, and output the downsampled scale features and the processed feature tensor; S4. Based on the multi-scale features output in steps S1-S3, construct cross-stage feature fusion. By applying spatiotemporal attention and channel dual attention in parallel, perform same-scale splicing, cross-scale splicing and 1×1 convolution fusion on the multi-scale features in sequence, and output the fused feature tensor. S5, input the fused feature tensor into the three-dimensional posture regression head, adjust the 3D human posture coordinate sequence, predict the three-dimensional coordinates of each joint of the human body, and output the optimized three-dimensional human posture coordinate sequence.

2. The 3D human pose estimation method in a dynamic occlusion scene according to claim 1, characterized in that, The coordinates of the human 2D pose sequence are mapped to the original scale features using one-dimensional convolution. The human 2D pose sequence is obtained from image or video frames, and its expression is: , In the formula, This is a 2D human pose sequence; The number of frames in a 2D human joint sequence. Let T be the 2D joint position of the T-th frame. , It is a real matrix. This represents the number of joints.

3. The 3D human pose estimation method in a dynamic occlusion scene according to claim 2, characterized in that, The location embedding for structural enhancement is obtained as follows: A relative position matrix between human joints is constructed, and this matrix is ​​mapped to a space with the same feature dimension through a learnable linear layer, expressed as: , In the formula, For structural reinforcement, embedding is required. Indicates a from Dimension mapping to A linear transformation of dimension; It is A dimensional vector; It is the dimension of the feature; The occlusion-aware attention embedding is obtained by predicting the occlusion confidence of each joint of the human body, and mapping the occlusion confidence to a space with the same feature dimension through a learnable linear layer, as expressed by: , In the formula, To embed attentional information for occlusion perception; Represents a mapping from J dimensions to A linear transformation of dimension; It is a J-dimensional vector; It is the dimension of the feature.

4. The 3D human pose estimation method in a dynamic occlusion scene according to claim 3, characterized in that, The location embedding of structure enhancement and the attention embedding of occlusion perception are added together to generate a guiding signal; The guiding signal is nonlinearly transformed using the SwiGLU activation function to output enhanced structure-aware position features; The enhanced structure-aware location features are added to the original features to output a high-dimensional feature processed by one-dimensional convolution. The high-dimensional features processed by the one-dimensional convolution are added to the enhanced structure-aware location features to output high-dimensional features. .

5. The 3D human pose estimation method in a dynamic occlusion scene according to claim 1, characterized in that, Construct the time-attention MVT2vs module using the following steps: S201, the high-dimensional feature is defined as The high-dimensional features are Remodeling into input features ; S202, the time encoder uses learnable position embeddings , with the input features Perform residual processing and output embedded features. ; S203, Set branch 1 to the embedded feature Perform Linear and Softmax operations sequentially to output attention weights. The expression is: , In the formula, Attention weights; For embedded features; The weights of the linear layer in branch 1; The bias for the linear layer of branch 1; S204, Set branch 2 parallel to branch 1 for the embedded feature. Perform Linear and Maxout activation function operations sequentially to output nonlinear transformation characteristics. The expression is: , In the formula, It exhibits nonlinear transformation characteristics; For embedded features; The weights of the linear layer in branch 2; The bias for the linear layer of branch 2; S205, the attention weights Characteristics of nonlinear transformation Perform element-wise multiplication and accumulation operations to output the fused features. ; S206, Set branch 3 to integrate the features Characteristics of nonlinear transformation Perform element-wise multiplication, then perform Linear and Maxout operations to output the enhanced features. The expression is: , In the formula, To enhance features; It exhibits nonlinear transformation characteristics; Features of fusion; The weights of the linear layer in branch 3; This is the bias for the linear layer of branch 3.

6. The 3D human pose estimation method in a dynamic occlusion scene according to claim 5, characterized in that, The output of the time encoder is calculated through the Transformer encoder layers at different levels, and the output feature tensor is as follows: When the number of time encoder layers At that time, the feature tensor output by the time encoder is The expression is: , When the number of time encoder layers At that time, the feature tensor output by the time encoder is The expression is: , In the formula, This is the feature tensor output when the time encoder has 1 layer. To enhance features; Number of layers for time encoder The feature tensor output at that time; For the time encoder The output of the layer preceding the previous layer; For the attention of multiple parties; For layer normalization; It is a feedforward network.

7. The 3D human pose estimation method in a dynamic occlusion scene according to claim 1, characterized in that, In step S3, the specific steps of spatial downsampling are as follows: S301, Construct a hierarchical downsampling unit. The hierarchical downsampling unit consists of cascaded processing layers. Each processing layer performs feature transformation through a unified calculation formula, as shown below: , In the formula, The features are those processed at the k-th layer; These are the features before processing at the (k-1)th layer; For max pooling; For the attention of multiple parties; For layer normalization; It is a convolutional feedforward network; S302, performs a recursive calculation mechanism on the hierarchical downsampling unit: when k=1, the input is When k≥2, the input is the output of the previous layer. .

8. The 3D human pose estimation method in a dynamic occlusion scene according to claim 1, characterized in that, Branch 1 is configured to perform layer normalization, linear layer and Sigmoid operation sequentially on the feature vector and / or the processed feature tensor, and output a weight vector. Branch 2, which runs parallel to branch 1, performs adaptive average pooling and Softmax normalization operations sequentially on the feature vector and / or the processed feature tensor, and outputs the weights. The feature vector and / or the processed feature tensor, weight vector, and weights are weighted and fused to output the fused feature vector, expressed as: , In the formula, The feature tensor after fusion; For feature vectors and / or processed feature tensors; This is a weighted fusion function; This is the weight vector; As weight.

9. The 3D human pose estimation method in a dynamic occlusion scene according to claim 1, characterized in that, Step S4 includes the following specific steps: a. Extract the original scale features of step S1 at Scale 1:1, and extract the 1 / 2 downsampled scale features and 1 / 4 downsampled scale features of step S3 at Scale 1:2 and Scale 1:4 to construct multi-scale spatiotemporal features; b. Input the multi-scale spatiotemporal features into the parallel spatiotemporal attention module and channel attention module, and output the multi-scale spatiotemporal attention enhanced feature map and the multi-scale channel attention enhanced feature map; c. For the multi-scale spatiotemporal features, their corresponding spatiotemporal attention enhancement feature maps and channel attention enhancement feature maps are concatenated along the channel dimension to output a multi-scale fusion feature map. d. The multi-scale fused feature map is spliced ​​along the channel dimension to form a cross-scale aggregated feature tensor; a 1×1 convolution operation is performed on the cross-scale aggregated feature tensor to realize cross-channel information fusion and feature dimension compression, and output the fused feature map. e. The fused feature map is added element-wise to the original scale feature map to generate the final output fused feature tensor.

10. A 3D human pose estimation method in a dynamic occlusion scene according to claim 9, characterized in that, The expression for the fused feature tensor is: , In the formula, To fuse feature tensors; For cross-scale splicing operations; Original scale features; Features at a 1 / 2 downsampling scale; Features at a 1 / 4 downsampling scale; The input features are the original scale.

Citation Information

Patent Citations

  • Sequential SAR terrain classification method based on multi-scale space-time self-attention

    CN120411790A

  • Lightweight human body posture estimation system and method fusing channel and space activation

    CN120580738A

  • Interactive behavior understanding method for posture reconstruction based on features of skeleton and image

    US20250022165A1

Cited By

  • Shielding perception vertex reasoning and semantic enhancement three-dimensional attitude estimation method and system

    CN121686575A