Human body posture estimation method based on dynamic attention and cross-branch dynamic fusion
By introducing dynamic attention and cross-branch dynamic fusion technology into the human posture estimation method, a model combining HRNet, CSAM, DAM and SimCC was constructed, and the problem of insufficient multi-resolution feature fusion and dynamic feature enhancement was solved, and high-precision and high-efficiency human posture estimation was achieved.
Patent Information
- Application Number
- CN202510094825.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing deep learning-based human pose estimation methods have shortcomings in multi-resolution feature fusion and dynamic feature enhancement, resulting in poor performance in human pose estimation, especially in complex poses and severe occlusion scenarios.
Using a human pose estimation method based on dynamic attention and cross-branch dynamic fusion, a key point prediction model combining HRNet, CSAM, DAM and SimCC is constructed to achieve efficient fusion and dynamic feature enhancement of multi-resolution features.
It significantly improves the accuracy and efficiency of human posture estimation, enhances the robustness and computing efficiency of the model, is suitable for real-time application scenarios, and provides an efficient and accurate human posture estimation solution.
Smart Images

Figure CN119992596A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human posture estimation, and in particular to a human posture estimation method based on dynamic attention and cross-branch dynamic fusion. Background Art
[0002] Human posture estimation is a key task in the field of computer vision and is widely used in scenarios such as motion capture, virtual reality, human behavior analysis, and human-computer interaction. With the development of deep learning technology, human posture estimation methods based on deep learning have become mainstream. For example, the Chinese invention patent application "A method for human posture estimation with small-scale perception enhancement" with application publication number CN 115830630 A first obtains a large number of pictures containing people and annotates the key point positions of the people in the pictures; then uses the obtained pictures and their annotation information to train the SSA-NET-based model framework, wherein the SSA-NET-based model framework includes a BackBone module for feature extraction of the input picture, a TAA module for feature enhancement of the extracted feature map, and a SimCC module for predicting the key point positions according to the enhanced feature map; then the picture to be estimated containing the person is input into the trained model, and the position coordinates of each key point of the person in the picture can be predicted. However, the SSA-NET-based model framework used in this method still faces the following challenges in practical applications: 1) Insufficient multi-resolution feature fusion: The BackBone module uses simple weighting or splicing operations to fuse feature map branches, fails to fully explore the complementarity between multi-resolution features, and lacks the ability to model complex associations between the global and local areas, especially in complex human postures or scenes with severe occlusion. 2) Insufficient dynamic feature enhancement: The TAA module uses fixed weights and cannot dynamically adjust the feature importance of each channel and spatial region according to the context of the input image. Background noise often interferes with key point prediction, resulting in reduced model robustness and accuracy. Summary of the invention
[0003] The present invention aims to solve the problem of poor human posture estimation effect caused by insufficient multi-resolution feature fusion and dynamic feature enhancement in existing human posture estimation methods based on deep learning, and provides a human posture estimation method based on dynamic attention and cross-branch dynamic fusion.
[0004] To solve the above problems, the present invention is achieved through the following technical solutions:
[0005] The human posture estimation method based on dynamic attention and cross-branch dynamic fusion includes the following steps:
[0006] Step 1: preprocessing the input image to obtain a preprocessed input image;
[0007] Step 2, the preprocessed input image is sent to the key point prediction model to obtain the predicted coordinates of the key points of the human body; the key point prediction model is composed of a multi-resolution feature extraction module, a cross-branch dynamic fusion module, a dynamic enhancement module and a key point coordinate classification module, the preprocessed input image is sent to the input of the multi-resolution feature extraction module, the output of the multi-resolution feature extraction module is connected to the input of the cross-branch dynamic fusion module, the output of the cross-branch dynamic fusion module is connected to the input of the dynamic enhancement module, the output of the dynamic enhancement module is connected to the input of the key point coordinate classification module, and the output of the key point coordinate classification module obtains the predicted coordinates of the key points of the human body;
[0008] Step 3: Map the predicted coordinates of the key points of the human body to the coordinate system of the original image, and connect the key points according to the topological structure of the human skeleton to generate a human posture diagram.
[0009] In the above scheme, the preprocessing of the input image includes normalization and image size adjustment.
[0010] In the above scheme, the specific processing process of the multi-resolution feature extraction module includes the following:
[0011] Step 2.1.1: Preprocessed input image I∈R W×H×3 After 3 convolution operations, the feature map F 1 ;
[0012] Step 2.1.2, feature map F 1 On the one hand, after two convolution operations, the feature map F is obtained 2 On the other hand, after downsampling and convolution operations, the feature map F is obtained 3 ;
[0013] Step 2.1.3, feature map F 2 On the one hand, after the convolution operation, the feature map F is obtained 4 On the other hand, after downsampling, the feature map F is obtained 5 ;
[0014] Step 2.1.4, feature map F 3 On the one hand, after upsampling, the feature map F is obtained 6 On the other hand, after the convolution operation, the feature map F is obtained 7 ;
[0015] Step 2.1.5: Feature map F 4 and feature map F 6 After the fusion operation, the feature map F is obtained 8 , feature map F 5 and feature map F 7 After the fusion operation, the feature map F is obtained 9;
[0016] Step 2.1.6, feature map F 8 After 4 convolution operations, the feature map F is obtained. 10 ;
[0017] Step 2.1.7. Feature map F 9 On the one hand, after 4 convolution operations, the feature map F is obtained 11 On the other hand, after downsampling and three convolution operations, the feature map F is obtained. 12 ;
[0018] Step 2.1.8, feature map F 10 On the one hand, after the convolution operation, the feature map F is obtained 13 On the other hand, after downsampling, the feature map F is obtained 14 ;
[0019] Step 2.1.9, feature map F 11 On the one hand, after upsampling, the feature map F is obtained 15 On the other hand, after the convolution operation, the feature map F is obtained 16 On the other hand, after downsampling, the feature map F is obtained 17 ;
[0020] Step 2.1.10, feature map F 12 On the one hand, after upsampling, the feature map F is obtained 18 On the other hand, after convolution, the feature map F is obtained 19 ;
[0021] Step 2.1.11. Feature map F 13 , feature map F 15 and feature map F 18 After the fusion operation, the feature map F is obtained 20 , feature map F 14 , feature map F 16 and feature map F 18 After the fusion operation, the feature map F is obtained 21 , feature map F 14 , feature map F 17 and feature map F 19 After the fusion operation, the feature map F is obtained 22 ;
[0022] Step 2.1.12, feature map F 20 After three convolution operations, the feature map F is obtained. 23 ;
[0023] Step 2.1.13. Feature map F 21 After three convolution operations, the feature map F is obtained. 24;
[0024] Step 2.1.14, feature map F 22 On the one hand, after three convolution operations, the feature map F is obtained 25 On the other hand, after downsampling and 2 convolution operations, the feature map F is obtained 26 ;
[0025] Step 2.1.15, feature map F 23 On the one hand, after the convolution operation, the feature map F is obtained 27 On the other hand, after downsampling, the feature map F is obtained 28 ;
[0026] Step 2.1.16, feature map F 24 On the one hand, after upsampling, the feature map F is obtained 29 On the other hand, after the convolution operation, the feature map F is obtained 30 On the other hand, after downsampling, the feature map F is obtained 31 ;
[0027] Step 2.1.17. Feature map F 25 On the one hand, after upsampling, the feature map F is obtained 32 On the other hand, after the convolution operation, the feature map F is obtained 33 On the other hand, after downsampling, the feature map F is obtained 34 ;
[0028] Step 2.1.18, feature map F 26 On the one hand, after upsampling, the feature map F is obtained 35 On the other hand, after convolution, the feature map F is obtained 36 ;
[0029] Step 2.1.19, feature map F 27 , feature map F 29 , feature map F 32 and feature map F 35 After the fusion operation, the feature map F is obtained high , feature map F 28 , feature map F 30 , feature map F 32 and feature map F 35 After the fusion operation, the feature map F is obtained mid , feature map F 28 , feature map F 31 , feature map F 33 and feature map F 35 After the fusion operation, the feature map F is obtained low ; Feature map F 28 , feature map F 31, feature map F 34 and feature map F 36 After the fusion operation, the feature map F is obtained lowest .
[0030] In the above scheme, the specific processing process of the cross-branch dynamic fusion module includes the following:
[0031] Step 2.2.1, for the four feature maps with different resolutions output by the resolution feature extraction module, adjust the three feature maps with lower resolutions to the same resolution as the feature map with the highest resolution by upsampling, and keep the feature map with the highest resolution unchanged, so as to obtain four feature maps with the same resolution;
[0032] Step 2.2.2, use 1×1 convolution to adjust the four feature maps of the same resolution to a uniform number of channels C, and obtain four response maps of the same channel; where C is the set number of channels;
[0033] Step 2.2.3: Use global average pooling to extract the global semantics of the four same channel response maps and obtain four global semantic vectors;
[0034] Step 2.2.4: Input the four global semantic vectors into the weight-sharing two-layer MLP respectively, and generate four dynamic weights through the Sigmoid activation function;
[0035] Step 2.2.5: Perform weighted summation on the 4 identical channel response maps according to the 4 dynamic weights to generate a fused feature map.
[0036] In the above scheme, the specific processing process of the dynamic enhancement module includes the following:
[0037] Step 2.3.1, perform global average pooling on the fusion feature map output by the cross-branch dynamic fusion module to obtain a channel response vector;
[0038] Step 2.3.2: Input the channel response vector into the weight-sharing two-layer MLP and generate channel weights through the Sigmoid activation function;
[0039] Step 2.3.3, apply the channel weight to each channel of the fused feature map, that is, perform weighted summation on each channel of the fused feature map according to the channel weight to obtain a channel enhanced feature map;
[0040] Step 2.3.4, perform global average pooling and global maximum pooling on the channel enhanced feature map to generate two single channel feature maps;
[0041] Step 2.3.5, concatenate the two single-channel feature maps in the channel dimension to obtain a concatenated feature map;
[0042] Step 2.3.6: The concatenated feature map is passed through a 7×7 convolution and a Sigmoid activation function to generate spatial weights;
[0043] Step 2.3.7: Apply the spatial weight to each channel of the channel enhancement feature map, that is, perform weighted summation on each pixel of the channel enhancement feature map according to the spatial weight to obtain the final dynamic enhancement feature map.
[0044] In the above scheme, the specific processing process of the key point coordinate classification module includes the following:
[0045] Step 2.4.1, use 1×1 convolution to convert the number of channels of the dynamic enhancement feature map from C to N to obtain the output feature map; where N is the set number of key points;
[0046] Step 2.4.2, divide the output feature map F into N groups along the channel dimension, each group contains a single channel feature map corresponding to a feature vector of a key point; then flatten the spatial dimension of the single channel feature map corresponding to each key point into a one-dimensional vector as the feature vector of the key point;
[0047] Step 2.4.3, for each key point's feature vector:
[0048] Use the horizontal coordinate classifier to map the feature vector of the key point to W 1 classification intervals, obtain the original output probability of the horizontal coordinate of the key point, and select the horizontal coordinate with the largest original output probability of the horizontal coordinate among the original output probabilities of the horizontal coordinate of the key point; at the same time, use the vertical coordinate classifier to map the feature vector of the key point to H 1 classification intervals, obtain the original output probability of the ordinate of the key point, select the ordinate with the maximum original output probability of the ordinate among the original output probabilities of the ordinate of the key point; then, output the selected abscissa and ordinate as the predicted coordinates of the key point;
[0049] The above W 1 represents the width of the dynamic enhancement feature map, H 1 Indicates the height of the dynamic enhancement feature map.
[0050] Compared with the prior art, the present invention has the following advantages:
[0051] 1. Construct a key point prediction model that combines HRNet (multi-resolution feature extraction module), CSAM (cross-branch dynamic fusion module), DAM (dynamic enhancement module) and SimCC (key point coordinate classification module). While significantly improving the accuracy of human posture estimation, it greatly reduces the computational complexity and is suitable for real-time application scenarios, providing a new solution for efficient and accurate human posture estimation.
[0052] 2. The cross-branch dynamic fusion module realizes efficient fusion of feature maps of different resolutions through a dynamic weighting mechanism, and fully exploits the complementarity of features of different resolutions by using channel alignment and global attention generation strategies;
[0053] 3. By combining channel attention and spatial attention, the dynamic enhancement module dynamically adjusts the importance of each channel and pixel area in the feature map, effectively suppressing the interference of background noise on key point prediction, while enhancing the feature expression ability of the key point area. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 Schematic diagram of the human pose estimation method based on dynamic attention and cross-branch dynamic fusion.
[0055] Figure 2 Schematic diagram of the multi-resolution feature extraction module.
[0056] Figure 3 The schematic diagram of the cross-branch dynamic fusion module.
[0057] Figure 4 This is the schematic diagram of the dynamic enhancement module.
[0058] Figure 5 This is the schematic diagram of the key point coordinate classification module. DETAILED DESCRIPTION
[0059] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in combination with specific examples and with reference to the accompanying drawings.
[0060] Human pose estimation methods based on dynamic attention and cross-branch dynamic fusion, such as Figure 1 As shown, it includes the following steps:
[0061] Step 1: Preprocessing the input image: preprocess the input image to obtain a preprocessed input image.
[0062] The input image is a color image H×W×3 in RGB format, where W is the width of the input image, H is the height of the input image, and 3 is the three RGB color channels of the input image.
[0063] Normalization: Scale the pixel values of the input image to the interval [0,1] to adapt to the input standard of the neural network. Assuming the original image is I, its normalization formula is:
[0064]
[0065] Among them, I min and I max are the minimum and maximum pixel values of the image, respectively.
[0066] Image resizing: The normalized input image is resized to a fixed size, such as 256×192 or 384×288, to meet the input requirements of subsequent models.
[0067] Step 2: Prediction of the predicted coordinates of the key points of the human body: The preprocessed input image is fed into the key point prediction model to obtain the predicted coordinates of the key points of the human body.
[0068] The key point prediction model of the present invention is composed of a multi-resolution feature extraction module, a cross-branch dynamic fusion module, a dynamic enhancement module and a key point coordinate classification module connected in series in sequence. The pre-processed input image is fed into the input of the multi-resolution feature extraction module, the output of the multi-resolution feature extraction module is connected to the input of the cross-branch dynamic fusion module, the output of the cross-branch dynamic fusion module is connected to the input of the dynamic enhancement module, the output of the dynamic enhancement module is connected to the input of the key point coordinate classification module, and the output of the key point coordinate classification module obtains the predicted coordinates of the key points of the human body.
[0069] Step 2.1, multi-resolution feature extraction module (HRNet), such as Figure 2 shown.
[0070] HRNet (High-Resolution Network) is a network framework widely used in human posture estimation. The HRNet backbone network consists of multiple parallel-connected branches of different resolutions, each of which contains multiple convolution modules. Through repeated feature fusion between branches of different resolutions (for example, adding the low-resolution features to the high-resolution features after upsampling, and adding the high-resolution features to the low-resolution features after downsampling), the high-resolution branches can continuously receive global context information from the low-resolution branches, thereby maintaining high resolution and retaining local details while also having rich global context information. The multi-resolution feature extraction module of the present invention directly adopts the backbone network of HRNet, and utilizes its parallel multi-branch structure and internal feature fusion mechanism to extract multi-resolution features.
[0071] The preprocessed input image I∈R WΔHΔ3 The multi-resolution feature extraction module extracts four feature maps with different resolutions through its unique multi-branch parallel architecture and internal feature fusion mechanism: and The high-resolution features retain the detailed information, and the low-resolution features contain richer global semantic vectors. W is the width of the input image, H is the height of the input image, and C m is the number of channels of the mth resolution feature map, m∈{high,mid,low,lowest}.
[0072] Step 2.1.1: Preprocessed input image I∈R W×H×3 After 3 convolution operations, the feature map F 1 ;
[0073] Step 2.1.2, feature map F 1 On the one hand, after two convolution operations, the feature map F is obtained 2 On the other hand, after downsampling and convolution operations, the feature map F is obtained 3 ;
[0074] Step 2.1.3, feature map F 2 On the one hand, after the convolution operation, the feature map F is obtained 4 On the other hand, after downsampling, the feature map F is obtained 5 ;
[0075] Step 2.1.4, feature map F 3 On the one hand, after upsampling, the feature map F is obtained 6 On the other hand, after the convolution operation, the feature map F is obtained 7 ;
[0076] Step 2.1.5: Feature map F 4 and feature map F 6 After the fusion operation, the feature map F is obtained 8 , feature map F 5 and feature map F 7 After the fusion operation, the feature map F is obtained 9 ;
[0077] Step 2.1.6, feature map F 8 After 4 convolution operations, the feature map F is obtained. 10 ;
[0078] Step 2.1.7. Feature map F 9 On the one hand, after 4 convolution operations, the feature map F is obtained 11 On the other hand, after downsampling and three convolution operations, the feature map F is obtained. 12 ;
[0079] Step 2.1.8, feature map F 10 On the one hand, after the convolution operation, the feature map F is obtained 13 On the other hand, after downsampling, the feature map F is obtained 14 ;
[0080] Step 2.1.9, feature map F 11 On the one hand, after upsampling, the feature map F is obtained 15 On the other hand, after the convolution operation, the feature map F is obtained16 On the other hand, after downsampling, the feature map F is obtained 17 ;
[0081] Step 2.1.10, feature map F 12 On the one hand, after upsampling, the feature map F is obtained 18 On the other hand, after convolution, the feature map F is obtained 19 ;
[0082] Step 2.1.11. Feature map F 13 , feature map F 15 and feature map F 18 After the fusion operation, the feature map F is obtained 20 , feature map F 14 , feature map F 16 and feature map F 18 After the fusion operation, the feature map F is obtained 21 , feature map F 14 , feature map F 17 and feature map F 19 After the fusion operation, the feature map F is obtained 22 ;
[0083] Step 2.1.12, feature map F 20 After three convolution operations, the feature map F is obtained. 23 ;
[0084] Step 2.1.13. Feature map F 21 After three convolution operations, the feature map F is obtained. 24 ;
[0085] Step 2.1.14, feature map F 22 On the one hand, after three convolution operations, the feature map F is obtained 25 On the other hand, after downsampling and 2 convolution operations, the feature map F is obtained 26 ;
[0086] Step 2.1.15, feature map F 23 On the one hand, after the convolution operation, the feature map F is obtained 27 On the other hand, after downsampling, the feature map F is obtained 28 ;
[0087] Step 2.1.16, feature map F 24 On the one hand, after upsampling, the feature map F is obtained 29 On the other hand, after the convolution operation, the feature map F is obtained 30 On the other hand, after downsampling, the feature map F is obtained 31 ;
[0088] Step 2.1.17. Feature map F 25 On the one hand, after upsampling, the feature map F is obtained 32 On the other hand, after the convolution operation, the feature map F is obtained 33 On the other hand, after downsampling, the feature map F is obtained 34 ;
[0089] Step 2.1.18, feature map F 26 On the one hand, after upsampling, the feature map F is obtained 35 On the other hand, after convolution, the feature map F is obtained 36 ;
[0090] Step 2.1.19, feature map F 27 , feature map F 29 , feature map F 32 and feature map F 35 After the fusion operation, the feature map F is obtained high , feature map F 28 , feature map F 30 , feature map F 32 and feature map F 35 After the fusion operation, the feature map F is obtained mid , feature map F 28 , feature map F 31 , feature map F 33 and feature map F 35 After the fusion operation, the feature map F is obtained low ; Feature map F 28 , feature map F 31 , feature map F 34 and feature map F 36 After the fusion operation, the feature map F is obtained lowest .
[0091] Step 2.2, cross-branch dynamic fusion module (CSAM), such as Figure 3 shown.
[0092] The CSAM module is integrated into the final stage of the HRNet backbone network, and dynamically weighted fuses the four feature maps of different resolutions output by the HRNet backbone network. Different from the original fusion mechanism of HRNet, CSAM uses channel alignment and global attention generation strategies to implement a dynamic weighting method, thereby more fully exploring the complementarity between features of different resolutions. In this way, CSAM generates a fusion of multi-scale information and the highest resolution feature map F high Feature maps F with the same resolution CSAM , which significantly improves the feature representation capability.
[0093] Step 2.2.1, resolution alignment: except for the feature map F high The other three feature maps F mid 、F mid and F lowest Through upsampling, it is uniformly adjusted to the feature map F high The same resolution is the highest resolution W / 4×H / 4, where the feature map F high Keep it unchanged and get 4 feature maps with the same resolution
[0094]
[0095] Step 2.2.2, channel alignment: adjust each resolution feature map F through 1×1 convolution m,aligned The number of channels is a uniform value C (C is the set value), and 4 identical channel response graphs F are obtained. m ' ,aligned ∈R W / 4×H / 4×C :
[0096] F′ m,aligned =Conv 1×1 (F m,aligned ),m∈{high,mid,low,lowest}
[0097] Step 2.2.3: Extract global semantic vector: Use global average pooling (GAP) to extract each feature map F m ' ,aligned The global semantics of , and obtain 4 one-dimensional global semantic vectors w of length C m :
[0098] w m = GAP(F′ m,aligned ),m∈{high,mid,low,lowest}
[0099] Step 2.2.4, dynamic weight generation: Each global semantic vector w m Input weight-sharing two-layer MLP (Multilayer Perceptron) and generate 4 dynamic weights through Sigmoid activation function
[0100]
[0101] in, and are the weight matrices of the two layers of MLP in the weight-sharing two-layer MLP, r is the channel compression ratio, C is the number of channels, ReLU is the activation function, and σ is the Sigmoid activation function.
[0102] Step 2.2.5: Feature weighted fusion: Based on 4 dynamic weights Response diagram F for 4 identical channels m ' ,aligned Perform weighted summation to generate fusion feature map F CSAM ∈R W / 4×H / 4×C :
[0103]
[0104] The generated fusion feature map F CSAM The global semantic information and detail expression capabilities are retained, and can be directly used in subsequent modules for accurate key point positioning.
[0105] Step 2.3, Dynamic Enhancement Module (DAM), such as Figure 4 shown.
[0106] The dynamic enhancement module further performs the fusion feature map F output by CSAM. CSAM Dynamic enhancement is performed. The channel features related to the key points are first enhanced using the channel attention mechanism, and then the features of the key point area are further enhanced using the spatial attention mechanism. By applying channel attention and spatial attention, the importance of each channel and pixel area in the feature map is dynamically adjusted, and the feature map is dynamically enhanced to effectively suppress the interference of background noise on key point prediction, while enhancing the feature expression ability of the key point area, and improving the accuracy and robustness of key point positioning.
[0107] Step 2.3.1, channel attention mechanism: This mechanism learns the importance weight of each channel by modeling the interdependence between feature map channels, thereby enhancing the channel features related to key point positioning and suppressing the responses of irrelevant channels.
[0108] Step 2.3.1.1. Average pooling generates channel response vector: The fusion feature map F is pooled by global average pooling. CSAM Perform global average pooling to obtain a one-dimensional channel response vector z of length C, where z corresponds to the response value z(c) of the cth channel, and the calculation formula is:
[0109]
[0110] Step 2.3.1.2, Multilayer Perceptron Generates Channel Weights: Input the channel response vector z into a weight-sharing two-layer MLP and generate channel weights w through the Sigmoid activation function 1 , where the channel weight w 1 The weight w corresponding to the cth channel 1 (c), the calculation formula is:
[0111] w1 (c) = σW 2 ReLU(W 1 z(c))),c=1,2,...C
[0112] in, and is the weight matrix of the two-layer MLP in the weight-sharing two-layer MLP, r is the channel compression ratio, C is the number of channels, ReLU is the activation function, and σ is the Sigmoid activation function.
[0113] Step 2.3.1.3: Apply weight to channel: Set channel weight w 1 Acting on the fusion feature map F CSAM On each channel of , that is, weighted summation is performed on each channel of the fusion feature map according to the channel weight to obtain the channel enhanced feature map F CA ∈R W / 4×H / 4×C , where the channel enhanced feature map F CA The eigenvalue F of the pixel (i, j) on the feature map of the cth channel CA (c,i,j) is:
[0114] F CA (c,i,j)=w 1 (c) F CSAM (c,i,j),i=1,2,...W / 4,j=1,2,...H / 4,c=1,2,...C
[0115] Step 2.3.2, spatial attention mechanism: This mechanism focuses on the importance of different spatial positions of the feature map, enhances the features of the area where the key points are located, and suppresses the interference of the background area.
[0116] Step 2.3.2.1: Pooling to generate a single channel feature map: Enhance the channel feature map F CA Perform global average pooling and global maximum pooling respectively to generate two single-channel feature maps F avg ∈R W / 4×H / 4 and F max ∈R W / 4×H / 4 :in
[0117] Single channel feature map F avg The eigenvalue F at pixel (i, j) avg (i,j) is:
[0118]
[0119] Single channel feature map F max The eigenvalue F at pixel (i, j) max (i,j) is:
[0120]
[0121] Step 2.3.2.2, Generate spatial weights: Substitute the two single-channel response vectors F avg and F max Splice in the channel dimension to get the spliced feature map F cat ∈R W / 4×H / 4×2 , where the concatenated feature map F cat The eigenvalue F at pixel (i, j) cat (i,j) is:
[0122] F cat (i,j)=Concat(F avg (i,j),F max (i,j)),i=1,2,...W / 4,j=1,2,...H / 4
[0123] Step 2.3.2.3: Concatenate the feature map F cat Through a 7×7 convolution and a Sigmoid activation function, the spatial weight w is generated. 2 , where the spatial weight w 2 The spatial weight value w at pixel (i, j) 2 (i,j) is:
[0124] w 2 (i,j)=σ(Conv 7×7 (F cat (i,j))),i=1,2,...W / 4,j=1,2,...H / 4
[0125] Step 2.3.2.4: Weights act on space: Apply spatial weights w 2 Acting on the channel enhancement feature map F CA On each channel of , that is, weighted summation is performed on each space (pixel point) of the channel enhancement feature map according to the spatial weight to obtain the dynamic enhancement feature map F DAM ∈R W / 4×H / 4×C , where the dynamic enhancement feature map F DAM The eigenvalue F at pixel (i, j) DAM (c,i,j) is:
[0126] F DAM (c,i,j)=w 2 (i,j)·F CA (c,i,j),i=1,2,...W / 4,j=1,2,...H / 4,c=1,2,...C
[0127] Step 2.4, key point coordinate classification module (SimCC), such as Figure 5 shown.
[0128] Considering that the heat map method relies on probability distribution to locate joint positions, it is difficult to achieve pixel-level accuracy, especially in scenes where the key points are closely spaced or occluded, which is prone to ambiguity and uncertainty, and the generation of high-resolution heat maps requires a lot of computing resources, especially when the number of key points is large. Subsequent analysis (such as maximum value search or coordinate regression) further increases the computational burden, which is not suitable for resource-constrained devices or real-time tasks. The present invention uses a key point coordinate classification module to replace the traditional heat map generation and analysis process, and uses a simple coordinate classification (SimCC) method to directly regress joint coordinates to achieve efficient and accurate joint positioning. Compared with the traditional heat map method, the key point coordinate classification module avoids the complex process of high-resolution heat map generation and analysis, achieves pixel-level precise positioning, and greatly improves the computational efficiency.
[0129] Step 2.4.1, channel adjustment: Through 1×1 convolution, the dynamic enhancement feature map F DAM The number of channels is converted from C to N, and the output feature map F∈R is obtained W / 4×H / 4×N :
[0130] F=Conv 1×1 (F DAM )
[0131] The above N is the number of key points set, and its value is set according to the dataset. For example, the COCO dataset defines 17 key points, and the MPII dataset defines 16 key points, which are located at the top of the head, neck, shoulders, elbows, wrists, hips, knees, ankles, etc.
[0132] Step 2.4.2, flattening features: Divide the output feature map F into N groups along the channel dimension, each group contains a single channel feature map corresponding to the feature vector of a key point. Then, flatten the spatial dimension (W×H) of the single channel feature map corresponding to each key point into a one-dimensional vector f n ∈R W / 4×H / 4 , as the feature vector of the key point.
[0133] f n =Flatten(F),n=1,2,...N
[0134] Step 2.4.3, coordinate prediction:
[0135] In order to predict the horizontal and vertical coordinates of each key point, two classifiers need to be designed respectively: the horizontal coordinate classifier and the vertical coordinate classifier. Both classifiers consist of a linear layer (fully connected layer) followed by a Softmax function activation layer. Assume that the feature map F DAM The width is W 1 and height is H 1 , where W 1 =W / 4,H 1 =H / 4, we divide the image into W 1 horizontal intervals and H 1 Each interval corresponds to a pixel coordinate, which is used to represent the horizontal and vertical coordinates of the key point.
[0136] For each key point n (n = 1, 2, ... N) feature vector f n Perform the following operations, where:
[0137] 1) Use the horizontal axis classifier to map to W 1 classification intervals, and obtain the original output probability of the horizontal coordinate of the key point n
[0138]
[0139] Use the ordinate classifier to map to H 1 classification interval, the original output probability of the vertical coordinate of the key point n
[0140]
[0141] Among them, W x and W y is the weight matrix of the linear layer, b x and b y is a bias term, and Softmax(·) is an activation function used to normalize the output to a probability distribution.
[0142] 2) Take the original output probability of the horizontal coordinate of the key point The horizontal axis with the largest horizontal axis original output probability in Take the original output probability of the vertical coordinate of the key point The ordinate with the largest ordinate original output probability in and will Output as the predicted coordinates of the key point.
[0143]
[0144] in, Represents the original output probability of the horizontal coordinate of key point n The original output probability of the horizontal coordinate at the horizontal coordinate i, The original output probability of the vertical coordinate of key point n The original output probability of the ordinate at ordinate j is,
[0145] Step 3: Posture generation: map the coordinates of each predicted key point to the original image coordinate system, and connect the key points according to the human skeleton topology to generate a human posture graph (skeleton graph), which can be directly used for subsequent analysis or display.
[0146] The present invention significantly improves the accuracy and efficiency of human key point positioning by introducing a cross-branch dynamic fusion module (CSAM) and a dynamic enhancement module (DAM) in combination with HRNet and SimCC technologies. The dynamic attention mechanism and multi-resolution dynamic fusion effectively reduce redundant calculations and significantly reduce the model inference time. The key point coordinates are directly regressed through SimCC to avoid the ambiguity introduced by the heat map and achieve pixel-level precise positioning. The method still performs well in complex scenes (such as severe occlusion and large changes in viewing angle) and has strong robustness.
[0147] It should be noted that although the embodiments of the present invention described above are illustrative, they are not intended to limit the present invention, and therefore the present invention is not limited to the above specific embodiments. Without departing from the principles of the present invention, any other embodiments obtained by those skilled in the art under the guidance of the present invention are deemed to be within the protection of the present invention.
Claims
1. A human pose estimation method based on dynamic attention and cross-branch dynamic fusion, characterized by: The steps include: Step 1: preprocessing the input image to obtain a preprocessed input image; Step 2, the preprocessed input image is sent to the key point prediction model to obtain the predicted coordinates of the key points of the human body; the key point prediction model is composed of a multi-resolution feature extraction module, a cross-branch dynamic fusion module, a dynamic enhancement module and a key point coordinate classification module, the preprocessed input image is sent to the input of the multi-resolution feature extraction module, the output of the multi-resolution feature extraction module is connected to the input of the cross-branch dynamic fusion module, the output of the cross-branch dynamic fusion module is connected to the input of the dynamic enhancement module, the output of the dynamic enhancement module is connected to the input of the key point coordinate classification module, and the output of the key point coordinate classification module obtains the predicted coordinates of the key points of the human body; Step 3: Map the predicted coordinates of the key points of the human body to the original image coordinate system, and connect the key points according to the topological structure of the human skeleton to generate a human posture diagram.
2. The method for human posture estimation based on dynamic attention and cross-branch dynamic fusion according to claim 1, characterized in that: The preprocessing of the input image includes normalization and image resizing.
3. The method for human posture estimation based on dynamic attention and cross-branch dynamic fusion according to claim 1, characterized in that: The specific processing of the multi-resolution feature extraction module includes the following: Step 2.1.1: Preprocessed input image I∈R W×H×3 After 3 convolution operations, the feature map F 1 ; Step 2.1.2, feature map F 1 On the one hand, after two convolution operations, the feature map F is obtained 2 On the other hand, after downsampling and convolution operations, the feature map F is obtained 3 ; Step 2.1.3, feature map F 2 On the one hand, after the convolution operation, the feature map F is obtained 4 On the other hand, after downsampling, the feature map F is obtained 5 ; Step 2.1.4, feature map F 3 On the one hand, after upsampling, the feature map F is obtained 6 On the other hand, after the convolution operation, the feature map F is obtained 7 ; Step 2.1.5: Feature map F 4 and feature map F 6 After the fusion operation, the feature map F is obtained 8 , feature map F 5 and feature map F 7 After the fusion operation, the feature map F is obtained 9 ; Step 2.1.6, feature map F 8 After 4 convolution operations, the feature map F is obtained. 10 ; Step 2.1.
7. Feature map F 9 On the one hand, after 4 convolution operations, the feature map F is obtained 11 On the other hand, after downsampling and three convolution operations, the feature map F is obtained. 12 ; Step 2.1.8, feature map F 10 On the one hand, after the convolution operation, the feature map F is obtained 13 On the other hand, after downsampling, the feature map F is obtained 14 ; Step 2.1.9, feature map F 11 On the one hand, after upsampling, the feature map F is obtained 15 On the other hand, after the convolution operation, the feature map F is obtained 16 On the other hand, after downsampling, the feature map F is obtained 17 ; Step 2.1.10, feature map F 12 On the one hand, after upsampling, the feature map F is obtained 18 On the other hand, after convolution, we get the feature map F 19 ; Step 2.1.
11. Feature map F 13 , feature map F 15 and feature map F 18 After the fusion operation, the feature map F is obtained 20 , feature map F 14 , feature map F 16 and feature map F 18 After the fusion operation, the feature map F is obtained 21 , feature map F 14 , feature map F 17 and feature map F 19 After the fusion operation, the feature map F is obtained 22 ; Step 2.1.12, feature map F 20 After three convolution operations, the feature map F is obtained. 23 ; Step 2.1.
13. Feature map F 21 After three convolution operations, the feature map F is obtained. 24 ; Step 2.1.14, feature map F 22 On the one hand, after three convolution operations, the feature map F is obtained 25 On the other hand, after downsampling and 2 convolution operations, the feature map F is obtained 26 ; Step 2.1.15, feature map F 23 On the one hand, after the convolution operation, the feature map F is obtained 27 On the other hand, after downsampling, the feature map F is obtained 28 ; Step 2.1.16, feature map F 24 On the one hand, after upsampling, the feature map F is obtained 29 On the other hand, after the convolution operation, the feature map F is obtained 30 On the other hand, after downsampling, the feature map F is obtained 31 ; Step 2.1.
17. Feature map F 25 On the one hand, after upsampling, the feature map F is obtained 32 On the other hand, after the convolution operation, the feature map F is obtained 33 On the other hand, after downsampling, the feature map F is obtained 34 ; Step 2.1.18, feature map F 26 On the one hand, after upsampling, the feature map F is obtained 35 On the other hand, after convolution, we get the feature map F 36 ; Step 2.1.19, feature map F 27 , feature map F 29 , feature map F 32 and feature map F 35 After the fusion operation, the feature map F is obtained high , feature map F 28 , feature map F 30 , feature map F 32 and feature map F 35 After the fusion operation, the feature map F is obtained mid , feature map F 28 , feature map F 31 , feature map F 33 and feature map F 35 After the fusion operation, the feature map F is obtained low ; Feature map F 28 , feature map F 31 , feature map F 34 and feature map F 36 After the fusion operation, the feature map F is obtained lowest .
4. The method for human posture estimation based on dynamic attention and cross-branch dynamic fusion according to claim 1, characterized in that: The specific processing process of the cross-branch dynamic fusion module includes the following: Step 2.2.1, for the four feature maps with different resolutions output by the resolution feature extraction module, adjust the three feature maps with lower resolutions to the same resolution as the feature map with the highest resolution by upsampling, and keep the feature map with the highest resolution unchanged, so as to obtain four feature maps with the same resolution; Step 2.2.2, use 1×1 convolution to adjust the four feature maps of the same resolution to a uniform number of channels C, and obtain four response maps of the same channel; where C is the set number of channels; Step 2.2.3: Use global average pooling to extract the global semantics of the four same channel response maps and obtain four global semantic vectors; Step 2.2.4: Input the four global semantic vectors into the weight-sharing two-layer MLP respectively, and generate four dynamic weights through the Sigmoid activation function; Step 2.2.5: Perform weighted summation on the 4 identical channel response maps according to the 4 dynamic weights to generate a fused feature map.
5. The method for human posture estimation based on dynamic attention and cross-branch dynamic fusion according to claim 1, characterized in that: The specific processing process of the dynamic enhancement module includes the following: Step 2.3.1, perform global average pooling on the fusion feature map output by the cross-branch dynamic fusion module to obtain a channel response vector; Step 2.3.2: Input the channel response vector into the weight-sharing two-layer MLP and generate channel weights through the Sigmoid activation function; Step 2.3.3, apply the channel weight to each channel of the fused feature map, that is, perform weighted summation on each channel of the fused feature map according to the channel weight to obtain a channel enhanced feature map; Step 2.3.4, perform global average pooling and global maximum pooling on the channel enhanced feature map to generate two single channel feature maps; Step 2.3.5, concatenate the two single-channel feature maps in the channel dimension to obtain a concatenated feature map; Step 2.3.6: The concatenated feature map is passed through a 7×7 convolution and a Sigmoid activation function to generate spatial weights; Step 2.3.7: Apply the spatial weight to each channel of the channel enhancement feature map, that is, perform weighted summation on each pixel of the channel enhancement feature map according to the spatial weight to obtain the final dynamic enhancement feature map.
6. The method for human posture estimation based on dynamic attention and cross-branch dynamic fusion according to claim 1, characterized in that: The specific processing process of the key point coordinate classification module includes the following: Step 2.4.1, use 1×1 convolution to convert the number of channels of the dynamic enhancement feature map from C to N to obtain the output feature map; where N is the set number of key points; Step 2.4.2, divide the output feature map F into N groups along the channel dimension, each group contains a single channel feature map corresponding to a feature vector of a key point; then flatten the spatial dimension of the single channel feature map corresponding to each key point into a one-dimensional vector as the feature vector of the key point; Step 2.4.3, for each key point's feature vector: The feature vector of the key point is mapped to W1 classification intervals by using the horizontal coordinate classifier, and the original output probability of the horizontal coordinate of the key point is obtained, and the horizontal coordinate with the maximum original output probability of the horizontal coordinate is selected from the original output probability of the horizontal coordinate of the key point; at the same time, the feature vector of the key point is mapped to H1 classification intervals by using the vertical coordinate classifier, and the original output probability of the vertical coordinate of the key point is obtained, and the vertical coordinate with the maximum original output probability of the vertical coordinate is selected from the original output probability of the vertical coordinate of the key point; Afterwards, the selected horizontal and vertical coordinates are output as the predicted coordinates of the key point; The above W1 represents the width of the dynamic enhancement feature map, and H1 represents the height of the dynamic enhancement feature map.
Citation Information
Patent Citations
Small-scale perception enhanced human body posture estimation method
CN115830630A
High-resolution lightweight human body posture estimation method combined with multispectral attention mechanism
CN113792641A
Human body posture estimation method based on dynamic information transmission
CN114299537A
Multi-branch 2D human body posture estimation method based on attention mechanism
CN116844186A
Method and system for detecting fundus image based on dynamic weighted attention mechanism
US20230377147A1