Multi-user attitude estimation method based on multilevel feature enhancement and cooperative calibration

By constructing the MLCPose model, combining the BSCA module and the MLTCC mechanism, the problem of difficult to balance the real-time and accuracy of human pose estimation in complex environments is solved, and high-precision and stable pose estimation in multi-person scenarios and occlusion conditions is achieved.

CN119964247APending Publication Date: 2025-05-09XINJIANG INST OF ENG
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510113190.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art faces challenges such as lighting changes, partial occlusion, dynamic backgrounds and multi-person scenes in human posture estimation in complex environments, making it difficult to balance real-time and precision.

Method used

A multi-person pose estimation method with multi-level feature enhancement and collaborative calibration is proposed. By constructing an MLCPose model, using backbone network, PAFPN and head modules, combining BSCA module and MLTCC mechanism, the global modeling and local detail capture capabilities are improved, and the stability and robustness of the model are enhanced through dynamic weight optimization and outlier point-aware loss function.

Benefits of technology

It realizes improving the positioning accuracy and stability of pose estimation in complex scenarios, balancing real-time and accuracy, and effectively dealing with multi-person scenarios and occlusion problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964247A_ABST
    Figure CN119964247A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-level feature enhancement and cooperative calibration multi-person attitude estimation method, which comprises an MLCpose model, and is characterized in that the MLCpose model comprises a backbone network, a PAFPN and a Head; the backbone network takes a multi-layer CSPLayer and a convolution module as a core, fuses a BSCA module, and improves global modeling and local detail capture capability through coordinate attention and bidirectional space attention; the PAFPN combines the high-level semantic features and the low-level space detail features through remodeling and connection operations; in the Head, an X-axis coordinate classifier and a Y-axis coordinate classifier are adopted to accurately position the coordinates of the key points, meanwhile, a dynamic MLTCC mechanism is introduced, and accurate positioning of the key points of each individual is predicted. According to the method, the feature reduction and boundary response capability of the occlusion area is improved through local detail reconstruction, noise interference is suppressed by dynamically recognizing the occlusion and labeling error and other outliers, and the positioning precision and stability of the model in a multi-target scene are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and in particular to a multi-person posture estimation method of multi-level feature enhancement and collaborative calibration. Background Art

[0002] Human pose estimation is a basic task in the field of computer vision, and is extended to other fields such as motion analysis, virtual reality, human-computer interaction, and intelligent monitoring. However, pose estimation in complex environments still faces many challenges such as illumination changes, partial occlusions, dynamic backgrounds, and multi-person scenes. Illumination changes can lead to inconsistencies in human texture features, affecting the model's feature extraction capabilities; occlusion problems often lead to missing key point information, reducing prediction accuracy; and dynamic backgrounds increase the difficulty of human region segmentation and positioning. In multi-person scenes, the detection and association of key points further increases the computational complexity, especially on resource-constrained devices, where real-time and accuracy are difficult to achieve at the same time. Although the heat map method and coordinate regression framework perform well in academic benchmarks, their real-time performance in complex scenes still cannot meet actual needs.

[0003] The current mainstream feature extraction modules have a significant bias in modeling local details and global dependencies, making it difficult to effectively capture the features of both at the same time. This problem is particularly evident in multi-person interaction and occlusion scenes. For example, when dealing with complex interactions, PifPaf and OpenPose have insufficient spatial relationship modeling capabilities and weak global geometric relationship capture capabilities, respectively, resulting in reduced key point prediction accuracy in multi-person scenes. Although HRNet significantly improves overall performance through high-resolution features, its high computational cost limits its application in real-time scenarios. While lightweight networks such as Lite-HRNet and LitePose have made some progress in improving real-time performance, the weakening of global feature modeling capabilities makes it difficult to ensure accuracy in complex scenes. Recent new methods such as LPSNet and MVSplat have made breakthroughs in multi-view detail modeling and global geometric modeling, respectively, but still face problems such as high computational cost and limited detail capture capabilities.

[0004] Existing models are not stable enough when facing complex probability distributions and abnormal samples (such as occlusion, annotation errors, etc.), and are easily disturbed by noise, which reduces the overall performance. For example, SimCC treats human posture estimation as a one-dimensional classification task, eliminating the expensive upsampling layer. However, its anti-interference ability is weak and it is easily affected by abnormal samples. Recently, although CineMPC combines the capabilities of zoom, focus, and scene composition optimization, its adaptability in complex environments is still insufficient. GDE-Pose improves performance through multi-scale feature fusion, but it is still difficult to ensure robustness in complex backgrounds. The hybrid spatiotemporal architecture introduced by PoseMagic enhances the stability of the model, but delays or inaccurate predictions may still occur in fast movements and drastic posture changes. Summary of the invention

[0005] In order to overcome the above problems existing in the prior art, the present invention proposes a multi-person posture estimation method with multi-level feature enhancement and collaborative calibration.

[0006] The technical solution adopted by the present invention to solve the technical problem is: a multi-person posture estimation method with multi-level feature enhancement and collaborative calibration, comprising the following steps: Step 1, construct an MLCPose model and train the constructed MLCPose model through the data set; Step 2, input the video stream data or the collected image to be predicted into the MLCPose model obtained in step 1, and the MLCPose model outputs an estimation result graph; MLCPose model, the MLCPose model includes a backbone network, PAFPN, and Head, the backbone network is used for feature extraction, the PAFPN is used for multi-scale fusion and optimization, and the Head is used for key point prediction; The backbone network is based on multi-layer CSPLayer and convolution modules, and integrates BSCA modules to improve global modeling and local detail capture capabilities through coordinate attention and bidirectional spatial attention; the PAFPN combines high-level semantic features and low-level spatial detail features through reshaping and connection operations; the Head uses X-axis and Y-axis coordinate classifiers to accurately locate the coordinates of key points, and introduces a dynamic MLTCC mechanism to predict the precise positioning of each individual key point.

[0007] The above-mentioned multi-person posture estimation method with multi-level feature enhancement and collaborative calibration, the BSCA module includes a coordinate attention mechanism and a bidirectional spatial attention mechanism, the coordinate attention mechanism is used to enhance the network to handle highly spatially correlated tasks, and the coordinate attention mechanism workflow includes coordinate information embedding and coordinate attention generation; the bidirectional spatial attention mechanism is used to enhance the model's responsiveness to key areas, and the bidirectional spatial attention mechanism workflow includes feature separation and attention generation.

[0008] The above-mentioned multi-level feature enhancement and collaborative calibration method for multi-person posture estimation, wherein the coordinate information embedding specifically includes: inputting a feature map , where C, H and W represent the number of channels, height and width respectively; perform one-dimensional global average pooling on the feature map in the height and width directions respectively, extract the features in the horizontal and vertical directions, and obtain the height Z h and width Z w Eigenvectors in directions; The coordinate attention generation specifically includes: and The features are concatenated and fully connected across channels through a 1×1 convolution layer, batch normalization is applied, and nonlinear characteristics are enhanced through nonlinear activation functions: ; Among them, Concat(·) means concatenation along the feature dimension. yes The convolution operation, BN(·) represents the batch normalization operation, ReLU(·) is the rectified linear unit activation function, and the obtained features ; Will Split into horizontal and vertical Two parts, apply Sigmoid activation function to each part to generate attention weights: ; Where σ(·) is the Sigmoid function, and the generated attention weight ; The generated attention weights are multiplied element-wise and Applied to the input feature map X: ; Among them, Y is the output feature map, .

[0009] The above-mentioned multi-person posture estimation method with multi-level feature enhancement and collaborative calibration, wherein the feature separation comprises: applying different convolution kernels to the input feature map X in the horizontal and vertical dimensions respectively, and independently extracting the features in the horizontal direction and vertical characteristics ; The attention generation includes: applying the Sigmoid activation function to the horizontal features and vertical characteristics Processing to generate a horizontal attention map And vertical attention map , Emphasize or suppress pixels with important edge information in the horizontal direction. Focus on important features in the vertical direction; horizontal and vertical attention maps are used to reweight the spatial features of the input feature map X, and finally output the feature map .

[0010] The above-mentioned multi-person posture estimation method with multi-level feature enhancement and collaborative calibration, the working process of the MLTCC mechanism is: the input batch data samples pass through the SimCC detection head to generate the predicted key point probability distribution and the real key point probability distribution; the loss calculation module combines the KL discrete distribution loss and the outlier perception loss, generates dynamic weights through an adaptive weight strategy, and optimizes the loss of each layer layer by layer; the weights are balanced through layer-by-layer normalization, and the high-confidence key point areas are optimized first to reduce noise interference; the joint loss function is used to optimize the model, which significantly improves the positioning accuracy and stability in the posture estimation task.

[0011] The above-mentioned multi-level feature enhancement and collaborative calibration multi-person pose estimation, the dynamic weight The calculation formula is: ; in, Used to normalize weights to ensure stability in the value space; ω i represents the contribution basis of the loss layer, Indicates the confidence of each key point; The joint loss function is specifically: ; in, Represents the sum of sub-losses from layer i to layer K; represents the sum of sub-losses from layer i to layer A, N represents the total number of layers, is a dynamic weight used to balance the loss of each layer.

[0012] In the above-mentioned multi-person posture estimation method with multi-level feature enhancement and collaborative calibration, the outlier perception loss introduces an adaptive adjustment module to dynamically correct the outliers when allocating loss weights; the loss function in the outlier perception loss is: ; in, and are the predicted distribution and the true target distribution, respectively, K is the number of key points, N is the total number of layers, and n is the layer reduction; is a dynamic weight factor, which is defined by the outlier perception mechanism as: ; M represents the set of key points that need to be shielded, and η is the scaling factor of the shielding points; the step size adjustment strategy is implemented by the following formula: ; Among them, α∈(0,1) is used to smooth high loss values, β>1 is used to enlarge the optimization step of normal values, and τ is the outlier detection threshold.

[0013] The beneficial effect of the present invention is that the present invention designs a dual-scale collaborative attention (BSCA) module, which realizes global dependency and local detail optimization. BSCA strengthens the global feature expression by capturing the long-range spatial relationship between key points, and improves the feature restoration and boundary response capabilities of the occluded area by reconstructing local details.

[0014] For the SimCC detection head, a multi-layer dynamic optimization (MLT) mechanism was developed, which was designed from the perspectives of high and low confidence and sparse and dense targets. By dynamically adjusting the loss weights of different scales, high confidence areas are optimized first, and a global normalization strategy is used to balance the feature distribution of sparse targets and dense targets, effectively improving the positioning accuracy and stability of the model in multi-target scenarios.

[0015] An outlier-aware loss function based on anomaly perception is proposed. By dynamically identifying outliers such as occlusion and annotation errors, the learning weights of credible samples are enhanced and noise interference is suppressed in combination with a weighted optimization strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic diagram of the MLCPose model framework of the present invention; Figure 2 is a schematic diagram of the BSCA module of the present invention; Figure 3 It is a schematic diagram of the MLTCC mechanism of the present invention; Figure 4The figure is a schematic diagram of the outlier perception loss of the present invention; wherein (a) describes in detail the core mechanism of outlier detection and step size optimization strategy; (b) demonstrates the efficiency and adaptability of the loss calculation process through modular steps, including weight initialization, KL divergence calculation and label matching; (c) is the overall process, emphasizing the mask mechanism of shielding key points to optimize the weights of different joints in a fine-grained manner; Figure 5 It is a schematic diagram of the gradual improvement design of the SimCC baseline model according to the present invention; Figure 6 Schematic diagram of comparison results of three different treatment methods of the present invention; Figure 7 It is a MLC Pose visualization diagram of human posture estimation in sports competition under different scenarios of the present invention. DETAILED DESCRIPTION

[0017] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0018] like Figure 1 As shown, this embodiment discloses a multi-person posture estimation method with multi-level feature enhancement and collaborative calibration, comprising the following steps: Step 1, construct an MLCPose model and train the constructed MLCPose model through the data set; Step 2: Input the video stream data or the collected image to be predicted into the MLCPose model obtained in step 1, and the MLCPose model outputs an estimation result graph.

[0019] MLCPose model, the MLCPose model includes a backbone network, PAFPN, and Head, the backbone network is used for feature extraction, the PAFPN is used for multi-scale fusion and optimization, and the Head is used for key point prediction; The backbone network is based on multi-layer CSPLayer and convolution modules, and integrates BSCA modules to improve global modeling and local detail capture capabilities through coordinate attention and bidirectional spatial attention; the PAFPN combines high-level semantic features and low-level spatial detail features through reshaping and connection operations; the Head uses X-axis and Y-axis coordinate classifiers to accurately locate the coordinates of key points, and introduces a dynamic MLTCC mechanism to predict the precise positioning of each individual key point.

[0020] The BSCA module is introduced to improve the feature expression ability and position accuracy of the key point detection model in complex spatial environments. For key point positioning tasks, the ability to capture spatial information is crucial, especially in the process of extracting fine-grained textures and spatial features, which determines the depth of the model's semantic understanding of human structure. The BSCA attention mechanism module is embedded behind the CSP interception network, which strengthens the long-range dependency between key points and significantly improves the ability to resolve texture details.

[0021] BSCA modules such as Figure 2 As shown in the figure, it includes the coordinate attention mechanism (CA) and the bidirectional spatial attention mechanism (BSA). The key point positioning is highly sensitive to spatial information. The coordinate attention (CA) mechanism is used to enhance the network to handle tasks with high spatial correlation. It operates through two main workflows: the coordinate information embedding process and the coordinate attention generation process.

[0022] Coordinate information embedding specifically includes: In the coordinate information embedding stage, input feature map , where C, H, and W represent the number of channels, height, and width, respectively. First, one-dimensional global average pooling (GAP) is performed on the feature map in the height and width directions to extract features in the horizontal and vertical directions. The process is defined as: ; in, and are the eigenvectors in the height and width directions respectively.

[0023] Coordinate attention generation specifically includes: after these features are extracted, and The features are concatenated and fully connected across channels through a 1×1 convolution layer, followed by a batch normalization (BN) operation, and finally the nonlinear characteristics are enhanced through a nonlinear activation function: ; Among them, Concat(∙) means concatenation along the feature dimension. yes The convolution operation, BN (∙) represents the batch normalization operation, ReLU (∙) is the rectified linear unit activation function, and the obtained features .

[0024] Then will Split into horizontal and vertical Two parts, apply Sigmoid activation function to each part to generate attention weights: ; Where σ(·) is the Sigmoid function, and the generated attention weight .

[0025] Finally, the generated attention weights are multiplied element-wise (⊙) and Applied to the input feature map X: ; in, To output feature maps, it combines horizontal and vertical attention information, significantly improving the model's responsiveness to key spatial locations. The integration of the CA mechanism improves the network's sensitivity to spatial details, which can effectively avoid feature distortion caused by resizing the network, while optimizing the shape of the feature map output by the network to adapt to the input requirements of the coordinate classification detection head.

[0026] The Bidirectional Spatial Attention (BSA) mechanism is designed to improve the performance of neural networks in processing visual tasks with complex spatial structures. The implementation of the BSA mechanism involves two main workflows: the feature separation process and the attention generation process.

[0027] Feature separation includes: In the feature separation stage, the bidirectional spatial attention (BSA) mechanism extracts horizontal and vertical information independently by performing specific operations on the input feature map X. This process is implemented through a directed convolution operation, that is, different convolution kernels are applied to the input feature map X in the horizontal and vertical dimensions respectively to independently extract horizontal features. and vertical characteristics , thereby enhancing the network's sensitivity to directional features and capturing richer spatial information. The separated feature maps represent the spatial features in the horizontal and vertical directions respectively: ; in, Represent the convolution operations performed in the horizontal and vertical directions, respectively, to extract spatial details in the corresponding directions.

[0028] Attention generation includes: In the attention generation stage, the and is further processed to generate a refined attention map. This is achieved by applying a Sigmoid activation function, which maps feature values ​​to the interval [0,1] to represent the importance of each pixel position: ; Among them, sigma represents the Sigmoid function, and The generated horizontal and vertical attention maps are respectively. These attention maps dynamically adjust the weights of each pixel position in the input feature map X. Specifically, Emphasize or suppress pixels with important edge information in the horizontal direction, such as lateral features of the torso, arms, or legs; The horizontal and vertical attention weights are used to reweight the spatial features of the input feature map. The specific reweighting process is as follows: ; Among them, ⊙ represents element-by-element multiplication. The final output feature map Y incorporates the attention modulation from two spatial directions, enhancing the network's responsiveness to key areas, thereby showing higher accuracy and efficiency in tasks that require high spatial sensitivity, such as object detection and key point detection.

[0029] By embedding the BSCA module into the CSP truncation module and combining the residual unit and linear projection design, the dimension of the input embedding is increased, allowing the model to adapt to a higher-dimensional feature space and thus capture finer and more complex features.

[0030] The MLTCC mechanism is an MLT dynamic optimization mechanism for the SimCC detection head. The probability distribution space of key points does not obey a single distribution, but exhibits complex multimodal distribution characteristics. The SimCC algorithm often inevitably leads to the loss of spatial information during the coordinate encoding process, which in turn affects the positioning accuracy of the model. To this end, this embodiment proposes a multi-layer dynamic optimization (MLT) mechanism for the SimCC detection head to impose multiple constraints on the key point positioning process of the model. The MLTCC mechanism is as follows: Figure 3 As shown in the figure, the working process is as follows: the input batch data samples pass through the SimCC detection head to generate the predicted key point probability distribution and the real key point probability distribution; the loss calculation module combines the KL discrete distribution loss and the outlier perception loss, generates dynamic weights through an adaptive weight strategy, and optimizes the loss of each layer layer by layer; the weights are balanced through layer-by-layer normalization, and the high-confidence key point areas are optimized first to reduce noise interference; the joint loss function is used to optimize the model, which significantly improves the positioning accuracy and stability in the posture estimation task.

[0031] The MLT mechanism dynamically allocates loss layer weights , combined with the initial weight ωi of the loss layer and the confidence of the key points , achieving fine-tuning of multi-layer losses. Dynamic weights The calculation formula is: ; in, Used to normalize weights to ensure stability in the value space; ω i represents the contribution basis of the loss layer, The confidence of each key point is expressed by prioritizing the key points in the high confidence area. By regulating the importance of different levels, it can effectively reduce the noise interference of the low confidence area on the model training, while not ignoring the potential weak key point information.

[0032] The specific joint loss function is: ; in, Represents the sum of sub-losses from layer i to layer K; represents the sum of sub-losses from layer i to layer A, N represents the number of layers, is a dynamic weight used to balance the loss of each layer.

[0033] Global normalization: To avoid imbalance in weight distribution, the MLT mechanism introduces a global normalization strategy.

[0034] ; Normalized weights This ensures that multiple layers of losses contribute evenly to the overall optimization objective, avoiding gradient instability or offset.

[0035] The outlier-aware loss significantly improves the accuracy of key point positioning and the ability to handle outliers by introducing dynamic weight adjustment and outlier correction mechanisms. The proposed loss function combines the adaptive adjustment mechanism with the high-dimensional feature weight optimization strategy, which enhances the dynamic response of the loss function to outliers in theory and practice, making it converge effectively during the optimization process, thereby reducing the negative impact of abnormal data on model performance. The mathematical expression of the loss function is: ; in, and are the predicted distribution and the true target distribution, respectively. is a dynamic weight factor, which is defined by the outlier perception mechanism as: ; M represents the set of key points that need to be shielded, and η is the scaling factor of the shielding points. The step size adjustment strategy is implemented by the following formula: ; Among them, α∈(0,1) is used to smooth high loss values, β>1 is used to enlarge the optimization step of normal values, and τ is the outlier detection threshold. Figure 4Part (a) intuitively demonstrates the operation of the outlier detection and step adjustment modules. By combining the predicted loss value and the dynamic adjustment factor, it realizes the layer-by-layer correction of different abnormality levels. Part (b) optimizes the model's adaptability to complex distribution features through the distribution alignment and weight initialization strategy of KL divergence. Part (c) introduces the key point shielding mechanism and the final loss normalization strategy to improve the ability to capture local errors.

[0036] The outlier-aware loss function effectively improves the model's adaptability under complex distributions through dynamic mask correction and outlier-aware mechanisms. Multi-scale feature alignment and optimized design for outliers enable it to show higher accuracy and model robustness when dealing with abnormal data interference.

[0037] The MLT mechanism can effectively avoid target rigidity and is particularly suitable for dynamic scene tasks such as video action recognition. In conjunction with the MLC Pose real-time target control strategy, the loss weight can be dynamically adjusted according to the confidence of the key points in each frame of the video to adapt to the temporal changes and spatial complexity of the action. Through multi-level optimization and normalization, the MLT mechanism balances the contribution of temporal information and spatial features, ensuring the synergy of different loss functions during training and avoiding the limitation of the model's expressive power by a single optimization target.

[0038] Figure 5 The paper demonstrates the gradual improvement design for the SimCC baseline model, combining refined parameter tuning and strategy optimization, and exploring the model's potential through multi-dimensional fine-tuning, which significantly improves the model's performance on the MPIIPCKh@0.5 index, ultimately achieving an excellent performance of 88.80. From learning rate adjustment to data enhancement strategy, each improvement has achieved significant gains through theoretical drive and experimental verification.

[0039] The AdamW optimizer is the best optimizer for the MLCPose model proposed in this embodiment. It effectively suppresses overfitting by introducing a weight decay mechanism and is more suitable for neural network training. Its update formula is: ; in, represents the parameter value of the tth round, η is the learning rate, and are the first-order momentum and second-order momentum estimates of the gradient, ϵ is a numerical stability term used to avoid the denominator being zero, and λ is the weight decay factor. −3 The AdamW optimizer performs best under the setting of . Typically, the parameters and bias terms of the BatchNorm layer do not need to be decayed, and the configuration avoids applying the same configuration to repeated parameters, which helps improve efficiency.

[0040] In data enhancement, sigma in SimCC sets a fixed value to control the intensity of Gaussian blur, and simcc_split_ratio=2.0 is the split ratio of SimCC. This paper uses the dual-channel sigma parameter (5.66, 5.66) to perform more precise control on multiple channels. The Gaussian distribution formula is: ; Among them, the position in the heat map The intensity value of , )Coordinates of key points; : Gaussian kernel width. Optimized Improved the ability to capture local features.

[0041] The EMA mechanism can capture long-term dependencies in time series. The SimCC output sequence can be directly post-processed to improve the smoothness and consistency of the time series. Its update formula is: ; in, is the parameter value after sliding average, is the parameter value of the current round, is the decay coefficient of the historical weight (0.99). The introduction of EMA effectively smoothes the parameter fluctuations during the training process.

[0042] The MLCPose model proposed in this embodiment is verified by the MPII dataset and compared with the performance and estimation results of three major models: 2D heat map, regression and Coord_cls_heads. The MPII dataset is widely used in the study of human posture estimation and behavior recognition tasks. The dataset contains about 24,987 natural scene images, covering complex human postures and diverse backgrounds in a variety of daily activities.

[0043] Each image is annotated with 16 key points (such as head, shoulders, elbows, wrists, hips, knees and ankles), and the key point annotations include 2D coordinates and visibility information. This information can help the model more effectively deal with complex situations such as occlusion and incomplete posture.

[0044] The dataset is divided into training set, validation set and test set: the training set contains 17,408 images, and the validation set and test set contain 7,579 images. The dataset contains a variety of activity categories (such as walking, running, jumping, weightlifting, etc.) and a variety of shooting perspectives (including front, side and oblique angles), which can effectively support the performance improvement of the model in different scenes and perspectives.

[0045] In the MPII dataset, multi-person detection tasks account for 60%, and single-person detection tasks account for 40%. This dataset focuses more on complex multi-person scene detection, and places higher requirements on the detection model's ability to handle occlusion and key point association. In terms of the distribution of detection methods, the heat map method accounts for the highest proportion (65.4%), showing its stability and maturity as the mainstream method for key point detection; regression methods account for 25.7%; classification head methods account for only 8.9%, but as a trending solution, they show strong development potential.

[0046] In the MPII dataset, the performance of pose estimation is usually evaluated by the PCK metric. Specifically, the MPII dataset uses the individual head diameter as the scale factor, which is the Euclidean distance between the upper left point and the lower right point of the head rectangle. The normalized pose estimation metric is called PCKh, which reflects the distance between the predicted key points and the true annotation points after normalization of the head size.

[0047] The normalized distance calculation formula is: ; Among them, the total number of key points is represented by A, and the indicator function (⋅) is used to determine whether the condition is met. If so, the value is 1, otherwise it is 0. The threshold ratio is determined by the parameter α, and the normalization factor S is defined as the head diameter.

[0048] In the experiments on the MPII dataset, random horizontal flipping and random bounding box transformations are used, and the SimCC label generation strategy is used. All input images are resized to 256×256 pixels by affine transformation for fair comparison with other methods. The training batch size is 16, the validation and test batch size is 32, and the total number of training epochs is set to 210. The AdamW optimizer with a learning rate of 1e-3 is used.

[0049] The results are shown in Table 1. @0.5 in Table 1 indicates that the threshold of the normalized distance is set to 0.5. The training input size of all methods in Table 1 is standardized to 256×256 pixels. As shown in Table 1, MLC Pose adopts a simcc-based method without an upsampling layer. The simcc paradigm plays a vital role in improving the performance of medium and small-scale human pose estimation. The proposed method not only outperforms the state-of-the-art 2D heatmap-based method, but also shows a better trade-off between accuracy and complexity compared with the simcc-based method. In the regression category, LAR-Pose (Lou et al., 2024) also cited a new loss function and network structure, achieving 89.7% PCKh@0.5, with a moderate number of parameters and computational complexity, taking into account both accuracy and efficiency, but still slightly inferior to the state-of-the-art regression methods such as MSRT.

[0050] Table 1 Comparison of results on the MPII test set

[0051] Specifically, MLCPose achieves higher accuracy than the state-of-the-art HRPVT-L while saving 46% of parameters and 30% of GFLOPS. Compared to SimCC†-based methods (such as those using HRNet-w32 as the backbone), MLCPose achieves higher accuracy than 0. 1PCKh@0.5 The model disclosed in this embodiment exceeds PPNetM4-D2-W32 with similar capacity.

[0052] Visualization diagram Figure 6-7 As shown, Figure 6 The original images shown here are selected from a variety of indoor and outdoor scenes, covering different detection difficulties. Figure 6 From top to bottom, the original image (Images), the visualization based on heatmap (heatmap), the visualization based on regression (regression), and the coordinate classification (MLCpose) and heatmap output results. In contrast, the limitation of the heatmap method is mainly reflected in the coarse granularity. The regression method obviously has poor positioning accuracy. The visualization obtained by the model disclosed in this embodiment is as follows: Figure 7 As shown, the granularity of the visualization image obtained by the disclosed model of this embodiment is more refined, which can not only accurately identify key points, but also further enhance the ability to capture and locate details through two one-dimensional heat maps.

[0053] The above embodiments are only exemplary embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art may make various modifications or equivalent substitutions to the present invention within the essence and protection scope of the present invention, and such modifications or equivalent substitutions shall also be deemed to fall within the protection scope of the present invention.

Claims

1. A multi-person pose estimation method with multi-level feature enhancement and collaborative calibration, characterized in that: The steps include: Step 1, construct an MLCPose model and train the constructed MLCPose model; Step 2, input the video stream data or the collected image to be predicted into the MLCPose model obtained in step 1, and the MLCPose model outputs an estimation result graph; The MLCPose model includes a backbone network, PAFPN, and Head. The backbone network is used for feature extraction, the PAFPN is used for multi-scale fusion and optimization, and the Head is used for key point prediction. The backbone network is based on multi-layer CSPLayer and convolution modules, and integrates BSCA modules to improve global modeling and local detail capture capabilities through coordinate attention and bidirectional spatial attention; the PAFPN combines high-level semantic features and low-level spatial detail features through reshaping and connection operations; the Head uses X-axis and Y-axis coordinate classifiers to accurately locate the coordinates of key points, and introduces a dynamic MLTCC mechanism to predict the precise positioning of each individual key point.

2. The method for multi-person posture estimation based on multi-level feature enhancement and collaborative calibration according to claim 1, characterized in that: The BSCA module includes a coordinate attention mechanism and a bidirectional spatial attention mechanism. The coordinate attention mechanism is used to enhance the network to handle highly spatially correlated tasks. The coordinate attention mechanism workflow includes coordinate information embedding and coordinate attention generation. The bidirectional spatial attention mechanism is used to enhance the model's responsiveness to key areas, and the bidirectional spatial attention mechanism workflow includes feature separation and attention generation.

3. The method for multi-person posture estimation based on multi-level feature enhancement and collaborative calibration according to claim 2, characterized in that: The embedding of coordinate information specifically includes: inputting a feature map , where C, H and W represent the number of channels, height and width respectively; perform one-dimensional global average pooling on the feature map in the height and width directions respectively, extract the features in the horizontal and vertical directions, and obtain the height Z h and width Z w Eigenvectors in directions; The coordinate attention generation specifically includes: and The features are concatenated and fully connected across channels through a 1×1 convolution layer, batch normalization is applied, and nonlinear characteristics are enhanced through nonlinear activation functions: ; Among them, Concat(·) means concatenation along the feature dimension. yes The convolution operation, BN(·) represents the batch normalization operation, ReLU(·) is the rectified linear unit activation function, and the obtained features ; Will Split into horizontal and vertical Two parts, apply Sigmoid activation function to each part to generate attention weights: ; Where σ(·) is the Sigmoid function, and the generated attention weight ; The generated attention weights are multiplied element-wise and Applied to the input feature map X: ; Among them, Y is the output feature map, .

4. The method for multi-person posture estimation based on multi-level feature enhancement and collaborative calibration according to claim 2, characterized in that: The feature separation includes: applying different convolution kernels to the input feature map X in the horizontal and vertical dimensions respectively, and independently extracting the features in the horizontal direction. and vertical characteristics ; The attention generation includes: applying a Sigmoid activation function to the horizontal features and vertical characteristics Processing to generate a horizontal attention map And vertical attention map , Emphasize or suppress pixels with important edge information in the horizontal direction. Focus on important features in the vertical direction; horizontal and vertical attention maps are used to reweight the spatial features of the input feature map X, and finally output the feature map .

5. The method for multi-person posture estimation based on multi-level feature enhancement and collaborative calibration according to claim 1, characterized in that: The working process of the MLTCC mechanism is as follows: the input batch data samples pass through the SimCC detection head to generate the predicted key point probability distribution and the real key point probability distribution; the loss calculation module combines the KL discrete distribution loss and the outlier perception loss, generates dynamic weights through an adaptive weight strategy, and optimizes the loss of each layer layer by layer; the weights are balanced through layer-by-layer normalization, and the high-confidence key point areas are optimized first to reduce noise interference; the joint loss function is used to optimize the model, which significantly improves the positioning accuracy and stability in the posture estimation task.

6. The method for multi-person posture estimation based on multi-level feature enhancement and collaborative calibration according to claim 5, characterized in that: The dynamic weight The calculation formula is: ; in, Used to normalize weights to ensure stability in the value space; ω i represents the contribution basis of the loss layer, Indicates the confidence of each key point; The joint loss function is specifically: ; in, Represents the sum of sub-losses from layer i to layer K; represents the sum of sub-losses from layer i to layer A; N represents the total number of layers, is a dynamic weight used to balance the loss of each layer.

7. The method for multi-person posture estimation with multi-level feature enhancement and collaborative calibration according to claim 5, characterized in that: The outlier-aware loss introduces an adaptive adjustment module to dynamically correct outliers when allocating loss weights. The loss function in the outlier-aware loss is: ; in, and are the predicted distribution and the true target distribution, respectively, K is the number of key points, N is the total number of layers, n is the layer index, is a dynamic weight factor, which is defined by the outlier perception mechanism as: ; M represents the set of key points that need to be shielded, and η is the scaling factor of the shielding points; the step size adjustment strategy is implemented by the following formula: ; Among them, α∈(0,1) is used to smooth high loss values, β>1 is used to enlarge the optimization step of normal values, and τ is the outlier detection threshold.

Citation Information

Cited By

  • Music performance posture real-time driving method and system based on time-space cooperation

    CN120411317A

  • Real-time driving method and system for music performance gesture based on time-space collaboration

    CN120411317B