Multi-person attitude estimation method based on improved YOLO11-pose
By improving the YOLO11-pose model, FasterNet network, Efficient RepGFPN feature pyramid and lightweight detection head were introduced, which solved the problem of difficult to balance speed and accuracy in multi-person pose estimation, and achieved efficient and accurate multi-person pose estimation.
Patent Information
- Application Number
- CN202510319442.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-27
AI Technical Summary
The existing multi-person pose estimation methods are difficult to achieve a balance between speed and accuracy. The reasoning time of the top-down method increases with the number of characters, the bottom-up method lags in accuracy, and the reasoning time of the method that relies on heat maps is relatively long.
A multi-person pose estimation method based on improved YOLO11-pose is proposed, and by introducing a more efficient FasterNet network to replace YOLO11-pose backbone network, reducing memory access and improving model accuracy. At the same time, the Efficient RepGFPN feature pyramid module and lightweight detail-enhanced detection head are introduced to optimize feature fusion and information transmission, reduce computing complexity and improve detection performance.
It significantly improves the accuracy and robustness of pose estimation, meets the needs of detection performance and multi-person detection in practical applications, and achieves a balance of speed and accuracy, which is suitable for efficient applications on the embedded end.
Smart Images

Figure CN120220188A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-person pose estimation, and in particular to an improved YOLO11-pose based multi-person pose estimation method. Background Art
[0002] With the rise of deep learning, various computer vision tasks have been solved using deep learning methods. Human pose estimation is a challenging detection task aimed at extracting human key point information from images. Multi-person pose estimation (MPE) is one of the important branches of human pose estimation, which can simultaneously capture the poses of multiple people in a single frame of image, and its application scope ranges from virtual reality to sports analysis. Although many multi-person pose estimation techniques have emerged, it is still challenging to achieve a balance between speed and accuracy.
[0003] Currently, there are mainly two types of multi-person pose estimation methods: top-down and bottom-up. The top-down method uses a pre-trained detector to create bounding boxes around the subjects, and then performs pose estimation for each individual. A key limitation is that their inference time varies with the number of people in the image. The bottom-up method directly predicts the key point positions of all individuals in the image. Then, the key points are assigned. However, compared with the top-down method, the current bottom-up method lags behind in terms of accuracy. In addition, there are some methods that rely on heatmap-based human pose estimation, but these methods rely heavily on post-processing techniques, resulting in a long inference time. Summary of the Invention
[0004] In order to overcome the above deficiencies, the present invention proposes a multi-person human pose estimation model based on improved YOLO11-pose. By improving the YOLO11-pose model structure, the accuracy of the model for key point detection in various scenarios is enhanced. The present invention aims to improve the accuracy and robustness of pose estimation, as well as the efficient application on the embedded side, to meet the requirements for detection performance and multi-person detection in practical applications.
[0005] According to one aspect of the present invention, there is provided an improved YOLO11-pose based multi-person pose estimation method. The YOLO11-pose network includes a FasterNet network module, an SPPF module, a C2PSA module, a feature pyramid module, a detection head module, and an output module. The method includes: Obtain an image to be processed; input the image to be processed into the FasterNet network module for feature extraction to obtain a first feature map; Input the first feature module into the SPPF module to obtain a multi-scale second feature map; input the multi-scale second feature map into the C2PSA module for feature enhancement; Input the enhanced multi-scale second feature map into the feature pyramid module for feature extraction and fusion to generate a multi-scale third feature map; Input the multi-scale third feature map into the detection head module to generate the prediction results of human pose key points; Input the prediction results into the output module for decoding to achieve multi-person pose estimation.
[0006] In the above technical solution, the present invention introduces a more efficient network, FasterNet, to replace the backbone network of YOLO11-pose, reducing memory access and improving the model accuracy. Specifically, FasterNet adopts efficient operators such as partial convolution (PConv) and pointwise convolution (PWConv), reducing redundant calculations and memory access, thus significantly improving the running speed of the network. In addition, FasterNet is optimized in feature fusion, capable of quickly fusing features of different scales, further enhancing the inference speed of the network and meeting the requirements of multi-person real-time detection. Moreover, through the combination of PConv and PWConv, FasterNet can more effectively extract spatial features and improve the feature expression ability. This enables the network to better capture the target information of different scales and complexities in the image, thereby improving the detection accuracy.
[0007] In some embodiments, the FasterNet network module includes at least 4 layers, each layer consisting of at least one embedding layer or merging layer and at least one FasterNet module; The FasterNet module is composed of a PConv layer connecting two PWConv layers; wherein, the PConv layer is configured to apply ordinary convolution to at least 50% of the input channels to extract spatial features, and the remaining channels remain unchanged.
[0008] In the above technical solution, on the basis of introducing the FasterNet network, the FasterNet network is further improved to significantly reduce the memory access volume during model inference and have stronger feature extraction capabilities. Specifically, FasterNet contains 4 stages, and there is an embedding layer or a merging layer before each stage for downsampling or increasing the number of channels. Each stage includes several FasterNet blocks. This design enables FasterNet to gradually extract high-level features of the image while maintaining computational efficiency. Each FasterNet block includes a PConv (Partial Convolution) layer followed by two PWConv (Point-Wise Convolution) layers. PConv only applies ordinary convolution to some of the input channels to extract spatial features, and the remaining channels remain unchanged, reducing the transmission of redundant information, retaining more original feature information, and improving the efficiency and accuracy of feature extraction. PWConv is used to further process the feature map after PConv, and it can effectively integrate channel information and enhance the feature representation ability while maintaining computational efficiency by point-wise convolution. Adding PWConv after PConv makes the input feature map more focused on the central position compared with the conventional convolutional layer. Since PConv only uses a part of the channels, PWConv can make full use of the channel information at this time. Through this design, the memory access volume during model inference can be significantly reduced, and it has stronger feature extraction capabilities.
[0009] In some embodiments, the feature pyramid module is an Efficient RepGFPN feature pyramid module; Among them, the ordinary convolutional layer in the feature pyramid module is replaced by a double convolutional layer; the upsampling operation module is replaced by a dynamic upsampling module; The double convolutional layer is configured to perform a convolutional operation on the input feature map to extract deep features; The dynamic upsampling module is configured to adaptively adjust the upsampling operation according to the resolution of the feature map.
[0010] In the above technical solution, on the basis of introducing the FasterNet network, the EfficientRepGFPN module is further introduced to replace the FPN feature pyramid module in the traditional YOLO11-pose network. Efficient RepGFPN (Efficient Reparameterized Generalized Feature Pyramid Network) is an improved feature pyramid network proposed in DAMO-YOLO, aiming to improve the accuracy and efficiency of object detection. This module fully draws on the advantages of the Generalized Feature Pyramid Network (GFPN) and makes a number of innovations on this basis to meet the design requirements of multi-person object detection. On this basis, the present invention further improves the EfficientRepGFPN module. Specifically, the ordinary convolution is replaced by a dual convolutional layer (Dual_Conv) to achieve effective compression of parameters and high efficiency of calculation while maintaining the feature depth extraction and representation ability. The upsampling operation module in Efficient RepGFPN is replaced by a dynamic upsampling module. The design purpose of the dynamic upsampling module is to adaptively adjust the upsampling operation according to the resolution of the feature map. Specifically, the dynamic upsampling module can automatically select an appropriate upsampling method according to the resolution of the input feature map, thereby reducing the computational cost while improving the flexibility and adaptability of the model. The application of the EfficientRepGFPN module in YOLO11-pose brings the following advantages: Through more efficient feature fusion and information transmission, the EfficientRepGFPN module significantly improves the accuracy of object detection, and at the same time maintains the real-time performance of the model by optimizing computational resources and reducing latency. The design of the dynamic upsampling module and the dual convolutional layer enables the model to maintain good performance under inputs of different resolutions, enhancing the robustness of the model.
[0011] In some embodiments, the detection head module includes a plurality of detection branches; For one detection branch, it is configured to process the third feature map of one scale; the same convolutional layer weights are used for each detection branch to perform a convolutional operation on the third feature maps of different branches to generate corresponding preliminary feature maps; the BN layers of each branch are calculated independently to normalize the preliminary feature maps to generate normalized feature maps; the Coord Attention mechanism is applied to the normalized feature maps to generate channel and spatial attention maps, and the feature maps are weighted to generate enhanced feature maps; after performing a scale operation on the enhanced feature maps, human pose keypoint prediction results are generated based on the feature maps.
[0012] In the above technical solution, on the basis of introducing the FasterNet network, the detection head is further improved. The shared convolutional layer weights reduce the scale of model parameters and computational complexity. At the same time, the independent BN layer and CoordAttention mechanism can further enhance the expressive ability and pertinence of the feature map while maintaining a certain number of parameters, improve the efficiency of feature extraction and utilization, and enhance the accuracy and performance of human pose keypoint prediction. The scale operation ensures that the model can correctly process feature maps of different scales, making the network sensitive to the target size on multi-scale features, thereby improving the detection performance. Specifically, the detection head module contains multiple detection branches, and each branch corresponds to processing a third feature map of a specific scale. This multi-branch structural design can effectively process features of different scales to meet the detection requirements of human pose keypoints at different scales. Each detection branch uses the same convolutional layer weights to perform convolutional operations on its corresponding third feature map to generate a preliminary feature map. The same convolutional layer weights can, on the one hand, share the feature extraction mode among different branches and reduce the number of parameters, and on the other hand, ensure the consistency of convolutional conversion for feature maps of different scales, which is helpful for subsequent feature fusion and processing. The batch normalization (BN) layers of each branch are calculated independently to normalize the preliminary feature map to obtain the normalized feature map. The independent calculation of the BN layer can be adapted to the feature distribution characteristics of different branches, making the features of each branch have more appropriate statistical characteristics after normalization, which is beneficial to improving the training effect and generalization ability of the model and avoiding affecting the stability and performance of the network due to excessive differences in feature distributions between different branches. Apply the Coord Attention mechanism to the normalized feature map to generate channel and spatial attention maps, and perform weighted processing on the feature map to obtain the enhanced feature map. The Coord Attention mechanism can simultaneously model the feature relationships in the channel and spatial dimensions, capture important channel information and spatial position information in the feature map, highlight the feature parts that are important for human pose keypoint detection through weighted means, and suppress irrelevant or interfering features, thereby enhancing the discriminative ability of the feature map and making the generated feature map more focused on the relevant features of human pose keypoints. Through the scale operation, feature maps of different scales can be appropriately adjusted and processed, enabling the network to better adapt to these scale changes and accurately capture the keypoint features of targets of different sizes. This helps to improve the accuracy and robustness of the model in the face of complex scenes and multi-scale target detection, and avoid detection omissions or errors caused by target scale changes. After the above series of processes, the enhanced feature map already better represents the key information of human poses. Finally, prediction results such as the position and category of human pose keypoints can be obtained through methods such as classification and regression, providing an accurate basis for human pose analysis.
[0013] According to another aspect of the present invention, there is provided a training method for an improved YOLO11-pose network, where the YOLO11-pose network includes a FasterNet network module, an SPPF module, a C2PSA module, a feature pyramid module, a detection head module, and an output module; the method includes: Obtain training images and their corresponding ground truth labels; Input the training image into the FasterNet network module for feature extraction to obtain a first feature map; Input the first feature module into the SPPF module to obtain a multi-scale second feature map; input the multi-scale second feature map into the C2PSA module for feature enhancement; Input the enhanced multi-scale second feature map into the feature pyramid module for feature extraction and fusion to generate a multi-scale third feature map; Input the multi-scale third feature map into the detection head module to generate a prediction result of human pose key points; Input the prediction result into the output module for decoding to generate a final prediction label; Based on the ground truth label and the prediction label of the training image, calculate a loss value; and Based on the loss value, train the YOLO11-pose network.
[0014] In the above technical solution, this training method corresponds one by one to the steps of the above-mentioned method for multi-person pose estimation based on improved YOLO11-pose, and the related advantages will not be elaborated.
[0015] According to yet another aspect of the present invention, there is provided an improved YOLO11-pose network, where the YOLO11-pose network includes a FasterNet network module, an SPPF module, a C2PSA module, a feature pyramid module, a detection head module, and an output module; this YOLO11-pose network is trained by the training method according to any one of claims 5-8.
[0016] In the above technical solution, in the above technical solution, this network is obtained by the training method of the above-mentioned improved YOLO11-pose network, and the related advantages will not be elaborated.
[0017] According to still another aspect of the present invention, there is provided an electronic device, including: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.
[0018] In the above technical solution, in order to better run and process the method, the above method is stored in a memory, and a processor is used to execute the stored method. It should be noted that the principle and effect of each step have been described above and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 is a schematic flowchart of an embodiment of a method for multi-person pose estimation based on improved YOLO11-pose of the present invention; Figure 2 is a schematic diagram of the FasterNet structure of an embodiment of a method for multi-person pose estimation based on improved YOLO11-pose of the present invention; Figure 3 is a schematic diagram of the feature pyramid structure of an embodiment of a method for multi-person pose estimation based on improved YOLO11-pose of the present invention; Figure 4 is a schematic diagram of the LSDECD-CA detection head structure of an embodiment of a method for multi-person pose estimation based on improved YOLO11-pose of the present invention; Figure 5 is a schematic diagram of the overall structure of improved YOLO11-pose of an embodiment of a method for multi-person pose estimation based on improved YOLO11-pose of the present invention; Figure 6 is a schematic flowchart of an embodiment of a training method for an improved YOLO11-pose network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The present invention will be further described in detail below with reference to the drawings and embodiments. It should be specifically noted that the following embodiments are only used to illustrate the present invention, but do not limit the scope of the present invention. Similarly, the following embodiments are only partial embodiments of the present invention rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0022] The present invention provides an improved YOLO11-pose multi-person pose estimation method. By improving the YOLO11-pose model structure, the accuracy of the model for key point detection in various scenarios is enhanced. The purpose of the present invention is to improve the accuracy and robustness of pose estimation, as well as the efficient application on the embedded side, to meet the requirements for detection performance and multi-person detection in practical applications.
[0023] One of the embodiments Please refer to Figure 1 , an improved YOLO11-pose multi-person pose estimation method. The YOLO11-pose network includes a FasterNet network module, an SPPF module, a C2PSA module, a feature pyramid module, a detection head module, and an output module. The method includes: A1. Obtain the image to be processed; input the image to be processed into the FasterNet network module for feature extraction to obtain a first feature map; In this embodiment, the present invention introduces a more efficient network, FasterNet, to replace the backbone network of YOLO11-pose, that is, the part before the SPPF module, reducing memory access and improving the model accuracy. Specifically, FasterNet adopts efficient operators such as partial convolution (PConv) and pointwise convolution (PWConv), reducing redundant calculations and memory access, thus significantly improving the running speed of the network. In addition, FasterNet is optimized in feature fusion, capable of quickly fusing features of different scales, further enhancing the inference speed of the network and meeting the real-time requirements. Moreover, through the combination of PConv and PWConv, FasterNet can extract spatial features more effectively, improving the feature expression ability. This enables the network to better capture the target information of different scales and complexities in the image, thereby improving the detection accuracy.
[0024] In this embodiment, please refer to Figure 2 , the FasterNet network module includes at least 4 layers, each layer consists of at least one embedding layer or merging layer and at least one FasterNet module. The FasterNet module is composed of a PConv layer connecting two PWConv layers. Among them, the PConv layer is configured to apply ordinary convolution only to some of the input channels to extract spatial features (it should be noted that these partial channels can be set in the code. Usually set to 0.5 or 0.75, that is, perform convolution calculations on 50% or 75% of the input, and the remaining channels remain unchanged. In this embodiment, it is 50%). Specifically: FasterNet consists of 4 stages, and before each stage, there is an embedding layer (a normal convolution with a stride of 4) or a merging layer (a normal convolution with a stride of 2) for downsampling or increasing the number of channels. Each stage includes several FasterNet blocks. The FasterNet blocks include a PConv (Partial Convolution) layer followed by two PWConv (Point-Wise Convolution) layers. The normal convolution directly slides the convolution kernel over the entire input feature map, with FLOPs of h×w×k2×c2; PConv only applies the normal convolution to a part of the input channels to extract spatial features, and the remaining channels remain unchanged. Through this design, redundant information transmission is reduced, and most of the original feature information is retained. The FLOPs are: h×w×k2×cp2 (cp is the number of partial channels); DWConv convolves only one channel with one convolution kernel, and the FLOPs are: h×w×k2×c. Adding PWConv after PConv reduces the computational cost while maintaining the feature expression ability. Compared with the conventional Conv, it pays more attention to the central position. Since PConv only uses a part of the channels, adding PWConv after PConv at this time can fully utilize the channel information. Through this design, the memory access volume during model inference can be significantly reduced, and it has a stronger feature extraction ability. Replace the part before the SPPF layer in the YOLO11-pose backbone network with FasterNet.
[0025] The advantage of this solution is that on the basis of introducing the FasterNet network, it further improves the FasterNet network to significantly reduce the memory access volume during model inference and has stronger feature extraction capabilities. Specifically, FasterNet contains 4 stages, and there is an embedding layer or a merging layer before each stage for downsampling or increasing the number of channels. Each stage includes several FasterNet blocks. This design enables FasterNet to gradually extract high-level features of the image while maintaining computational efficiency. Each FasterNet block includes a PConv (Partial Convolution) layer followed by two PWConv (Point-Wise Convolution) layers. PConv applies ordinary convolution only to some of the input channels to extract spatial features, and the remaining channels remain unchanged, reducing the transmission of redundant information, retaining more original feature information, and improving the efficiency and accuracy of feature extraction. PWConv is used to further process the feature map after PConv and integrate channel information through point-wise convolution, which can effectively integrate channel information, enhance the feature representation ability, and maintain computational efficiency. Compared with the conventional convolution layer, it pays more attention to the central position. Since PConv only uses a part of the channels, connecting PWConv after PConv at this time can make full use of the channel information. Through this design, the memory access volume during model inference can be significantly reduced, and it has stronger feature extraction capabilities.
[0026] A2. Input the first feature module into the SPPF module to obtain a multi-scale second feature map; input the multi-scale second feature map into the C2PSA module for feature enhancement; A3. Input the enhanced multi-scale second feature map into the feature pyramid module for feature extraction and fusion to generate a multi-scale third feature map; In this embodiment, please refer to Figure 3, the feature pyramid module is an Efficient RepGFPN feature pyramid module; by optimizing the Efficient RepGFPN feature pyramid, an Efficient RepGDFPN feature pyramid is composed of CSPStage, convolution, Dual_Conv, and dynamic upsampling. The present invention introduces Dual_Conv, replaces the ordinary convolutional layer and adopts the dynamic upsampling technology to replace the traditional upsampling operation, obtains the Efficient RepGDFPN feature pyramid, and replaces the FPN feature pyramid in the YOLO11-pose model. Among them, the ordinary convolutional layer in the feature pyramid module is replaced by a double convolutional layer; the upsampling operation module is replaced by a dynamic upsampling module; the double convolutional layer is configured to perform a convolutional operation on the input feature map to extract deep features; the dynamic upsampling module is configured to adaptively adjust the upsampling operation according to the resolution of the feature map. Specifically, some ordinary convolutions of the Efficient RepGFPN feature pyramid are replaced by Dual_Conv, which realizes effective compression of parameters and high efficiency of calculation while maintaining the feature depth extraction and representation ability. The traditional upsampling in the Efficient RepGFPN feature pyramid is replaced by dynamic upsampling, which reduces the calculation cost and better adapts to the data features.
[0027] The advantage of this solution is that on the basis of introducing the FasterNet network, the EfficientRepGFPN module is further introduced to replace the FPN (Feature Pyramid Network) in the traditional YOLO11-pose network. Efficient RepGFPN (Efficient Reparameterized Generalized Feature Pyramid Network) is an improved feature pyramid network proposed in DAMO-YOLO, aiming to improve the accuracy and efficiency of object detection. This module fully draws on the advantages of the Generalized Feature Pyramid Network (GFPN) and makes a number of innovations on this basis to meet the design requirements of multi-person object detection. On this basis, the present invention further improves the EfficientRepGFPN module. Specifically, ordinary convolutions are replaced with dual convolutional layers (Dual_Conv) to achieve effective compression of parameters and high efficiency of computation while maintaining the ability to extract and represent features in depth. The upsampling operation module in Efficient RepGFPN is replaced by a dynamic upsampling module. The design purpose of the dynamic upsampling module is to adaptively adjust the upsampling operation according to the resolution of the feature map. Specifically, the dynamic upsampling module can automatically select an appropriate upsampling method according to the resolution of the input feature map, thereby reducing the computational cost while improving the flexibility and adaptability of the model. The application of the EfficientRepGFPN module in YOLO11-pose brings the following advantages: Through more efficient feature fusion and information transmission, the EfficientRepGFPN module significantly improves the accuracy of object detection, and at the same time maintains the real-time performance of the model by optimizing computational resources and reducing latency. The design of the dynamic upsampling module and the dual convolutional layer enables the model to maintain good performance under inputs of different resolutions, enhancing the robustness of the model.
[0028] A4. Input the multi-scale third feature maps into the detection head module to generate the prediction results of human pose key points; In this embodiment, please refer to Figure 4, the detection head module includes several detection branches; for one detection branch, it is configured to process the third feature map of one scale; for each detection branch, the same convolutional layer weights are used to perform convolution operations on the third feature maps of different branches to generate corresponding preliminary feature maps; the BN layers of each branch are calculated independently to normalize the preliminary feature maps to generate normalized feature maps; the CoordAttention mechanism is applied to the normalized feature maps to generate channel and spatial attention maps, and the feature maps are weighted to generate enhanced feature maps; after performing a scale operation on the enhanced feature maps, human pose key point prediction results are generated based on the feature maps. Specifically, the present invention proposes a lightweight detail-enhanced detection head that uses a shared convolutional layer for each branch, while the BN is calculated independently. A scale operation is added to the output of each detection head to obtain the LSDECD-CA detection head, and the detection head in the YOLO11-pose model is replaced. The specific steps are as follows: Optimize the feature extraction part of the YOLO11-pose detection head, use shared convolution and the attention mechanism (Coord Attention) for feature extraction, and the BN of each branch is calculated independently. While keeping the model lightweight, the ability to capture details is maintained. Apply a scale operation to the output of the detection head. By appropriately scaling and normalizing the outputs of feature maps of different scales, it is ensured that the model can correctly process feature maps of different scales, enabling the network to maintain sensitivity to target sizes on multi-scale features, thereby improving the detection performance. Use LSDECD-CA to replace the detection head of the YOLO11-pose.
[0029] The advantages of this solution are as follows. Based on the introduction of the FasterNet network, the detection head is further improved. The shared convolutional layer weights reduce the scale of model parameters and the computational complexity. At the same time, the independent BN layer and the CoordAttention mechanism can further enhance the expressive ability and pertinence of the feature map while maintaining a certain number of parameters, improve the efficiency of feature extraction and utilization, and enhance the accuracy and performance of human pose keypoint prediction. The scale operation ensures that the model can correctly process feature maps of different scales, making the network sensitive to the target size on multi-scale features, thereby improving the detection performance. Specifically, the detection head module contains multiple detection branches, and each branch corresponds to processing a third feature map of a specific scale. This multi-branch structural design can effectively process features of different scales to meet the detection requirements of human pose keypoints at different scales. Each detection branch uses the same convolutional layer weights to perform convolutional operations on its corresponding third feature map to generate a preliminary feature map. The same convolutional layer weights can, on the one hand, share the feature extraction pattern among different branches and reduce the number of parameters, and on the other hand, ensure the consistency of convolutional transformation for feature maps of different scales, which is helpful for subsequent feature fusion and processing. The batch normalization (BN) layers of each branch are calculated independently to normalize the preliminary feature map to obtain the normalized feature map. The independent calculation of the BN layer can be adapted to the characteristics of the feature distribution of different branches, making the features of each branch have more appropriate statistical characteristics after normalization, which is beneficial to improving the training effect and generalization ability of the model and avoiding affecting the stability and performance of the network due to excessive differences in feature distributions between different branches. The Coord Attention mechanism is applied to the normalized feature map to generate channel and spatial attention maps, and the feature map is weighted to obtain the enhanced feature map. The Coord Attention mechanism can simultaneously model the feature relationships in the channel and spatial dimensions, capture important channel information and spatial position information in the feature map, highlight the feature parts that are important for human pose keypoint detection through weighting, and suppress irrelevant or interfering features, thereby enhancing the discriminative ability of the feature map and making the generated feature map more focused on the features related to human pose keypoints. Through the scale operation, feature maps of different scales can be appropriately adjusted and processed, enabling the network to better adapt to these scale changes and accurately capture the keypoint features of targets of different sizes. This helps to improve the accuracy and robustness of the model in the face of complex scenes and multi-scale target detection, and avoid detection omissions or errors caused by target scale changes. After the above series of processes, the enhanced feature map has better represented the key information of human poses. Finally, the position, category, etc. of human pose keypoints can be obtained through methods such as classification and regression, providing an accurate basis for human pose analysis.
[0030] A5. Input the prediction result into the output module for decoding to achieve multi-person pose estimation.
[0031] In this embodiment, the finally improved YOLO11-pose network of the present invention is as Figure 5 shown. From the above description of the present invention, compared with the prior art, the improvements made to some algorithms of the present invention are as follows: 1. The present invention introduces a more efficient network, FasterNet, to replace the backbone network of YOLO11-pose, reducing memory access and improving the model accuracy. 2. The present invention introduces Dual_Conv into the Efficient RepGFPN feature pyramid to replace the ordinary convolutional layer, achieving an improvement in the model's computational efficiency while maintaining the depth feature extraction and representation capabilities. In addition, the present invention also uses dynamic upsampling technology to replace the traditional upsampling operation, reducing the computational cost and better adapting to the data features. 3. The present invention proposes a lightweight detail enhancement detection head. By using a shared convolutional layer for each branch and calculating BN separately, the ability to capture details is maintained while keeping the model lightweight. Adding a scale operation to the output of each detection head helps the model better fit targets of different scales. In this embodiment, in order to verify the effectiveness of each module of the present invention's solution, YOLO11n-pose is used as the baseline, and ablation experiments are conducted on each key improvement part. The experimental results are shown in the following table:
[0032] The experimental results show that after replacing the backbone network before the SPPF module of the original model with FasterNet, the computational amount of the model increases, and at the same time, the AP increases by 3.0%. On this basis, replacing the Neck part of the original model with Efficient-RepGDFPN, the computational amount increases slightly, but the AP further increases by 2.0%. Finally, using the LSDECD-CA detection head to replace the Head part of the original model, through the design of shared convolution and learnable factors, the AP slightly decreases by 0.4%, but the AP 50 increases by 0.1%, and at the same time, the computational amount decreases. This shows that this detection head can maintain the detail capture ability while effectively reducing the computational amount.
[0033] In summary, the ablation experiments verified that the model proposed in this patent has a significant improvement in accuracy compared to the baseline model. Meanwhile, the increase in computational cost is reasonable, and it has better comprehensive performance. This benefits from the stronger feature extraction ability of FasterNet, the global semantic fusion and transmission ability of Efficient-RepGDFPN, and the lightweight and detail-preserving ability of LSDECD-CA.
[0034] In this embodiment, to verify the effectiveness of the present invention, the present invention compared some mainstream pose estimation models. The experimental results are shown in the following table:
[0035] The present invention was compared with other mainstream human pose estimation methods on the COCO-Pose validation set. These methods include classical bottom-up and top-down models. The model of the present invention is comparable to more complex models in terms of accuracy. Although the number of parameters is much less than that of Hourglass, the model of the present invention exceeds it in terms of accuracy. Compared with YOLOv5s6-pose_640_ti_lite, the model of the present invention has higher accuracy with a 33.55% reduction in computational cost. In addition, compared with YOLOv8n-pose and YOLO11n-pose, the model of the present invention achieves the best performance with the highest accuracy. These experimental results verify the effectiveness of the model of the present invention.
[0036] Embodiment 2 Please refer to Figure 6 , a training method for an improved YOLO11-pose network, where the YOLO11-pose network includes a FasterNet network module, an SPPF module, a C2PSA module, a feature pyramid module, a detection head module, and an output module; the method includes: B1. Obtain training images and their corresponding ground truth labels; B2. Input the training image into the FasterNet network module for feature extraction to obtain a first feature map; B3. Input the first feature module into the SPPF module to obtain a multi-scale second feature map; input the multi-scale second feature map into the C2PSA module for feature enhancement; B4. Input the enhanced multi-scale second feature map into the feature pyramid module for feature extraction and fusion to generate a multi-scale third feature map; B5. Input the multi-scale third feature map into the detection head module to generate a prediction result of human pose key points; B6. Input the prediction result into the output module for decoding to generate the final prediction label; B7. Calculate the loss value based on the true label and the predicted label of the training image; and B8. Train the YOLO11-pose network based on the loss value.
[0037] In the above technical solution, this training method corresponds one by one to the steps of a multi-person pose estimation method based on the improved YOLO11-pose described in one of the embodiments, and the relevant advantages will not be elaborated. The hardware environment for experimental training and testing includes an Intel(R) Xeon(R) W3-2423 CPU processor and an NVIDIA GeForce 4090 GPU. The software environment includes the CUDA 11.2 parallel computing framework, Python 3.8.1, and the PyTorch 1.13.1 deep learning framework. The operating system is Windows 11. The SGD (Stochastic Gradient Descent) optimizer is used for network training. The batch size is set to 32, the initial learning rate is set to 1e-2, the weight decay is 0.0005, and the momentum factor is 0.937. All tests are evaluated using Pycocotools, with a batch size of 16 and an input image size of 640x640.
[0038] The FasterNet network module contains at least 4 layers, and each layer consists of at least one embedding layer or merging layer and at least one FasterNet module; the FasterNet module is composed of a PConv layer connecting two PWConv layers; wherein, the PConv layer is configured to apply ordinary convolution only to some of the input channels to extract spatial features, and the remaining channels remain unchanged.
[0039] The feature pyramid module is an Efficient RepGFPN feature pyramid module; Among them, the ordinary convolution layer in this feature pyramid module is replaced by a double convolution layer; the upsampling operation module is replaced by a dynamic upsampling module; The double convolution layer is configured to perform a convolution operation on the input feature map to extract deep features; The dynamic upsampling module is configured to adaptively adjust the upsampling operation according to the resolution of the feature map.
[0040] The detection head module contains several detection branches; For a detection branch, it is configured to process the third feature map of one scale; using the same convolutional layer weights for each detection branch, perform convolutional operations on the third feature maps of different branches to generate corresponding preliminary feature maps; the BN layers of each branch are calculated independently to normalize the preliminary feature maps to generate normalized feature maps; apply the Coord Attention mechanism to the normalized feature maps to generate channel and spatial attention maps, perform weighted processing on the feature maps to generate enhanced feature maps; after performing a scale operation on the enhanced feature maps, generate human pose keypoint prediction results based on the feature maps.
[0041] Embodiment Three Please refer to Figure 5 , to provide an improved YOLO11-pose network, the YOLO11-pose network includes a FasterNet network module, an SPPF module, a C2PSA module, a feature pyramid module, a detection head module, and an output module; this YOLO11-pose network is trained by the training method described in Embodiment Two. This network is obtained by the training method of an improved YOLO11-pose network described in Embodiment Two, and the related advantages will not be elaborated here.
[0042] Embodiment Three An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-8.
[0043] In the above technical solution, in order to better run and process this method, the above method is stored in the memory, and the processor is used to execute the stored method. It should be noted that the principle and effect of each step have been described above and will not be elaborated here.
[0044] The above are only partial embodiments of the present invention, and thus do not limit the protection scope of the present invention. Any equivalent device or equivalent process transformation made using the content of the specification and drawings of the present invention, directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A multi-person pose estimation method based on improved YOLO11-pose, characterized in that: The YOLO11-pose network includes a FasterNet network module, an SPPF module, a C2PSA module, a feature pyramid module, a detection head module, and an output module; the method includes: Acquire an image to be processed; input the image to be processed into the FasterNet network module to extract features to obtain a first feature map; Input the first feature module into the SPPF module to obtain a multi-scale second feature map; input the multi-scale second feature map into the C2PSA module for feature enhancement; Inputting the enhanced multi-scale second feature map into a feature pyramid module for feature extraction and fusion to generate a multi-scale third feature map; Input the multi-scale third feature map into the detection head module to generate a human body posture key point prediction result; The prediction result is input into the output module for decoding to achieve multi-person posture estimation.
2. A multi-person posture estimation method based on improved YOLO11-pose as claimed in claim 1, characterized in that: The FasterNet network module comprises at least 4 layers, each layer consisting of at least one embedding layer or merging layer and at least one FasterNet module; The FasterNet module consists of a PConv layer connected to two PWConv layers; the PConv layer is configured to apply ordinary convolution to only at least 50% of the input channels to extract spatial features, and the remaining channels remain unchanged.
3. A multi-person posture estimation method based on improved YOLO11-pose as claimed in claim 1, characterized in that: The feature pyramid module is an Efficient RepGFPN feature pyramid module; Among them, the ordinary convolution layer in the feature pyramid module is replaced by a double convolution layer; the upsampling operation module is replaced by a dynamic upsampling module; The double convolutional layer is configured to perform convolution operations on the input feature map to extract deep features; The dynamic upsampling module is configured to adaptively adjust the upsampling operation according to the resolution of the feature map.
4. A multi-person posture estimation method based on improved YOLO11-pose as claimed in claim 1, characterized in that: The detection head module includes a number of detection branches; For one detection branch, it is configured to process a third feature map of one scale; for each detection branch, the same convolution layer weight is used to perform convolution operations on the third feature maps of different branches to generate corresponding preliminary feature maps; the BN layer of each branch is calculated independently, the preliminary feature map is normalized, and a normalized feature map is generated; Apply the Coord Attention mechanism to the normalized feature map to generate channel and spatial attention maps, perform weighted processing on the feature map, and generate an enhanced feature map; After the enhanced feature map is scaled, the human body posture key point prediction result is generated based on the feature map.
5. An improved YOLO11-pose network training method, characterized in that: The YOLO11-pose network includes a FasterNet network module, an SPPF module, a C2PSA module, a feature pyramid module, a detection head module, and an output module; the method includes: Get training images and their corresponding true labels; Input the training image into the FasterNet network module to extract features to obtain a first feature map; Input the first feature module into the SPPF module to obtain a multi-scale second feature map; input the multi-scale second feature map into the C2PSA module for feature enhancement; Inputting the enhanced multi-scale second feature map into a feature pyramid module for feature extraction and fusion to generate a multi-scale third feature map; Input the multi-scale third feature map into the detection head module to generate a human body posture key point prediction result; The prediction result is input into the output module for decoding to generate the final prediction label; Calculating a loss value based on the true label and the predicted label of the training image; and Based on the loss value, the YOLO11-pose network is trained.
6. The training method of an improved YOLO11-pose network as claimed in claim 5, characterized in that: The FasterNet network module comprises at least 4 layers, each layer consisting of at least one embedding layer or merging layer and at least one FasterNet module; The FasterNet module consists of a PConv layer connected to two PWConv layers; the PConv layer is configured to apply ordinary convolution to only at least 50% of the input channels to extract spatial features, and the remaining channels remain unchanged.
7. The training method of an improved YOLO11-pose network as claimed in claim 5, characterized in that: The feature pyramid module is an Efficient RepGFPN feature pyramid module; Among them, the ordinary convolution layer in the feature pyramid module is replaced by a double convolution layer; the upsampling operation module is replaced by a dynamic upsampling module; The double convolutional layer is configured to perform convolution operations on the input feature map to extract deep features; The dynamic upsampling module is configured to adaptively adjust the upsampling operation according to the resolution of the feature map.
8. The training method of an improved YOLO11-pose network as claimed in claim 5, characterized in that: The detection head module includes a number of detection branches; For one detection branch, it is configured to process a third feature map of one scale; for each detection branch, the same convolution layer weight is used to perform convolution operations on the third feature maps of different branches to generate corresponding preliminary feature maps; the BN layer of each branch is calculated independently, the preliminary feature map is normalized, and a normalized feature map is generated; Apply the Coord Attention mechanism to the normalized feature map to generate channel and spatial attention maps, perform weighted processing on the feature map, and generate an enhanced feature map; After the enhanced feature map is scaled, the human body posture key point prediction result is generated based on the feature map.
9. An improved YOLO11-pose network, characterized in that: The YOLO11-pose network includes a FasterNet network module, an SPPF module, a C2PSA module, a feature pyramid module, a detection head module, and an output module; the YOLO11-pose network is trained by the training method described in any one of claims 5 to 8.
10. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.