Lightweight crowded scene pedestrian detection method based on improved YOLOv11n

By improving the backbone network, feature fusion network and head network of the YOLOv11n model, the problems of low detection accuracy and high computational complexity of occlusion and multi-scale targets in crowded scenarios are solved, and lightweight and efficient pedestrian detection is achieved.

CN120298961AActive Publication Date: 2025-07-11NORTHEAST DIANLI UNIVERSITY

Patent Information

Application Number
CN202510348485.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The existing pedestrian detection models have problems with low detection accuracy and high computational complexity when occluding and multi-scale target processing in crowded scenarios, making it difficult to deploy in real-time on resource-constrained devices.

Method used

Using the improved YOLOv11n model, by building an improved backbone network, a lightweight feature fusion network, and an improved head network, combined with the C3k2_EMBC module, Enhance_FPN module and DySample module, we optimize feature extraction and fusion to reduce the calculation amount.

Benefits of technology

It improves the accuracy and speed of pedestrian detection in crowded scenarios, reduces the calculation overhead and parameter amount of the model, and makes it suitable for resource-constrained equipment deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298961A_ABST
    Figure CN120298961A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight crowded scene pedestrian detection method based on improved YOLOv11n. The method specifically comprises the following steps: (1) establishing a crowded scene pedestrian image data set; (2) adding annotation information to the images in the data set; (3) constructing a lightweight crowded scene pedestrian detection model based on the improved YOLOv11n; (4) training the model by adopting the training set and the verification set, and storing the trained model; and (5) the test set is adopted to test the model, the precision of the model meets the generalization requirement, and the final lightweight crowded scene pedestrian detection model based on the improved YOLOv11n is obtained. Compared with the prior art, the lightweight crowded scene pedestrian detection method based on the improved YOLOv11n disclosed by the invention has the advantages that the accuracy of crowded scene pedestrian detection can be effectively improved, and meanwhile, the model has a better lightweight characteristic and is convenient to deploy on a mobile hardware platform with limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and specifically relates to a lightweight pedestrian detection method for crowded scenes based on improved YOLOv11n. Background Art

[0002] With the expansion of urban scale and the growth of population density, public safety issues have become increasingly prominent. Especially in crowded scenes such as transportation hubs, commercial centers, and large-scale event venues, public safety accidents such as congestion and stampedes occur frequently. This not only seriously threatens the lives and property safety of citizens, but also poses higher requirements for maintaining social order. However, the traditional manual inspection method has obvious limitations in crowded scenes. It is not only time-consuming and laborious, but also difficult to achieve full-range, whole-process, and real-time inspection management. With the rapid development of intelligent and digital technologies, object detection technology based on computer vision has gradually become an important means for monitoring crowded scenes. By combining advanced monitoring devices, this technology realizes real-time monitoring of crowded scenes, which not only significantly improves the efficiency and accuracy of public safety management, but also promotes the development of public safety management towards intelligence and refinement.

[0003] Currently, pedestrian detection, as an important means for safety monitoring in crowded scenes, is mainly divided into two categories: traditional pedestrian detection methods and deep learning-based pedestrian detection methods. Traditional pedestrian detection methods use manual methods to extract features to train classifiers for distinguishing pedestrians from the background. For example, the method of using Histogram of Oriented Gradients (HOG) to extract features and then classifying and detecting the features through a Support Vector Machine (SVM); although traditional methods have achieved certain results in the field of pedestrian detection, there are problems such as weak feature extraction ability and poor generalization.

[0004] In recent years, deep learning methods have gradually become the mainstream of pedestrian detection. Object detection algorithms based on deep learning are roughly divided into two categories: two-stage object detection algorithms and one-stage object detection algorithms. Two-stage object detection algorithms are mainly the Region-based Convolutional Neural Network (R-CNN) series. The two-stage algorithm is divided into two steps. The specific process is to first generate candidate boxes and then perform classification and regression. Although this method has considerable accuracy, it is necessary to first locate the target and then classify the target. The complex processes of these two stages will result in a slower detection speed; in contrast, one-stage object detection algorithms combine object detection and classification tasks, have higher detection accuracy and faster detection speed, and occupy the mainstream in object detection tasks, such as the YOLO series.

[0005] However, in crowded scenes, the occlusion phenomenon among pedestrians leads to the loss of feature information, affecting the detection accuracy. Traditional feature extraction methods are difficult to effectively process the occluded areas. Due to the diversity of occlusion and shooting angles, there are significant scale differences in pedestrian targets. Existing feature fusion methods have limitations in dealing with multi-scale features. Existing detection models have a high computational complexity when dealing with large-scale pedestrian targets, resulting in low inference efficiency and difficulty in being applied to edge devices and real-time scenarios. Therefore, it is necessary to propose a pedestrian detection method that can effectively address challenges such as occlusion, multi-scale targets, and the difficulty of deploying existing models in crowded scenes. Summary of the Invention

[0006] Aiming at the problems existing in the prior art, the present invention provides a lightweight pedestrian detection method for crowded scenes based on improved YOLOv11n. By improving the YOLOv11n object detection model, it solves the problems of false detection and missed detection in the pedestrian detection model for crowded scenes caused by factors such as severe occlusion, diverse scale changes, and large amount of calculation among pedestrians in crowded scenes. At the same time, it reduces the volume and the number of parameters of the pedestrian detection model for crowded scenes, enabling it to be deployed in practical application scenarios with limited resources.

[0007] The technical solution provided by the present invention includes the following steps:

[0008] Step 1: Establish a pedestrian image dataset for crowded scenes to form a first dataset;

[0009] Step 2: Add annotation information to the images in the first dataset to form a second dataset, and divide the second dataset into a training set, a validation set, and a test set;

[0010] Step 3: Construct a lightweight pedestrian detection model for crowded scenes based on improved YOLOv11n; the construction of the model further includes steps 3.1 to 3.3:

[0011] Step 3.1: Construct an improved backbone network of a lightweight pedestrian detection model for crowded scenes based on improved YOLOv11n. The improved backbone network is composed of a convolutional layer 1, a convolutional layer 2, a C3k2_EMBC module 1, a convolutional layer 3, a C3k2_EMBC module 2, a convolutional layer 4, a C3k2_EMBC module 3, a convolutional layer 5, a C3k2_EMBC module 4, an SPPF module, and a C2PSA module 1 connected in sequence;

[0012] The improved backbone network outputs four different scales of pedestrian feature information through the C3k2_EMBC module 1, the C3k2_EMBC module 2, the C3k2_EMBC module 3, and the C2PSA module 1 respectively;

[0013] Step 3.2: Construct the lightweight feature fusion network LDEFPN of the lightweight crowded scene pedestrian detection model based on the improved YOLOv11n;

[0014] The lightweight feature fusion network LDEFPN specifically includes 8 convolutional layers, 6 Enhance_FPN modules, 2 DySample modules and 5 C3k2 modules;

[0015] The output of the C3k2_EMBC module 1 is used as the input of the convolutional layer 6, and the output of the convolutional layer 6 is used as the input of the Enhance_FPN module 3;

[0016] The output of the C3k2_EMBC module 2 is used as the input of the convolutional layer 7, and the output of the convolutional layer 7 is used as the input of the Enhance_FPN module 3; the output of the convolutional layer 7 is used as the input of the Enhance_FPN module 1, the output of the Enhance_FPN module 1 is used as the input of the C3k2 module 2, the output of the C3k2 module 2 is used as the input of the Enhance_FPN module 3, the output of the Enhance_FPN module 3 is used as the input of the C3k2 module 3, and the output of the C3k2 module 3 is used as the input of the convolutional layer 10;

[0017] The output of the C3k2_EMBC module 3 is used as the input of the convolutional layer 8, the output of the convolutional layer 8 is used as the input of the Enhance_FPN module 2, the output of the Enhance_FPN module 2 is used as the input of the C3k2 module 1, and the output of the C3k2 module 1 is fused with the output of the convolutional layer 7 in the Ehance_FPN module 1 through the DySample module 1; the outputs of the convolutional layer 8, the C3k2 module 1 and the convolutional layer 10 are fused in the Enhance_FPN module 4, the output of the Enhance_FPN module 4 is used as the input of the C3k2 module 4, and the output of the C3k2 module 4 is used as the input of the convolutional layer 11;

[0018] The output of the C2PSA module 1 is used as the input of the convolutional layer 9, and the output of the convolutional layer 9 is fused with the output of the convolutional layer 8 in the Enhance_FPN module 2 through the DySample module 2; the output of the convolutional layer 9 is fused with the output of the convolutional layer 11 in the Enahnce_FPN module 5, the output of the Enahnce_FPN module 5 is used as the input of the C3k2 module 5, the output of the C3k2 module 5 is used as the input of the convolutional layer 13, the output of the convolutional layer 9 is used as the input of the convolutional layer 12, and the outputs of the convolutional layer 12 and the convolutional layer 13 are fused in the Enhance_FPN module 6;

[0019] Step 3.3: Construct an improved head network for the lightweight pedestrian detection model based on the improved YOLOv11n. The improved head network includes 4 detection heads;

[0020] The output of the C3k2 module 3 is used as the input of detection head 1, the output of the C3k2 module 4 is used as the input of detection head 2, the output of the C3k2 module 5 is used as the input of detection head 3, and the output of the Enhance_FPN module 6 is used as the output of detection head 4;

[0021] Step 4: Use the training set and the validation set to train the lightweight pedestrian detection model based on the improved YOLOv11n described in Step 3, and save the trained model. Step 4 further includes Steps 4.1 to 4.4:

[0022] Step 4.1: Set the training parameters of the lightweight pedestrian detection model based on the improved YOLOv11n. The training parameters include: number of iterations, batch size, optimizer, learning rate, momentum, weight decay, and number of threads;

[0023] Step 4.2: Input the training set and validation set images and their corresponding labels into the lightweight pedestrian detection model based on the improved YOLOv11n. Use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters, and update the model parameters according to the gradient to gradually reduce the loss function;

[0024] Step 4.3: Monitor the loss function value and performance metrics during the training process. When the loss functions of the training set and the validation set no longer decrease, and at the same time, evaluation metrics such as accuracy P, recall R, and mean average precision mAP no longer improve, stop the training to avoid overfitting;

[0025] Step 4.4: After the training is completed, save the trained model and select the optimal model from it;

[0026] Step 5: Use the pedestrian test set in the crowded scene to test the optimal model obtained from the training, and obtain the final lightweight pedestrian detection model based on the improved YOLOv11n.

[0027] Furthermore, in Step 1, the images in the first dataset can be collected through the network, taken with a digital camera, or obtained from surveillance videos;

[0028] Preferably, in Step 2, the LableImg annotation tool can be used to add annotation information to pedestrians; the training set, validation set, and test set can be divided in a ratio of 6:2:2.

[0029] Further, in the C3k2_EMBC module in step 3.1, when C3k_EMBC = False, the C3k2_EMBC module replaces the Bottleneck module in the original C3k2 module with the EMB convolution disclosed in this invention patent. When C3k_EMBC = True, the C3k2_EMBC module replaces the C3k module in the original C3k2 with the C3k_EMBC module. The C3k_EMBC module replaces the Bottleneck module in the C3k module with the EMB convolution disclosed in this invention patent;

[0030] The EMB convolution includes two convolutional layers, one depthwise separable convolution, one eSE module, and one Droupout module;

[0031] The feature map input to the EMB convolution first passes through convolutional layer 1. The feature map output by convolutional layer 1 passes through normalization and the Swish activation function and is input into the depthwise separable convolution. The feature map output by the depthwise separable convolution passes through normalization and the Swish activation function and is input into the eSE module. The feature map output by the eSE module is used as the input to convolutional layer 2. The feature map output by convolutional layer 2 passes through normalization and is used as the input to the Droupout module. The feature map output by the Droupout module is aggregated with the original input feature map, and the aggregated feature map is used as the output feature map of the EMB convolution.

[0032] Further, the Enhance_FPN module in step 3.2 includes a weighted fusion module and a lightweight spatial attention mechanism; the Enhance_FPN module 1, Enhance_FPN module 2, Enhance_FPN module 5, and Enhance_FPN module 6 have 2 inputs, and the Enhance_FPN module 3 and Enhance_FPN module 4 have 3 inputs; the internal processes of the Enhance_FPN module with 2 inputs and the Enhance_FPN module with 3 inputs further include steps 3.2.1 to 3.2.5:

[0033] Step 3.2.1: For the Enhance_FPN module with 2 inputs, the input feature maps are feature map X0 and feature map X1 respectively; for the Enhance_FPN module with 3 inputs, the input feature maps are feature map X0, feature map X1, and feature map X2 respectively;

[0034] Step 3.2.2: For the Enhance_FPN module with 2 inputs, assign weights W0 and W1 to it according to the importance of X0 and X1, multiply X0 by its weight W0, multiply X1 by its weight W1, and add the weighted feature maps to obtain the weighted sum. The calculation formula for the weighted sum is:

[0035] O = X0·W0 + X1·W1 (1)

[0036] In formula (1), O is the output feature after weighted summation, X0 and X1 are the input feature maps of the Enhance_FPN module with 2 inputs respectively, W0 is the weight of feature map X0, and W1 is the weight of feature map X1;

[0037] For the Enhance_FPN module with 3 inputs, assign weights W0, W1, and W2 to it according to the importance of X0, X1, and X2, multiply X0 by its weight W0, multiply X1 by its weight W1, multiply X2 by its weight W2, and add the weighted feature maps to obtain the weighted sum. The calculation formula for the weighted sum is:

[0038] O = X0·W0 + X1·W1 + X2·W2 (2)

[0039] In formula (2), O is the output feature after weighted summation, X0, X1, and X2 are the input feature maps of the Enhance_FPN module with 3 inputs respectively, W0 is the weight of feature map X0, W1 is the weight of feature map X1, and W2 is the weight of feature map X2;

[0040] Step 3.2.3: Introduce non-linear features to the weighted sum through the Swish activation function. The calculation formula after being processed by the Swish activation function is:

[0041] fused_feature = Swish(O) (3)

[0042] In formula (3), O is the output feature after weighted summation, Swish is the Swish activation function, and fused_feature is the feature map processed by the Swish activation function;

[0043] Step 3.2.4: Use the feature map fused_feature processed by the Swish activation function as the input feature map of the convolutional layer in the lightweight spatial attention mechanism. The output feature map LSA of the convolutional layer is processed by the Sigmoid activation function. The calculation formula after being processed by the Sigmoid activation function is:

[0044] attention_map = Sigmoid(LSA) (4)

[0045] In formula (4), LSA is the output feature map of the convolutional layer in the lightweight spatial attention mechanism, Sigmoid is the Sigmoid activation function, and attention_map is the feature map processed by the Sigmoid activation function;

[0046] Step 3.2.5: Multiply the feature map attention_map processed by the Sigmoid activation function element-wise with the feature map fused_feature processed by the Swish activation function. The feature map obtained after the element-wise multiplication is used as the final output feature map of the Enhance_FPN module. The specific calculation formula for the element-wise multiplication is:

[0047] Enhance_feature = fused_feature * attention_map (5)

[0048] In formula (5), fused_feature is the feature map processed by the Swish activation function, attention_map is the feature map processed by the Sigmoid activation function; Enhance_feature is the final output feature map of the Enhance_FPN module;

[0049] Furthermore, before each DySample module in step 3.2 performs the fusion operation in the lightweight feature fusion network LDEFPN, it can choose to first reconstruct the input feature map into a high-resolution feature map using DySample upsampling with a static range factor added, or choose to first reconstruct the input feature map into a high-resolution feature map using DySample upsampling with a dynamic range factor added;

[0050] Among them, before the lightweight feature fusion network LDEFPN performs the fusion operation, first reconstructing the input feature map into a high-resolution feature map using DySample upsampling with a static range factor added is expressed as:

[0051] x′ = grid_sample(x, s) (6)

[0052] s = g + o (7)

[0053] o = 0.25linear(x) (8)

[0054] In formulas (6) to (8), x represents the input feature map, s represents the upsampling scale factor; grid_sample represents the grid sampling operation, x′ represents the reconstructed high-resolution feature map; o represents the offset, g represents the original sampling grid; linear represents the linear layer operation;

[0055] Before the lightweight feature fusion network LDEFPN performs the fusion operation, first use DySample upsampling with a dynamic range factor added to reconstruct the input feature map into a high-resolution feature map, which is expressed as:

[0056] x′ = grid_sample(x, s) (9)

[0057] s = g + o (10)

[0058] o = 0.5sigmoid(linear1(x))·linear2(x) (11)

[0059] In formulas (9) to (11), x represents the input feature map, s represents the upsampling scale factor; grid_sample represents the grid sampling operation, x′ represents the reconstructed high-resolution feature map; o represents the offset, g represents the original sampling grid; linear represents the linear layer operation, and sigmoid is the sigmoid activation function;

[0060] Further, step 5 specifically includes steps 5.1 to 5.3:

[0061] Step 5.1: Input the test set into the optimal model in step 4.4 for testing;

[0062] Step 5.2: Calculate the model performance metrics: accuracy P, recall R, mean average precision mAP, number of parameters, computational complexity GFLOPs, frames per second FPS, and model size. The specific calculation formulas for the accuracy P, recall R, and mean average precision mAP are as follows:

[0063]

[0064] In formulas (12) to (15), P is the accuracy, R is the recall, mAP is the mean average precision of all classes, AP is the average precision, m is the total number of pedestrian label classes, TP represents the number of positive samples correctly identified as positive samples, FP represents the number of negative samples misidentified as positive samples, FN represents the number of positive samples misidentified as negative samples, and AP i is the AP of the i-th class of pedestrians, and P(R) is the function of P changing with R;

[0065] Step 5.3: Evaluate the performance metrics of the test set. If the accuracy of the test set is similar to that of the training set, it indicates that the model meets the generalization requirements, and the final lightweight pedestrian detection model based on the improved YOLOv11n is obtained.

[0066] Compared with the prior art, the beneficial effects of the present invention are:

[0067] (1) A lightweight pedestrian detection method in crowded scenes based on improved YOLOv11n disclosed by the present invention. This method enhances the model's ability to restore details of occluded pedestrians and improve the detection accuracy in complex occlusion patterns by introducing the C3k2_EMBC module disclosed in the present invention's patent to finely capture the details and shapes of occluded pedestrian targets.

[0068] (2) This method introduces the lightweight feature fusion network LDEFPN disclosed in the present invention's patent to improve the detection performance of multi-scale pedestrian targets and occluded pedestrian targets while reducing the computational overhead. Description of the Drawings

[0069] Figure 1 It is a flowchart of the lightweight pedestrian detection method in crowded scenes based on improved YOLOv11n of the present invention;

[0070] Figure 2 It is a schematic structural diagram of the lightweight pedestrian detection model in crowded scenes based on improved YOLOv11n of the present invention;

[0071] Figure 3 It is a schematic structural diagram of the C3k2_EMBC module (C3k2_EMBC = False);

[0072] Figure 4 It is a schematic structural diagram of the C3k2_EMBC module (C3k2_EMBC = True);

[0073] Figure 5 It is a schematic structural diagram of the C3k_EMBC module;

[0074] Figure 6 It is a schematic structural diagram of the EMB convolution;

[0075] Figure 7 It is a schematic structural diagram of the Enhance_FPN module with 2 inputs;

[0076] Figure 8 It is a schematic structural diagram of the Enhance_FPN module with 3 inputs;

[0077] Figure 9 It is a schematic structural diagram of the DySample module, where (a) is the overall schematic structural diagram of DySample, (b) is the sampling point generator schematic diagram of DySample with a static range factor, and (c) is the sampling point generator schematic diagram of DySample with a dynamic range factor; Detailed Embodiment

[0078] In order to make the technical solutions, structural features, achieved objectives and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments and accompanied by drawings. It should be noted that the specific embodiments described herein are only used to more clearly explain the present invention and are not used to limit the present invention.

[0079] Figure 1 Figure 4 is a flowchart of the lightweight pedestrian detection method for crowded scenes based on the improved YOLOv11n of the present invention, and its implementation process is as follows:

[0080] Step 1: Establish a pedestrian image dataset for crowded scenes to form a first dataset; the images in the first dataset can be collected through the network, taken with a digital camera, or obtained from surveillance videos.

[0081] In this embodiment, in order to better evaluate the detection effect of a lightweight pedestrian detection method for crowded scenes based on the improved YOLOv11n disclosed by the present invention, the publicly available dataset CrowedHumen dataset is adopted; the CrowedHumen dataset is used to form the first dataset.

[0082] Step 2: Add annotation information to the images in the first dataset to form a second dataset, and divide the second dataset into a training set, a validation set, and a test set.

[0083] Preferably, the LableImg annotation tool is used to add annotation information to the pedestrians in the images of the first dataset; since there is already annotation information in the publicly available dataset CrowedHumen adopted in this embodiment, this step is skipped.

[0084] Convert the label files in odgt format in the first dataset in this embodiment into txt format label files required by YOLOv11n to form a second dataset.

[0085] Preferably, divide the second dataset into a training set, a validation set, and a test set according to 6:2:2; the divided training set includes 15,000 pictures, the validation set includes 4,370 pictures, and the test set includes 4,370 pictures. The image resolution in the second dataset in this embodiment is 640×640.

[0086] Step 3: Construct a lightweight pedestrian detection model for crowded scenes based on the improved YOLOv11n; the structure of the improved YOLOv11n model is as Figure 2 shown, and the model construction process specifically includes steps 3.1 to 3.3:

[0087] Step 3.1: Construct an improved backbone network for the lightweight pedestrian detection model in crowded scenes based on the improved YOLOv11n. The improved backbone network consists of a convolutional layer 1, a convolutional layer 2, a C3k2_EMBC module 1, a convolutional layer 3, a C3k2_EMBC module 2, a convolutional layer 4, a C3k2_EMBC module 3, a convolutional layer 5, a C3k2_EMBC module 4, an SPPF module, and a C2PSA module 1 connected in sequence;

[0088] The improved backbone network outputs four different scales of feature information through the C3k2_EMBC module 1, the C3k2_EMBC module 2, the C3k2_EMBC module 3, and the C2PSA module 1 respectively.

[0089] Furthermore, the structure of the C3k2_EMBC module (C3k_EMBC = False) is as Figure 3 shown, and the structure of the C3k2_EMBC module (C3k_EMBC = True) is as Figure 4 shown, and the structure of the C3k_EMBC module is as Figure 5 shown; when C3k_EMBC = False, the Bottleneck module in the original C3k2 module in the C3k2_EMBC module is replaced with the EMB convolution disclosed in this invention patent. When C3k_EMBC = True, the C3k module in the original C3k2 in the C3k2_EMBC module is replaced with the C3k_EMBC module; the C3k_EMBC module is to replace the Bottleneck module in the C3k module with the EMB convolution disclosed in this invention patent.

[0090] Furthermore, the structure of the EMB convolution is as Figure 6 shown; the EMB convolution includes two convolutional layers, a depthwise separable convolution, an eSE module, and a Droupout module;

[0091] The feature map input to the EMB convolution first passes through convolutional layer 1. The feature map output by convolutional layer 1 passes through normalization and the Swish activation function and is input into the depthwise separable convolution. The feature map output by the depthwise separable convolution passes through normalization and the Swish activation function and is input into the eSE module. The feature map output by the eSE module is used as the input of convolutional layer 2. The feature map output by convolutional layer 2 passes through normalization and is used as the input of the Droupout module. The feature map output by the Droupout module is aggregated with the original input feature map, and the aggregated feature map is used as the output feature map of the EMB convolution.

[0092] Step 3.2: Construct a lightweight feature fusion network LDEFPN for the lightweight pedestrian detection model in crowded scenes based on the improved YOLOv11n;

[0093] The lightweight feature fusion network LDEFPN specifically includes 8 convolutional layers, 6 Enhance_FPN modules, 2 DySample modules, and 5 C3k2 modules;

[0094] The output of the C3k2_EMBC module 1 serves as the input to convolutional layer 6, and the output of convolutional layer 6 serves as the input to Enhance_FPN module 3;

[0095] The output of the C3k2_EMBC module 2 serves as the input to convolutional layer 7, and the output of convolutional layer 7 serves as the input to Enhance_FPN module 3; meanwhile, the output of convolutional layer 7 serves as the input to Enhance_FPN module 1, the output of Enhance_FPN module 1 serves as the input to C3k2 module 2, the output of C3k2 module 2 serves as the input to Enhance_FPN module 3, the output of Enhance_FPN module 3 serves as the input to C3k2 module 3, and the output of C3k2 module 3 serves as the input to convolutional layer 10;

[0096] The output of the C3k2_EMBC module 3 serves as the input to convolutional layer 8, the output of convolutional layer 8 serves as the input to Enhance_FPN module 2, the output of Enhance_FPN module 2 serves as the input to C3k2 module 1, and the output of C3k2 module 1 is fused with the output of convolutional layer 7 in Enhance_FPN module 1 through DySample module 1; the outputs of convolutional layer 8, C3k2 module 1, and convolutional layer 10 are fused in Enhance_FPN module 4, the output of Enhance_FPN module 4 serves as the input to C3k2 module 4, and the output of C3k2 module 4 serves as the input to convolutional layer 11;

[0097] The output of the C2PSA module 1 serves as the input to convolutional layer 9, and the output of convolutional layer 9 is fused with the output of convolutional layer 8 in Enhance_FPN module 2 through DySample module 2; the output of convolutional layer 9 is fused with the output of convolutional layer 11 in Enhance_FPN module 5, the output of Enhance_FPN module 5 serves as the input to C3k2 module 5, the output of C3k2 module 5 serves as the input to convolutional layer 13, the output of convolutional layer 9 serves as the input to convolutional layer 12, and the outputs of convolutional layer 12 and convolutional layer 13 are fused in Enhance_FPN module 6;

[0098] Further, the structure of the Enhance_FPN module with 2 inputs is as Figure 7As shown in; the structure of the Enhance_FPN module with 3 inputs is as follows Figure 8 As shown in; the Enhance_FPN module includes a weighted fusion module and a lightweight spatial attention mechanism; the Enhance_FPN module 1, Enhance_FPN module 2, Enhance_FPN module 5, and Enhance_FPN module 6 have 2 inputs, and the Enhance_FPN module 3 and Enhance_FPN module 4 have 3 inputs; the internal processes of the Enhance_FPN module with 2 inputs and the Enhance_FPN module with 3 inputs further include steps 3.2.1 to 3.2.5:

[0099] Step 3.2.1: For the Enhance_FPN module with 2 inputs, the input feature maps are feature map X0 and feature map X1 respectively; for the Enhance_FPN module with 3 inputs, the input feature maps are feature map X0, feature map X1, and feature map X2 respectively;

[0100] Step 3.2.2: For the Enhance_FPN module with 2 inputs, according to the importance of X0 and X1, assign weights W0 and W1 to it, multiply X0 by its weight W0, multiply X1 by its weight W1, and add the weighted feature maps to obtain a weighted sum. The specific calculation formula for the weighted sum is:

[0101] O = X0·W0 + X1·W1 (1)

[0102] In formula (1), O is the output feature after weighted summation, X0 and X1 are the input feature maps of the Enhance_FPN module with 2 inputs respectively, W0 is the weight of feature map X0, and W1 is the weight of feature map X1;

[0103] For the Enhance_FPN module with 3 inputs, according to the importance of X0, X1, and X2, assign weights W0, W1, and W2 to it, multiply X0 by its weight W0, multiply X1 by its weight W1, multiply X2 by its weight W2, and add the weighted feature maps to obtain a weighted sum. The specific calculation formula for the weighted sum is:

[0104] O = X0·W0 + X1·W1 + X2·W2 (2)

[0105] In formula (2), O is the output feature after weighted summation, X0, X1, and X2 are the input feature maps of the Enhance_FPN module with 3 inputs respectively, W0 is the weight of feature map X0, W1 is the weight of feature map X1, and W2 is the weight of feature map X2;

[0106] Step 3.2.3: Introduce non-linear features to the weighted sum through the Swish activation function. The calculation formula after being processed by the Swish activation function is:

[0107] fused_feature = Swish(O) (3)

[0108] In formula (3), O is the output feature after weighted summation, Swish is the Swish activation function, and fused_feature is the feature map processed by the Swish activation function;

[0109] Step 3.2.4: Take the feature map fused_feature processed by the Swish activation function as the input feature map of the convolutional layer in the lightweight spatial attention mechanism. The output feature map LSA of the convolutional layer is processed by the Sigmoid activation function. The calculation formula after being processed by the Sigmoid activation function is:

[0110] attention_map = Sigmoid(LSA) (4)

[0111] In formula (4), LSA is the output feature map of the convolutional layer in the lightweight spatial attention mechanism, Sigmoid is the Sigmoid activation function, and attention_map is the feature map processed by the Sigmoid activation function;

[0112] Step 3.2.5: Element-wise multiply the feature map attention_map processed by the Sigmoid activation function with the feature map fused_feature processed by the Swish activation function. The feature map after element-wise multiplication is used as the final output feature map of the Enhance_FPN module. The specific calculation formula for element-wise multiplication is:

[0113] Enhance_feature = fused_feature * attention_map (5)

[0114] In formula (5), fused_feature is the feature map processed by the Swish activation function, attention_map is the feature map processed by the Sigmoid activation function; Enhance_feature is the final output feature map of the Enhance_FPN module;

[0115] Furthermore, the structure of DySample is as Figure 9As shown in the figure, where (a) is the overall structure diagram of DySample, (b) is the schematic diagram of the sampling point generator with a static range factor for DySample, and (c) is the schematic diagram of the sampling point generator with a dynamic range factor for DySample; before each DySample module in step 3.2 performs a fusion operation in the lightweight feature fusion network LDEFPN, it is possible to first use the DySample upsampling with a static range factor added to reconstruct the high-resolution feature map for the input feature map, or first use the DySample upsampling with a dynamic range factor added to reconstruct the high-resolution feature map for the input feature map;

[0116] Among them, before the lightweight feature fusion network LDEFPN performs a fusion operation, first using the DySample upsampling with a static range factor added to reconstruct the high-resolution feature map for the input feature map is expressed as:

[0117] x′ = grid_sample(x, s) (6)

[0118] s = g + o (7)

[0119] o = 0.25linear(x) (8)

[0120] In formulas (6) to (8), x represents the input feature map, s represents the upsampling scale factor; grid_sample represents the grid sampling operation, x′ represents the reconstructed high-resolution feature map; o represents the offset, g represents the original sampling grid; linear represents the linear layer operation;

[0121] Among them, before the lightweight feature fusion network LDEFPN performs a fusion operation, first using the DySample upsampling with a dynamic range factor added to reconstruct the high-resolution feature map for the input feature map is expressed as:

[0122] x′ = grid_sample(x, s) (9)

[0123] s = g + o (10)

[0124] o = 0.5sigmoid(linear1(x))·linear2(x) (11)

[0125] In formulas (9) to (11), x represents the input feature map, s represents the upsampling scale factor, grid_sample represents the grid sampling operation, x′ represents the reconstructed high-resolution feature map, o represents the offset, g represents the original sampling grid; linear represents the linear layer operation, and sigmoid is the sigmoid activation function;

[0126] Step 3.3: Construct an improved head network for the lightweight pedestrian detection model based on the improved YOLOv11n. The improved head network includes 4 detection heads;

[0127] The output of the C3k2 module 3 is used as the input of detection head 1, the output of the C3k2 module 4 is used as the input of detection head 2, the output of the C3k2 module 5 is used as the input of detection head 3, and the output of the Enhance_FPN module 6 is used as the output of detection head 4;

[0128] Step 4: Input the training set and the validation set into the improved model for training and save the trained model; Step 4 further includes steps 4.1 to 4.4:

[0129] Step 4.1: Set the training parameters of the lightweight pedestrian detection model based on the improved YOLOv11n. The training parameters include: number of epochs, batch size, optimizer, learning rate, momentum, weight decay, and number of threads;

[0130] In this embodiment, the training parameters include: the number of epochs Epoch is 300, the batch size batchsize is 16, the optimizer optimizer is SGD, the initial learning rate 1r0 is 0.01, the momentum momentum is 0.937, the weight decay weight_decay is 0.0005, and the number of threads workers is 8;

[0131] Step 4.2: Input the training set and validation set images and their corresponding labels into the lightweight pedestrian detection model based on the improved YOLOv11n, and use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters.

[0132] Specifically, in each training iteration, first, calculate the error between the predicted output of the model and the true label through forward propagation, and then use the error to calculate the loss function.

[0133] Secondly, use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters.

[0134] Thirdly, use the optimization algorithm to update the model parameters.

[0135] By repeatedly iterating this process, the model parameters will gradually be adjusted to the position that minimizes the loss function, thereby enabling the model to more accurately detect pedestrians in crowded scenes.

[0136] Step 4.3: Monitor the loss function value and performance metrics during training. The performance metrics include accuracy P, recall R, and mean average precision mAP. When the loss functions of the training set and the validation set no longer decrease, and at the same time, the evaluation metrics such as accuracy P, recall R, and mean average precision mAP no longer improve, stop the training to avoid overfitting the training data and ensure the generalization ability of the model on the test data.

[0137] Step 4.4: After training is completed, save the trained model and select the optimal model from it. Ensure that the model can be used at any time for pedestrian detection tasks in crowded scenes and can also be conveniently compared and evaluated with other models.

[0138] Step 5: Input the test set into the optimal model for testing. When the accuracy of the model meets the generalization requirements, obtain the final model. The said Step 5 further includes Steps 5.1 to 5.3:

[0139] Step 5.1: Input the said test set into the optimal model selected in Step 4 for testing.

[0140] Step 5.2: Calculate the model performance metrics: accuracy P, recall R, mean average precision mAP, number of parameters, computational complexity GFLOPs, frames per second FPS, and model size. The specific calculation formulas for the accuracy P, recall R, and mean average precision mAP are as follows:

[0141]

[0142] In Formulas (12) to (15), P is the accuracy, R is the recall, mAP is the mean average precision of all classes, AP is the average precision, m is the total number of pedestrian label classes, TP represents the number of positive samples correctly identified as positive samples, FP represents the number of negative samples misidentified as positive samples, FN represents the number of positive samples misidentified as negative samples, and AP i is the AP of the i-th class of pedestrians, and P(R) is the function of P varying with R.

[0143] Step 5.3: Evaluate the performance metrics of the test set. If the accuracy of the test set is similar to that of the training set, it indicates that the model meets the generalization requirements, and the final lightweight pedestrian detection model for crowded scenes based on the improved YOLOv11n is obtained.

[0144] In this embodiment, in order to verify the effect of the lightweight pedestrian detection model for crowded scenes based on the improved YOLOv11n disclosed in the present invention, the YOLOv5n model, GS-YOLOv5 model, YOLOv8n model, YOLOv10n model and YOLOv11n model are used to test with the lightweight pedestrian detection model for crowded scenes based on the improved YOLOv11n disclosed in the present invention on the CrowedHumen dataset, and the evaluation results are shown in Table 1.

[0145] Table 1 Comparison experiment results

[0146]

[0147] According to the data in Table 1, it can be seen that the lightweight pedestrian detection model for crowded scenes based on the improved YOLOv11n disclosed in the present invention has been improved in terms of accuracy P, recall rate R, mAP@0.5, mAP@0.5:0.95 and other indicators, indicating that the model has higher detection accuracy. At the same time, the improved model performs lower in terms of the number of parameters, model size, computational complexity GFLOPs and other performance indicators compared with the original YOLOv11n model. This improvement enables the lightweight pedestrian detection model for crowded scenes based on the improved YOLOv11n disclosed in the invention patent to have better lightweight characteristics and is easier to be deployed on mobile devices.

[0148] In this embodiment, in order to verify the effectiveness of the lightweight feature extraction network LDEFPN in the lightweight pedestrian detection model for crowded scenes based on the improved YOLOv11n disclosed in the present invention, taking YOLOv11n as the benchmark model, the lightweight feature extraction network LDEFPN disclosed in the present invention is used to compare with several improved neck networks. The improved neck networks include: the neck network of YOLOv11n, BiFPN, HSFPN, AFPN, GSPAN and the lightweight feature extraction network LDEFPN disclosed in the present invention. The experimental results are shown in Table 2.

[0149] Table 2 Comparison of LDEFPN effectiveness

[0150]

[0151] According to the data in Table 2, it can be seen that compared with other improved neck networks, using the lightweight feature extraction network LDEFPN disclosed in this invention patent to replace the original neck network in YOLOv11n, the indicators such as accuracy P, recall rate R, mAP@0.5, mAP@0.5:0.95, FPS, etc. have all improved, indicating that the lightweight feature extraction network LDEFPN disclosed in this invention has higher detection speed and accuracy. At the same time, the performance indicators such as the number of parameters, model size, computational complexity GFLOPs, etc. are lower, indicating that the lightweight feature extraction network LDEFPN disclosed in this invention has better lightweight characteristics and is easier to be deployed on mobile devices.

[0152] The above is only one embodiment of the present invention, and it does not limit the patent scope of the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A lightweight pedestrian detection method for crowded scenes based on improved YOLOv11n, characterized in that, Specifically, it includes the following steps: Step 1: Establish a pedestrian image dataset for crowded scenes to form a first dataset; In the first dataset, pedestrian images in crowded scenes can be collected through the network, taken with a digital camera, or obtained from surveillance videos; Step 2: Add annotation information to the images in the first dataset to form a second dataset, and divide the second dataset into a training set, a validation set, and a test set according to the ratio of 6:2:2; Step 3: Construct a lightweight pedestrian detection model for crowded scenes based on the improved YOLOv11n. The construction of this model further includes steps 3.1 to 3.3: Step 3.1: Construct an improved backbone network for the lightweight pedestrian detection model for crowded scenes based on the improved YOLOv11n. The improved backbone network consists of convolutional layer 1, convolutional layer 2, C3k2_EMBC module 1, convolutional layer 3, C3k2_EMBC module 2, convolutional layer 4, C3k2_EMBC module 3, convolutional layer 5, C3k2_EMBC module 4, SPPF module, and C2PSA module 1 connected in sequence; The improved backbone network outputs four different scales of feature information through C3k2_EMBC module 1, C3k2_EMBC module 2, C3k2_EMBC module 3, and C2PSA module 1 respectively; Step 3.2: Construct a lightweight feature fusion network LDEFPN for the lightweight pedestrian detection model for crowded scenes based on the improved YOLOv11n; The lightweight feature fusion network LDEFPN specifically includes 8 convolutional layers, 6 Enhance_FPN modules, 2 DySample modules, and 5 C3k2 modules; The output of C3k2_EMBC module 1 is used as the input of convolutional layer 6, and the output of convolutional layer 6 is used as the input of Enhance_FPN module 3; The output of C3k2_EMBC module 2 is used as the input of convolutional layer 7, and the output of convolutional layer 7 is used as the input of Enhance_FPN module 3; the output of convolutional layer 7 is used as the input of Enhance_FPN module 1, the output of Enhance_FPN module 1 is used as the input of C3k2 module 2, the output of C3k2 module 2 is used as the input of Enhance_FPN module 3, the output of Enhance_FPN module 3 is used as the input of C3k2 module 3, and the output of C3k2 module 3 is used as the input of convolutional layer 10; The output of the C3k2_EMBC module 3 serves as the input to the convolutional layer 8. The output of the convolutional layer 8 serves as the input to the Enhance_FPN module 2. The output of the Enhance_FPN module 2 serves as the input to the C3k2 module 1. The output of the C3k2 module 1 is fused with the output of the convolutional layer 7 in the Ehance_FPN module 1 through the DySample module 1. The outputs of the convolutional layer 8, the C3k2 module 1, and the convolutional layer 10 are fused in the Enhance_FPN module 4. The output of the Enhance_FPN module 4 serves as the input to the C3k2 module 4. The output of the C3k2 module 4 serves as the input to the convolutional layer 11. The output of the C2PSA module 1 serves as the input to the convolutional layer 9. The output of the convolutional layer 9 is fused with the output of the convolutional layer 8 in the Enhance_FPN module 2 through the DySample module 2. The output of the convolutional layer 9 is fused with the output of the convolutional layer 11 in the Enahnce_FPN module 5. The output of the Enahnce_FPN module 5 serves as the input to the C3k2 module 5. The output of the C3k2 module 5 serves as the input to the convolutional layer 13. The output of the convolutional layer 9 serves as the input to the convolutional layer 12. The outputs of the convolutional layer 12 and the convolutional layer 13 are fused in the Enhance_FPN module 6. Step 3.3: Construct an improved head network for the lightweight pedestrian detection model in crowded scenes based on the improved YOLOv11n. The improved head network includes 4 detection heads. The output of the C3k2 module 3 serves as the input to the detection head 1. The output of the C3k2 module 4 serves as the input to the detection head 2. The output of the C3k2 module 5 serves as the input to the detection head 3. The output of the Enhance_FPN module 6 serves as the input to the detection head 4. Step 4: Train the lightweight pedestrian detection model in crowded scenes based on the improved YOLOv11n described in Step 3 using the training set and the validation set, and save the trained model. Step 4 further includes Steps 4.1 to 4.4: Step 4.1: Set the training parameters of the lightweight pedestrian detection model in crowded scenes based on the improved YOLOv11n. The model training parameters include: number of iterations, batch size, optimizer, learning rate, momentum, weight decay, and number of threads. Step 4.2: Input the training set and validation set images and their corresponding labels into the lightweight pedestrian detection model in crowded scenes based on the improved YOLOv11n. Use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters, and update the model parameters according to the gradient to gradually reduce the loss function. Step 4.3: Monitor the loss function values and performance metrics during the training process. Stop the training when the loss functions of the training set and the validation set no longer decrease, and at the same time, evaluation metrics such as the accuracy P, recall rate R, and mean average precision mAP no longer improve, to avoid overfitting of the model. Step 4.4: After the training is completed, save the trained model and select the optimal model from it; Step 5: Use the test set to test the optimal model selected in Step 4.4, evaluate the test results of the test set, and if the accuracy of the model meets the generalization requirements, the final lightweight pedestrian detection model based on the improved YOLOv11n can be obtained.

2. The lightweight pedestrian detection method in a crowded scene based on the improved YOLOv11n according to claim 1, wherein In the C3k2_EMBC module in Step 3.1, when C3k_EMBC = False, the Bottleneck module in the original C3k2 module is replaced by the EMB convolution disclosed in this invention patent; when C3k_EMBC = True, the C3k module in the original C3k2 is replaced by the C3k_EMBC module; the C3k_EMBC module is to replace the Bottleneck module in the C3k module with the EMB convolution disclosed in this invention patent; The EMB convolution includes two convolutional layers, one depthwise separable convolution, one eSE module, and one Droupout module; The feature map input to the EMB convolution first passes through convolutional layer 1. The feature map output by convolutional layer 1 passes through normalization and the Swish activation function and is input into the depthwise separable convolution. The feature map output by the depthwise separable convolution passes through normalization and the Swish activation function and is input into the eSE module. The feature map output by the eSE module is used as the input of convolutional layer 2. The feature map output by convolutional layer 2 passes through normalization and is used as the input of the Droupout module. The feature map output by the Droupout module is aggregated with the original input feature map, and the aggregated feature map is used as the output feature map of the EMB convolution.

3. A lightweight pedestrian detection method in crowded scenes based on the improved YOLOv11n according to claim 1, characterized in that, The Enhance_FPN module in Step 3.2 includes a weighted fusion module and a lightweight spatial attention mechanism; the Enhance_FPN module 1, Enhance_FPN module 2, Enhance_FPN module 5, and Enhance_FPN module 6 have 2 inputs, and the Enhance_FPN module 3 and Enhance_FPN module 4 have 3 inputs; the internal processes of the Enhance_FPN module with 2 inputs and the Enhance_FPN module with 3 inputs further include Steps 3.2.1 to 3.2.5: Step 3.2.1: For the Enhance_FPN module with 2 inputs, the input feature maps are feature map X0 and feature map X1 respectively; for the Enhance_FPN module with 3 inputs, the input feature maps are feature map X0, feature map X1, and feature map X2 respectively; Step 3.2.2: For the Enhance_FPN module with 2 inputs, according to the importance of X0 and X1, assign weights W0 and W1 to them, multiply X0 by its weight W0, multiply X1 by its weight W1, and add the weighted feature maps to obtain the weighted sum. The specific calculation formula for the weighted sum is: O = X0·W0 + X1·W1 (1) In formula (1), O is the output feature after weighted summation, X0 and X1 are the input feature maps of the Enhance_FPN module with 2 inputs respectively, W0 is the weight of feature map X0, and W1 is the weight of feature map X1; For the Enhance_FPN module with 3 inputs, according to the importance of X0, X1, and X2, weights W0, W1, and W2 are assigned to it. Multiply X0 by its weight W0, multiply X1 by its weight W1, multiply X2 by its weight W2, and add the weighted feature maps to obtain the weighted sum. The specific calculation formula for the weighted sum is: O = X0·W0 + X1·W1 + X2·W2 (2) In formula (2), O is the output feature after weighted summation, X0, X1, and X2 are the input feature maps of the Enhance_FPN module with 3 inputs respectively, W0 is the weight of feature map X0, W1 is the weight of feature map X1, and W2 is the weight of feature map X2; Step 3.2.3: The weighted sum introduces non-linear features through the Swish activation function. The calculation formula after being processed by the Swish activation function is: fused_feature = Swish(O) (3) In formula (3), O is the output feature after weighted summation, Swish is the Swish activation function, and fused_feature is the feature map processed by the Swish activation function; Step 3.2.4: Use the feature map fused_feature processed by the Swish activation function as the input feature map of the convolutional layer in the lightweight spatial attention mechanism. The output feature map LSA of the convolutional layer is processed by the Sigmoid activation function. The calculation formula after being processed by the Sigmoid activation function is: attention_map = Sigmoid(LSA) (4) In formula (4), LSA is the output feature map of the convolutional layer in the lightweight spatial attention mechanism, Sigmoid is the Sigmoid activation function, and attention_map is the feature map processed by the Sigmoid activation function; Step 3.2.5: Element-wise multiply the feature map attention_map processed by the Sigmoid activation function by the feature map fused_feature processed by the Swish activation function. The feature map obtained after element-wise multiplication is used as the final output feature map of the Enhance_FPN module. The specific calculation formula for element-wise multiplication is: Enhance_feature = fused_feature * attention_map (5) In formula (5), fused_feature is the feature map processed by the Swish activation function, and attention_map is the feature map processed by the Sigmoid activation function; Enhance_feature is the final output feature map of the Enhance_FPN module.

4. A lightweight pedestrian detection method in crowded scenes based on the improved YOLOv11n according to claim 1, characterized in that, In step 3.2, before each DySample module in the lightweight feature fusion network LDEFPN performs the fusion operation, it is possible to first use the DySample upsampling with a static range factor added to reconstruct the high-resolution feature map of the input feature map, or first use the DySample upsampling with a dynamic range factor added to reconstruct the high-resolution feature map of the input feature map; Among them, before the lightweight feature fusion network LDEFPN performs the fusion operation, first using the DySample upsampling with a static range factor added to reconstruct the high-resolution feature map of the input feature map is expressed as: x′ = grid_sample(x, s) (6) s = g + o (7) o = 0.25linear(x) (8) In formulas (6) to (8), x represents the input feature map, s represents the upsampling scale factor; grid_sample represents the grid sampling operation, x′ represents the reconstructed high-resolution feature map; o represents the offset, g represents the original sampling grid; linear represents the linear layer operation; Among them, before the lightweight feature fusion network LDEFPN performs the fusion operation, first using the DySample upsampling with a dynamic range factor added to reconstruct the high-resolution feature map of the input feature map is expressed as: x′ = grid_sample(x, s) (9) s = g + o (10) o = 0.5sigmoid(linear1(x)) · linear2(x) (11) In formulas (9) to (11), x represents the input feature map, s represents the upsampling scale factor; grid_sample represents the grid sampling operation, x′ represents the reconstructed high-resolution feature map; o represents the offset, g represents the original sampling grid; linear represents the linear layer operation, and sigmoid is the sigmoid activation function.

5. A lightweight pedestrian detection method for crowded scenes based on the improved YOLOv11n according to claim 1, characterized in that, Step 5 further includes steps 5.1 to 5.3: Step 5.1: Input the test set into the optimal model selected in step 4 for testing; Step 5.2: Calculate the model performance metrics: accuracy P, recall R, mean average precision mAP, number of parameters, computational complexity GFLOPs, frames per second FPS, and model size. The specific calculation formulas for the accuracy P, recall R, and mean average precision mAP are as follows: In formulas (12) to (15), P is the accuracy rate, R is the recall rate, mAP is the mean average precision of all categories, AP is the average precision, m is the total number of pedestrian label categories, TP represents the number of positive samples correctly identified as positive samples, FP represents the number of negative samples misidentified as positive samples, FN represents the number of positive samples misidentified as negative samples, and AP i is the AP of the i-th category of pedestrians, and P(R) is the function of P changing with R; Step 5.3: Evaluate the performance metrics of the test set. If the accuracy of the test set is similar to that of the training set, it indicates that the model meets the generalization requirements, and the final lightweight pedestrian detection model based on the improved YOLOv11n is obtained.

Citation Information

Patent Citations

  • Dense pedestrian detection system based on improved lightweight YOLOv7

    CN116612427A

  • Pedestrian detection method based on YOLOv7

    CN117542082A

  • Improved YOLOv8n-based crowded scene pedestrian detection method

    CN119540869A

Cited By

  • Communication big data-driven sewage plant peak clipping and valley filling low-carbon scheduling method

    CN121504031A