A Pedestrian Detection Method in Crowded Scenes Based on Improved YOLOv8n
By improving the YOLOv8n algorithm, a pedestrian detection model of the feature fusion network DCPAN and DiFPN modules was constructed, which solved the missed detection and misdetection problems caused by occlusion and multi-scale problems in crowded scenarios, and achieved higher detection accuracy.
Patent Information
- Application Number
- CN202411691449.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-11-25
AI Technical Summary
In crowded scenarios, existing pedestrian detection models are prone to missed or missed detection when dealing with occlusion and multi-scale problems, and it is difficult to accurately identify small-scale occluded pedestrians.
By improving the YOLOv8n algorithm, a pedestrian detection model based on the improved YOLOv8n is built, and the feature fusion network DCPAN and DiFPN modules are used to enhance the network's detection ability of pedestrians at different scales.
The model's sensitivity to small-scale pedestrian targets and detection ability of obscured pedestrians is significantly improved, the missed detection and missed detection are reduced, and the detection accuracy is improved.
Smart Images

Figure CN119540869B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to a pedestrian detection method in crowded scenes based on improved YOLOv8n. Background Art
[0002] In recent years, due to the improvement of people's material living standards and the acceleration of the urbanization process, the number of urban residents has increased rapidly, resulting in frequent crowd gatherings in shopping malls, stations, streets, hospitals, scenic spots and other places. There are many potential safety hazards in these crowded environments. For example, stampede accidents, collision injuries, traffic congestion, etc. These potential safety hazards often cause serious property losses and casualties. In order to effectively address these problems and ensure public safety, it has become particularly important to develop a model that can accurately and quickly identify pedestrian targets in crowded scenes.
[0003] However, pedestrian detection in crowded scenes faces multiple challenges. In addition to the influence of common factors such as lighting changes, complex and diverse backgrounds, and different shooting angles, the characteristics of the human body itself, such as structural similarity, pose diversity, scale differences, and different clothing, all increase the difficulty of detection. What is more difficult is that there is generally occlusion among pedestrians, which makes it difficult for the model to accurately separate the target object from the background and other pedestrians, thereby affecting the quality of the detection results.
[0004] Currently, the safety monitoring methods for people in crowded scenes are mainly divided into the following three types, namely manual inspection, traditional pedestrian detection methods, and deep learning-based pedestrian detection methods. Manual inspection relies on security personnel to patrol on site or observe in real time through a video monitoring system. Although this method is intuitive and flexible, its efficiency is low, and it is easy to cause negligence or omission due to human factors. Traditional pedestrian detection methods obtain candidate regions by traversing with a sliding window, extract pedestrian features manually, train a classifier, and use the trained classifier to classify the scanned window and output the target detection result. Classic algorithms include HOG+SVM, Haar+Adaboost, etc.; these detection algorithms have disadvantages such as high dimension, large computational amount, time-consuming and laborious, weak feature extraction ability, and poor generalization. With the rapid development of deep learning, deep learning models have been widely used in the field of pedestrian detection. Deep learning methods are a class of neural network algorithms that autonomously learn internal hierarchical features of images from a large amount of training data, avoiding the cumbersome and inaccuracy of manual feature extraction in traditional methods. Representative algorithms include CNN, YOLO, SSD, etc.
[0005] In a crowded scene, when there is severe occlusion between pedestrians, the occluded pedestrians at a small scale only contain a single pixel or a few pixels in the feature map, which causes the model to be more inclined to capture the features of large-scale pedestrians and ignore the features of small-scale pedestrians. When existing models handle such problems, there will still be cases of missed detection or false detection. Therefore, it is necessary to propose a pedestrian detection method to reduce the impact of occlusion and multi-scale problems in crowded scenes. Summary of the Invention
[0006] Aiming at the problems existing in the prior art, the present invention provides a pedestrian detection method for crowded scenes based on improved YOLOv8n. By improving the YOLOv8n algorithm, it reduces the problems of false detection and missed detection in the pedestrian detection model for crowded scenes caused by factors such as severe occlusion and diverse scale changes between pedestrians in crowded scenes.
[0007] The technical solution provided by the present invention includes the following steps:
[0008] Step 1: Obtain pedestrian images in a crowded scene to form a first data set;
[0009] Step 2: Add annotation information to the images in the first data set to form a second data set, and divide the second data set into a training set, a validation set, and a test set;
[0010] Step 3: Construct a pedestrian detection model for crowded scenes based on improved YOLOv8n. The model includes a backbone network, a feature fusion network DCPAN, and a head network. The construction of the model further includes steps 3.1 to 3.3:
[0011] Step 3.1: The backbone network consists of a convolutional layer 1, a convolutional layer 2, a C2f module 1, a convolutional layer 3, a C2f module 2, a convolutional layer 4, a C2f module 3, a convolutional layer 5, a C2f module 4, and an SPPF module connected in sequence;
[0012] Use the training set and the validation set in the second data set as the input of the backbone network;
[0013] The backbone network outputs four different scales of feature information through the C2f module 1, the C2f module 2, the C2f module 3, and the SPPF module respectively;
[0014] Step 3.2: The feature fusion network DCPAN, on the basis of the YOLOv8n neck network, adds a convolutional layer 7, a convolutional layer 8, a convolutional layer 9, and a new feature scale fusion layer, and uses the DiFPN module proposed by this invention patent to replace the Concat module of the YOLOv8n neck network;
[0015] The new feature scale fusion layer consists of a convolutional layer 6, an upsampling layer 1, a DiFPN module 3, a C2f module 7, and a convolutional layer 10 in sequence;
[0016] The output of the SPPF module is used as the input of the convolutional layer 9, and the output of the convolutional layer 9 is transmitted to the top-down path. At the same time, the convolutional layer 9 is horizontally connected to the DiFPN module 6 in the bottom-up path;
[0017] The output of the C2f module 3 is used as the input of the convolutional layer 8, and the output of the convolutional layer 8 is used as the input of the DiFPN module 2. At the same time, the convolutional layer 8 is horizontally connected to the DiFPN module 5 in the bottom-up path;
[0018] The output of the C2f module 2 is used as the input of the convolutional layer 7, and the output of the convolutional layer 7 is used as the input of the DiFPN module 1. At the same time, the convolutional layer 7 is horizontally connected to the DiFPN module 4 in the bottom-up path;
[0019] The output of the C2f module 1 is used as the input of the convolutional layer 6, and the convolutional layer 6 is horizontally connected to the DiFPN module 3 in the bottom-up path. At the same time, the output of the upsampling layer 1 is also used as the input of the DiFPN module 3; the output of the DiFPN module 3 is used as the input of the C2f module 7, the output of the C2f module 7 is used as the input of the convolutional layer 10, and the output of the convolutional layer 10 is used as the input of the DiFPN module 4;
[0020] Step 3.3: The head network includes 4 detection heads; the output of the C2f module 10 is used as the input of the detection head 4; the output of the C2f module 9 is used as the input of the detection head 3; the output of the C2f module 8 is used as the input of the detection head 2; the output of the C2f module 7 is used as the input of the detection head 1;
[0021] Step 4: Use the training set and the validation set to train the pedestrian detection model based on the improved YOLOv8n and save the trained model parameters as the optimal model;
[0022] Step 5: Use the test set to test the optimal model, evaluate the test results of the test set with objective evaluation indicators, meet the accuracy requirements, and obtain the final pedestrian detection model based on the improved YOLOv8n for crowded scenes.
[0023] Furthermore, in step 1, the images in the first dataset can be collected through the network, taken with a digital camera, or obtained from surveillance videos;
[0024] Preferably, in step 2, the training set, the validation set, and the test set can be divided in a ratio of 6:2:2.
[0025] Further, the DiFPN module 1, DiFPN module 2, DiFPN module 3, and DiFPN module 6 in step 3.2 have 2 inputs. The DiFPN module with 2 inputs includes 2 DCNv4 modules and 1 weighted fusion module; the DiFPN module 4 and DiFPN module 5 have 3 inputs, and the DiFPN module with 3 inputs includes 3 DCNv4 modules and 1 weighted fusion module. The internal processes of the DiFPN module with 2 inputs and the DiFPN module with 3 inputs further include steps 3.2.1 to 3.2.7:
[0026] Step 3.2.1: For the DiFPN module with 2 inputs, the input feature maps are feature map X 0 and feature map X 1 ; the feature map X 0 and feature map X 1 are respectively input into the 2 DCNv4 modules in the DiFPN module with 2 inputs;
[0027] For the DiFPN module with 3 inputs, the input feature maps are feature map X 0 , feature map X 1 and feature map X 2 ; the feature map X 0 , feature map X 1 and feature map X 2 are respectively input into the 3 DCNv4 modules in the DiFPN module with 3 inputs;
[0028] Step 3.2.2: The DCNv4 module uses 1 convolutional layer to predict the offset of each pixel position in the input feature map;
[0029] Step 3.2.3: According to the offset, dynamically adjust the position of the sampling points. For the position (i, j) of each convolutional kernel, calculate the new sampling position ( , ):
[0030] (1)
[0031] In formula (1), offset(i, j) is the predicted offset;
[0032] Step 3.2.4: For the DiFPN module with 2 inputs, the feature map X 0 , feature map X 1 and offset are jointly used as the input of the deformable convolutional layer. Finally, generate 2 deformed feature maps and ;
[0033] For the DiFPN module with 3 inputs, the feature maps X 0 , the feature map X 1 , the feature map X 2 and the offset are jointly used as the inputs of the deformable convolutional layer. Finally, 3 deformed feature maps are generated , and ;
[0034] Step 3.2.5: For the DiFPN module with 2 inputs, input the 2 deformed feature maps and into the weighted fusion module. According to the importance of and , assign weights W 0 and W 1 to them. Multiply by its weight W 0 , multiply by its weight W 1 , and add the weighted feature maps to obtain the weighted sum. The calculation formula of the weighted sum is:
[0035] (2)
[0036] In formula (2), O is the output feature after weighted fusion, W i is the weight, is the output feature of DCNv4, and i = 0, 1;
[0037] For the DiFPN module with 3 inputs, input the 3 deformed feature maps , and into the weighted fusion module. According to the importance of , and , assign weights W 0 , W 1 and W 2 to them. Multiply by its weight W 0 , multiply by its weight W 1 , multiply by its weight W 2 , and add the weighted feature maps to obtain the weighted sum. The calculation formula of the weighted sum is:
[0038] (3)
[0039] In formula (3), O is the output feature after weighted fusion, Wi is the weight, is the output feature of DCNv4, i = 0, 1, 2;
[0040] Step 3.2.6: Introduce non - linear features to the weighted sum through the Swish activation function;
[0041] Step 3.2.7: For the DiFPN module with 2 inputs, use a 1×1 convolution to operate on the activated feature map to obtain the output feature map of the DiFPN module with 2 inputs;
[0042] For the DiFPN module with 3 inputs, use a 1×1 convolution to operate on the activated feature map to obtain the output feature map of the DiFPN module with 3 inputs.
[0043] Furthermore, the feature fusion method of the lateral cross - connection in Step 3.2 is as follows:
[0044] For the i - th layer, obtain the i - th layer feature map input to this layer and the (i + 1)-th layer feature map according to the top - down and bottom - up feature propagation and the (i + 1)-th layer feature map , upsample the (i + 1)-th layer feature map to adjust it to the same size as the i - th layer feature map , perform weighted fusion and normalization on the upsampled (i + 1)-th layer feature map and the i - th layer feature map to obtain the intermediate feature map of the i - th layer in the top - down path and the output feature map of the i - th layer in the bottom - up path;
[0045] The feature fusion expression of the intermediate feature map of the i - th layer in the top - down path is:
[0046] (4)
[0047] In formula (4), ω 1 represents the weight of the i - th layer feature map, ω 2 represents the weight of the upsampled (i + 1)-th layer feature map, ε is a parameter used to prevent the denominator from being 0, and Conv represents feature fusion;
[0048] The feature fusion expression of the output feature map of the i - th layer in the bottom - up path is:
[0049] (5)
[0050] In formula (5), represents the output feature map of the (i - 1)-th layer in the top - down path, represents adjusting the output feature map of the (i - 1)-th layer by upsampling to the same size as the i - th layer feature map, denotes the weight of the i-th layer feature map of the input denotes the weight of the intermediate feature map of the i-th layer in the bottom-up path denotes the weight of the upsampled (i - 1)-th layer feature map
[0051] Furthermore, the training process of the pedestrian detection model for crowded scenes based on the improved YOLOv8n in step 4 specifically includes steps 4.1 to 4.4:
[0052] Step 4.1: Set the training parameters of the pedestrian detection model for crowded scenes based on the improved YOLOv8n;
[0053] The model training parameters include: number of iterations, batch size, optimizer, learning rate, momentum, weight decay, and number of threads;
[0054] Step 4.2: Input the training set, validation set, and corresponding labels into the pedestrian detection model for crowded scenes based on the improved YOLOv8n, and use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters;
[0055] Step 4.3: Use the optimizer to update the model parameters in the direction of gradient descent; until the loss functions of the training set and validation set no longer decrease, and at the same time, evaluation metrics such as accuracy P, recall R, and mAP no longer improve;
[0056] Step 4.4: Save the trained model parameters as the optimal model.
[0057] Furthermore, step 5 specifically includes steps 5.1 to 5.3:
[0058] Step 5.1: Input the test set into the optimal model in step 4;
[0059] Step 5.2: Calculate the model performance metrics: The specific calculation formulas for accuracy P, recall R, and mean average precision mAP are as follows:
[0060] (6)
[0061] (7)
[0062] (8)
[0063] (9)
[0064] In Formulas (6) to (9), P is the accuracy rate, R is the recall rate, mAP is the mean average precision of all categories, AP is the average precision, m is the total number of pedestrian label categories, TP represents the number of positive samples correctly identified as positive samples, FP represents the number of negative samples misidentified as positive samples, and FN represents the number of positive samples misidentified as negative samples;
[0065] Step 5.3: Evaluate the objective evaluation indicators for the test results of the test set. If the accuracy requirements are met, the final pedestrian detection model based on the improved YOLOv8n for crowded scenes is obtained.
[0066] Compared with the prior art, the beneficial effects of the present invention are:
[0067] In the feature fusion network DCPAN disclosed in this invention patent, the horizontal cross-connection effectively retains and utilizes the important detailed information in the shallow layer, enabling the model to not only maintain good detection of large-scale pedestrian targets but also significantly improve the sensitivity to small-scale pedestrian targets;
[0068] The DiFPN module disclosed in this invention patent can fully fuse the feature information from different levels, enhancing the adaptability of the network to the shape and position of small-scale pedestrians and occluded pedestrians;
[0069] In the feature fusion network DCPAN disclosed in this invention patent, 4 feature scale fusion layers are used to detect pedestrian targets of different scales, and the strategy of using different feature scale fusion layers to process pedestrian targets of different scales is adopted to improve the detection accuracy of the model in detecting pedestrian targets of different scales. Brief Description of the Drawings
[0070] Figure 1 It is a flowchart of the pedestrian detection method for crowded scenes based on the improved YOLOv8n of the present invention;
[0071] Figure 2 It is a schematic diagram of the model structure of the pedestrian detection method for crowded scenes based on the improved YOLOv8n of the present invention;
[0072] Figure 3 It is a schematic diagram of the DiFPN structure with 2 inputs;
[0073] Figure 4 It is a schematic diagram of the DiFPN structure with 3 inputs; Detailed Description of the Invention
[0074] In order to make the technical solutions, structural features, achieved purposes and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in combination with specific embodiments and with reference to the accompanying drawings; it should be noted that the specific embodiments described herein are only used to explain the present invention more clearly and are not used to limit the present invention.
[0075] Figure 1 This is a flowchart of a pedestrian detection method in crowded scenes based on improved YOLOv8n disclosed by the present invention, and its implementation process is as follows:
[0076] Step 1: Obtain pedestrian images in crowded scenes to form a first dataset; in the first dataset, the pedestrian images in crowded scenes can be obtained by collecting through the network, taking pictures with a digital camera, or obtaining from surveillance videos.
[0077] In this embodiment, in order to better evaluate the detection effect of a pedestrian detection method in crowded scenes based on improved YOLOv8n disclosed by the present invention, the publicly available dataset CrowedHumen dataset is adopted; the first dataset is formed by using the CrowedHumen dataset.
[0078] Step 2: Add annotation information to the images in the first dataset to form a second dataset, and divide the second dataset into a training set, a validation set, and a test set.
[0079] Since there is already annotation information in the publicly available dataset CrowedHumen adopted in this embodiment, this step is skipped; the label files in odgt format in the first dataset in this embodiment are converted into txt format label files required by YOLOv8n to form a second dataset, and the second dataset is divided into a training set, a validation set, and a test set according to 6:2:2; the divided training set includes 15,000 pictures, the validation set includes 4,370 pictures, and the test set includes 4,370 pictures. The image resolution in the second dataset in this embodiment is 640×640.
[0080] Step 3: Build a pedestrian detection model in crowded scenes based on improved YOLOv8n. The improved YOLOv8n model structure is as Figure 2 shown. The model includes a backbone network, a feature fusion network DCPAN, and a head network. The construction of the model further includes steps 3.1 to 3.3:
[0081] Step 3.1: The backbone network is composed of a convolutional layer 1, a convolutional layer 2, a C2f module 1, a convolutional layer 3, a C2f module 2, a convolutional layer 4, a C2f module 3, a convolutional layer 5, a C2f module 4, and an SPPF module connected in sequence.
[0082] Take the training set and the validation set in the second dataset as the input of the backbone network.
[0083] The backbone network outputs 4 different scales of feature information through the C2f module 1, the C2f module 2, the C2f module 3, and the SPPF module respectively.
[0084] Step 3.2: Based on the YOLOv8n neck network, the feature fusion network DCPAN adds convolutional layer 7, convolutional layer 8, convolutional layer 9, and a new feature scale fusion layer, and uses the DiFPN module disclosed in this invention patent to replace the Concat module of the YOLOv8n neck network;
[0085] The new feature scale fusion layer consists of convolutional layer 6, upsampling 1, DiFPN module 3, C2f module 7, and convolutional layer 10 in sequence;
[0086] The output of the SPPF module is used as the input of convolutional layer 9, and the output of convolutional layer 9 is transmitted to the top - down path. At the same time, convolutional layer 9 is horizontally connected to DiFPN module 6 in the bottom - up path;
[0087] The output of C2f module 3 is used as the input of convolutional layer 8, the output of convolutional layer 8 is used as the input of DiFPN module 2, and at the same time, convolutional layer 8 is horizontally connected to DiFPN module 5 in the bottom - up path;
[0088] The output of C2f module 2 is used as the input of convolutional layer 7, the output of convolutional layer 7 is used as the input of DiFPN module 1, and at the same time, convolutional layer 7 is horizontally connected to DiFPN module 4 in the bottom - up path;
[0089] The output of C2f module 1 is used as the input of convolutional layer 6, convolutional layer 6 is horizontally connected to DiFPN module 3 in the bottom - up path, and at the same time, the output of upsampling 1 is also used as the input of DiFPN module 3; the output of DiFPN module 3 is used as the input of C2f module 7, the output of C2f module 7 is used as the input of convolutional layer 10, and the output of convolutional layer 10 is used as the input of DiFPN module 4;
[0090] Further, the structure of the DiFPN module with 2 inputs is as Figure 3 shown, and the structure of the DiFPN module with 3 inputs is as Figure 4 shown; DiFPN module 1, DiFPN module 2, DiFPN module 3, and DiFPN module 6 in Step 3.2 have 2 inputs. The DiFPN module with 2 inputs contains 2 DCNv4 modules and 1 weighted fusion module; DiFPN module 4 and DiFPN module 5 have 3 inputs. The DiFPN module with 3 inputs contains 3 DCNv4 modules and 1 weighted fusion module; The internal processes of the DiFPN module with 2 inputs and the DiFPN module with 3 inputs further include steps 3.2.1 to 3.2.7:
[0091] Step 3.2.1: For the DiFPN module with 2 inputs, the input feature maps are feature map X 0 and feature map X 1 ; Input the feature map X 0 and feature map X 1 into the 2 DCNv4 modules in the DiFPN module with 2 inputs respectively;
[0092] For the DiFPN module with 3 inputs, the input feature maps are feature map X 0 , feature map X 1 and feature map X 2 ; Input the feature map X 0 , feature map X 1 and feature map X 2 into the 3 DCNv4 modules in the DiFPN module with 3 inputs respectively;
[0093] Step 3.2.2: The DCNv4 module uses 1 convolutional layer to predict the offset of each pixel position in the input feature map;
[0094] Step 3.2.3: Dynamically adjust the position of the sampling points according to the offset. For the position (i, j) of each convolution kernel, calculate the new sampling position ( , ):
[0095] (1)
[0096] In formula (1), offset(i, j) is the predicted offset;
[0097] Step 3.2.4: For the DiFPN module with 2 inputs, the feature map X 0 , feature map X 1 and offset are jointly used as the input of the deformable convolutional layer. Finally, generate 2 deformed feature maps and ;
[0098] For the DiFPN module with 3 inputs, the feature map X 0 , feature map X 1 , feature map X 2 and offset are jointly used as the input of the deformable convolutional layer. Finally, generate 3 deformed feature maps , and ;
[0099] Step 3.2.5: For the DiFPN module with 2 inputs, the 2 deformed feature maps and Input into the weighted fusion module, according to and importance, assign weights W 0 and W 1 , multiply by its weight W 0 , multiply by its weight W 1 , and add the weighted feature maps to obtain a weighted sum. The calculation formula of the weighted sum is:
[0100] (2)
[0101] In formula (2), O is the output feature after weighted fusion, W i is the weight, is the output feature of DCNv4, i = 0, 1;
[0102] For the DiFPN module with 3 inputs, input the 3 deformed feature maps , and into the weighted fusion module. According to , and importance, assign weights W 0 , W 1 and W 2 , multiply by its weight W 0 , multiply by its weight W 1 , multiply by its weight W 2 , and add the weighted feature maps to obtain a weighted sum. The calculation formula of the weighted sum is:
[0103] (3)
[0104] In formula (3), O is the output feature after weighted fusion, W i is the weight, is the output feature of DCNv4, i = 0, 1, 2;
[0105] Step 3.2.6: Introduce non-linear features by passing the weighted sum through the Swish activation function;
[0106] Step 3.2.7: For the DiFPN module with 2 inputs, perform operations on the activated feature map using a 1×1 convolution to obtain the output feature map of the DiFPN module with 2 inputs;
[0107] For the DiFPN module with three inputs, a 1×1 convolutional layer is used to operate on the activated feature maps to obtain the output feature maps of the DiFPN module with three inputs.
[0108] Further, the lateral skip connections are as shown by the dashed lines in Figure 2 . The feature fusion method of the lateral skip connections is as follows:
[0109] For the i-th layer, the i-th layer feature map input to this layer and the (i + 1)-th layer feature map are obtained according to the top-down and bottom-up feature propagations. The (i + 1)-th layer feature map is upsampled to adjust to the same size as the i-th layer feature map. The upsampled (i + 1)-th layer feature map and the i-th layer feature map are weighted fused and normalized to obtain the intermediate feature map of the i-th layer in the top-down path and the output feature map of the i-th layer in the bottom-up path.
[0110] The feature fusion expression of the intermediate feature map of the i-th layer in the top-down path is:
[0111] (4)
[0112] In formula (4), ω 1 represents the weight of the i-th layer feature map, ω 2 represents the weight of the upsampled (i + 1)-th layer feature map, ε is a parameter used to prevent the denominator from being 0, and Conv represents feature fusion.
[0113] The feature fusion expression of the output feature map of the i-th layer in the bottom-up path is:
[0114] (5)
[0115] In formula (5), represents the output feature map of the (i - 1)-th layer in the top-down path, represents the upsampling of the output feature map of the (i - 1)-th layer to adjust to the same size as the i-th layer feature map, represents the weight of the input i-th layer feature map, represents the weight of the intermediate feature map of the i-th layer in the bottom-up path, represents the weight of the upsampled (i - 1)-th layer feature map.
[0116] Step 3.3: The head network includes 4 detection heads; the output of the C2f module 10 is used as the input of detection head 4; the output of the C2f module 9 is used as the input of detection head 3; the output of the C2f module 8 is used as the input of detection head 2; the output of the C2f module 7 is used as the input of detection head 1;
[0117] Step 4: Use the training set and the validation set to train the pedestrian detection model for crowded scenes based on the improved YOLOv8n, and save the trained model as the optimal model; the training process of the pedestrian detection model for crowded scenes based on the improved YOLOv8n further includes steps 4.1 to 4.4:
[0118] Step 4.1: Set the training parameters of the pedestrian detection model for crowded scenes based on the improved YOLOv8n;
[0119] In this embodiment, the training parameters include: the number of epochs Epoch is 300, the batch size batchsize is 4, the optimizer optimizer is SGD, the initial learning rate 1r0 is 0.01, the momentum momentum is 0.937, the weight decay weight_decay is 0.0005, and the number of threads workers is 4;
[0120] Step 4.2: Input the training set and validation set images and their corresponding labels into the pedestrian detection model for crowded scenes based on the improved YOLOv8n, and use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters; the backpropagation algorithm is an effective method for calculating gradients, which uses the chain rule to calculate the gradient of each parameter with respect to the loss function; specifically, backpropagation calculates the gradient of each parameter layer by layer by propagating the loss function backward from the output layer; in this process, the gradient of each parameter represents the rate of change of the loss function with respect to that parameter, that is, how the loss function changes as the parameter changes; by minimizing the loss function, the model parameters are adjusted to gradually approach the optimal solution;
[0121] Step 4.3: After calculating the gradients of the model parameters, use the optimizer to update these parameters; the optimizer updates the parameters in the opposite direction of the gradient according to the gradient information of the parameters; parameters with larger gradients will be updated with larger steps, while parameters with smaller gradients will be updated with smaller steps; by continuously iteratively updating the model parameters, the value of the loss function can be gradually reduced; by minimizing the loss function, the values of the model parameters are adjusted to gradually approach the optimal solution, that is, the parameter values at which the loss function reaches the minimum value, and the difference between the prediction results of the model and the true values is minimized, and at the same time, evaluation metrics such as mAP, recall rate R, and accuracy P no longer improve;
[0122] Step 4.4: Save the trained model parameters as the optimal model;
[0123] Step 5: Use the test set to test the optimal model. If the test results meet the accuracy requirements, the final model is obtained. Specifically, Step 5 further includes Steps 5.1 to 5.3:
[0124] Step 5.1: Input the test set into the optimal model described in Step 4;
[0125] Step 5.2: Calculate the model performance metrics: accuracy P, recall R, and mean average precision mAP. The specific calculation formulas are as follows:
[0126] (6)
[0127] (7)
[0128] (8)
[0129] (9)
[0130] In formulas (6) to (9), P is the accuracy, R is the recall, mAP is the mean average precision of all classes, AP is the average precision, m is the total number of pedestrian label classes, TP represents the number of positive samples correctly identified as positive samples, FP represents the number of negative samples misidentified as positive samples, and FN represents the number of positive samples misidentified as negative samples;
[0131] Step 5.3: When the performance metrics meet the accuracy requirements, obtain the final pedestrian detection model for crowded scenes based on the improved YOLOv8n.
[0132] In this embodiment, to verify the effectiveness of the detection model disclosed in this invention patent, this article uses the YOLOv8n model, YOLOv8n model + BiFPN, YOLOv8n model + GlodFPN, YOLOv8n model + RepGFPN, YOLOv7-tiny model, YOLOv10n model, YOLOv11n model, GR-yolo, and the detection model disclosed in this invention patent to conduct tests on the CrowedHumen dataset. The evaluation results are shown in Table 1. Among them, the proposed pedestrian detection model for crowded scenes based on the improved YOLOv8n is superior to other comparison models in terms of the evaluation metrics of accuracy P, recall R, mAP@0.5, and mAP@0.5:0.95:
[0133]
[0134] The above is only one embodiment of the present invention, and thus does not limit the patent scope of the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A pedestrian detection method in crowded scenes based on improved YOLOv8n, characterized in that: The specific steps include: Step 1: Obtain pedestrian images in a crowded scene to form a first data set; the images in the first data set can be collected through the Internet, photographed with a digital camera, or obtained from surveillance videos; Step 2: adding annotation information to the images in the first data set to form a second data set, and dividing the second data set into a training set, a validation set, and a test set; Step 3: Construct a crowded scene pedestrian detection model based on improved YOLOv8n, the model includes a backbone network, a feature fusion network DCPAN and a head network, and the construction of the model further includes steps 3.1 to 3.3: Step 3.1: The backbone network consists of a convolutional layer 1, a convolutional layer 2, a C2f module 1, a convolutional layer 3, a C2f module 2, a convolutional layer 4, a C2f module 3, a convolutional layer 5, a C2f module 4 and an SPPF module connected in sequence; Using the training set and the validation set in the second data set as input of the backbone network; The backbone network outputs feature information of four different scales through C2f module 1, C2f module 2, C2f module 3 and SPPF module respectively; Step 3.2: The feature fusion network DCPAN adds convolutional layer 7, convolutional layer 8, convolutional layer 9 and a new feature scale fusion layer on the basis of the YOLOv8n neck network, and uses the DiFPN module to replace the Concat module of the YOLOv8n neck network; The new feature scale fusion layer is composed of a convolution layer 6, an upsampling layer 1, a DiFPN module 3, a C2f module 7, and a convolution layer 10 in sequence; The output of the SPPF module is used as the input of the convolutional layer 9, and the output of the convolutional layer 9 is transmitted to the top-down path. At the same time, the convolutional layer 9 is laterally connected to the DiFPN module 6 in the bottom-up path; The output of the C2f module 3 is used as the input of the convolutional layer 8, and the output of the convolutional layer 8 is used as the input of the DiFPN module 2. At the same time, the convolutional layer 8 is laterally connected to the DiFPN module 5 in the bottom-up path; The output of the C2f module 2 is used as the input of the convolutional layer 7, and the output of the convolutional layer 7 is used as the input of the DiFPN module 1. At the same time, the convolutional layer 7 is laterally connected to the DiFPN module 4 in the bottom-up path; The output of the C2f module 1 is used as the input of the convolution layer 6, and the convolution layer 6 is laterally connected to the DiFPN module 3 in the bottom-up path. At the same time, the output of the upsampling 1 is also used as the input of the DiFPN module 3; the output of the DiFPN module 3 is used as the input of the C2f module 7, and the output of the C2f module 7 is used as the input of the convolution layer 10, and the output of the convolution layer 10 is used as the input of the DiFPN module 4; The DiFPN module 1, DiFPN module 2, DiFPN module 3 and DiFPN module 6 have 2 inputs, and the DiFPN module with 2 inputs includes 2 DCNv4 modules and 1 weighted fusion module; DiFPN module 4 and DiFPN module 5 have 3 inputs, and the DiFPN module with 3 inputs includes 3 DCNv4 modules and 1 weighted fusion module; Step 3.3: The head network includes 4 detection heads; the output of C2f module 10 is used as the input of detection head 4; the output of C2f module 9 is used as the input of detection head 3; the output of C2f module 8 is used as the input of detection head 2; the output of C2f module 7 is used as the input of detection head 1; Step 4: Use the training set and the validation set to train the crowded scene pedestrian detection model based on the improved YOLOv8n, and save the trained model as the optimal model; the crowded scene pedestrian detection model training process based on the improved YOLOv8n further includes steps 4.1 to 4.4: Step 4.1: Setting the training parameters of the crowded scene pedestrian detection model based on improved YOLOv8n; Model training parameters include: number of iterations, batch size, optimizer, learning rate, momentum, weight decay, and number of threads; Step 4.2: Input the training set and the validation set and the corresponding labels into the crowded scene pedestrian detection model based on the improved YOLOv8n, and use the back propagation algorithm to calculate the gradient of the loss function to the model parameters; Step 4.3: Use the optimizer to update the model parameters in the direction of gradient descent until the loss function of the training set and the validation set no longer decreases, and the accuracy P, recall R, and mAP evaluation indicators no longer improve; Step 4.4: Save the trained model parameters as the optimal model; Step 5: Use the test set to test the optimal model, evaluate the test results of the test set by objective evaluation indicators, meet the accuracy requirements, and obtain the final crowded scene pedestrian detection model based on the improved YOLOv8n.
2. The method for pedestrian detection in crowded scenes based on improved YOLOv8n according to claim 1, characterized in that: The internal process of the DiFPN module with 2 inputs and the DiFPN module with 3 inputs further includes steps 3.2.1 to 3.2.7: Step 3.2.1: The DiFPN module with two inputs, the input feature maps are feature map X0 and feature map X1; the feature map X0 and feature map X1 are respectively input into the two DCNv4 modules in the DiFPN module with two inputs; The DiFPN module with three inputs, the input feature maps are feature map X0, feature map X1 and feature map X2 respectively; the feature map X0, feature map X1 and feature map X2 are respectively input into the three DCNv4 modules in the DiFPN module with three inputs; Step 3.2.2: The DCNv4 module uses 1 convolutional layer to predict the offset of each pixel position in the input feature map; Step 3.2.3: Dynamically adjust the position of the sampling point according to the offset, and for each convolution kernel position (i, j), calculate the new sampling position ( , ): (1) In formula (1), offset(i,j) is the predicted offset; Step 3.2.4: For the DiFPN module with 2 inputs, feature map X0, feature map X1 and offset are used as inputs of the deformable convolution layer. Finally, 2 deformed feature maps are generated. and ; For the DiFPN module with 3 inputs, feature map X0, feature map X1, feature map X2 and offset are used as the input of the deformable convolution layer. Finally, 3 deformed feature maps are generated. , and ; Step 3.2.5: For the DiFPN module with two inputs, the two deformed feature maps are and Input to the weighted fusion module, according to and The importance of assigning weights W0 and W1 to them will Multiplying by its weight W0, Multiply it by its weight W1, and add the weighted feature maps to obtain a weighted sum. The calculation formula of the weighted sum is: (2) In formula (2), O is the output feature after weighted fusion, W i is the weight, Output features for DCNv4, i=0, 1; For the DiFPN module with 3 inputs, the 3 deformed feature maps are , and Input to the weighted fusion module, according to , and The importance of assigning weights W0, W1 and W2 to them. Multiplying by its weight W0, Multiply it by its weight W1, Multiply it by its weight W2, and add the weighted feature maps to obtain a weighted sum. The calculation formula of the weighted sum is: (3) In formula (3), O is the output feature after weighted fusion, W i is the weight, Output features for DCNv4, i=0, 1, 2; Step 3.2.6: Introduce nonlinear features into the weighted sum through the Swish activation function; Step 3.2.7: For the DiFPN module with 2 inputs, use a 1×1 convolutional layer to operate the activated feature map to obtain an output feature map of the DiFPN module with 2 inputs; For the DiFPN module with 3 inputs, a 1×1 convolutional layer is used to operate the activated feature map to obtain an output feature map of the DiFPN module with 3 inputs.
3. The method for pedestrian detection in crowded scenes based on improved YOLOv8n according to claim 1, characterized in that: The feature fusion method for each horizontal span connection in step 3.2 is: For the i-th layer, the i-th layer feature map input to this layer is obtained according to the top-down and bottom-up feature propagation. and the i+1th layer feature map , the i+1th layer feature map Upsample and adjust to the same size as the feature map of the i-th layer , weightedly fuse and normalize the upsampled i+1th layer feature map and the i-th layer feature map to obtain the intermediate feature map of the i-th layer in the top-down path and the output feature map of the i-th layer in the bottom-up path; The feature fusion expression of the intermediate feature map of the i-th layer in the top-down path is: (4) In formula (4), ω1 represents the weight of the feature map of the i-th layer, ω2 represents the weight of the feature map of the i+1-th layer after upsampling, ε is a parameter used to prevent the denominator from being 0, and Conv represents feature fusion; The feature fusion expression of the output feature map of the i-th layer in the bottom-up path is: (5) In formula (5), represents the output feature map of the i-1th layer in the top-down path, It means that the output feature map of the i-1th layer is upsampled to the same size as the feature map of the i-th layer. represents the weight of the input i-th layer feature map, represents the weight of the intermediate feature map of the i-th layer in the bottom-up path, Represents the weight of the i-1th layer feature map after upsampling.
4. The method for pedestrian detection in crowded scenes based on improved YOLOv8n according to claim 1, characterized in that: In step 2, the second data set is divided into a training set, a validation set and a test set according to a preset ratio.
5. The method for pedestrian detection in crowded scenes based on improved YOLOv8n according to claim 1, characterized in that: The step 5 further includes steps 5.1 to 5.3: Step 5.1: Input the test set into the optimal model described in step 4; Step 5.2: Calculate the model performance indicators: accuracy P, recall R, average precision mAP. The specific calculation formula is as follows: (6) (7) (8) (9) In formula (6) to formula (9), P is the precision, R is the recall, mAP is the average precision of all categories, AP is the average precision, m is the total number of pedestrian label categories, TP is the number of positive samples correctly identified as positive samples, FP is the number of negative samples incorrectly identified as positive samples, and FN is the number of positive samples incorrectly identified as negative samples; Step 5.3: Evaluate the test results of the test set using objective evaluation indicators to meet the accuracy requirements, thereby obtaining the final crowded scene pedestrian detection model based on the improved YOLOv8n.
Citation Information
Patent Citations
YOLOv8 target detection method based on attention mechanism and multi-scale feature fusion
CN116883801A
Improved YOLOv8 dense pedestrian detection method based on GSConv + VOV-GSCSP
CN118015539A