A Dense Pedestrian Detection System Based on Improved Lightweight YOLOv7
Through the improved lightweight YOLOv7 detection model, combined with technical means such as PC-ELAN module, coordinate attention mechanism and long-term spatial attention mechanism, the problems of large amount of intensive pedestrian detection and large amount of parameters in the existing technology have been solved, and the pedestrian detection effect with high accuracy and real-time performance has been achieved.
Patent Information
- Application Number
- CN202310505340.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-05-08
AI Technical Summary
The existing deep learning object detection technology has a large amount of calculation and parameters in intensive pedestrian detection, making it difficult to deploy in embedded terminals, and the detection accuracy and real-time performance are difficult to meet the actual application needs.
The improved lightweight YOLOv7 detection model is adopted, and the model is lightweight and structural optimization is carried out through technical means such as PC-ELAN module, coordinate attention mechanism and long-term spatial attention mechanism. Combined with data enhancement and feature pyramid network, a system that can be efficiently detected in dense pedestrian scenarios is built.
It realizes pedestrian detection with high accuracy and strong real-time performance in dense pedestrian scenarios, which can meet pedestrian detection tasks in different application scenarios, and has high robustness and generalization performance.
Smart Images

Figure CN116612427B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning object detection, in particular to a dense pedestrian detection system based on an improved lightweight YOLOv7. Background Art
[0002] In various places in daily life, such as relatively crowded areas like large supermarkets, stations, traffic intersections, entertainment venues, and tourist attractions, surveillance devices such as cameras are required to detect people in real time, evaluate the density of people, and record the behavior information of pedestrians, so as to promptly disperse the crowd and take reasonable safety prevention measures.
[0003] With the development of deep learning technology, more and more deep learning technologies are applied to fields such as object detection. However, for the problem of dense pedestrian detection, it is still in the research stage. At the same time, deep learning models often have a large amount of computation and a large number of parameters, and are not easily deployed to embedded terminals. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a dense pedestrian detection system based on an improved lightweight YOLOv7, which applies the improved lightweight YOLOv7 detection model to the dense pedestrian detection task, and at the same time satisfies pedestrian detection in both dense pedestrian scenarios and sparse pedestrian scenarios, meets the real-time requirements of practical applications, has a fast processing speed, and high detection accuracy.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A dense pedestrian detection system based on an improved lightweight YOLOv7, including the following steps:
[0006] Step S1: Collect pedestrian images in different scenarios, including images with a high density of pedestrians;
[0007] Step S2: Perform information annotation on the pedestrian images collected in Step S1, and label the pedestrian position information, image pedestrian density information, and image pedestrian scale information according to the position of the pedestrian in the image and the pixel size occupied by the pedestrian in the image;
[0008] Step S3: Perform data augmentation on the pedestrian images collected in Step S1;
[0009] Step S4: Construct an improved YOLOv7 dense pedestrian detection model, and at the same time achieve pedestrian detection in dense pedestrian scenarios and non-dense pedestrian scenarios;
[0010] Step S5: Perform lightweight processing on the improved YOLOv7 dense pedestrian detection model in Step S4;
[0011] Step S6: Deploy the model after lightweight processing in step S5 to a terminal device and build a dense pedestrian detection system;
[0012] Step S7: Use the dense pedestrian detection system built in step S6 to perform real-time detection on the pedestrians in the preprocessed pedestrian images in step S3 and output the detection results.
[0013] In a preferred embodiment, the acquisition methods of the pedestrian images in step S1 include collecting pedestrian image information through surveillance cameras in public places, taking aerial photos of public places through drone cameras, collecting pedestrian image information through on-vehicle cameras of public transportation vehicles, or collecting pedestrian image information through researchers using cameras in public places.
[0014] In a preferred embodiment, the annotation method in step S2 is as follows: Use DarkLabel software to annotate the collected pedestrian images, and respectively annotate the position information of the pedestrians in the images, the density information of the pedestrians in the images, and the scale information of the pedestrians in the images, that is, the pixel size occupied by the pedestrians in the images. Pedestrian targets with a ratio of the pixels occupied by the pedestrians to the image pixels less than 1% are regarded as small-scale targets; According to the pedestrian density information of the images and the pedestrian scale information of the images, classify the pedestrian images, and classify the images with a larger pedestrian density or smaller pedestrian scale into the difficult-to-classify first level.
[0015] In a preferred embodiment, the data augmentation operations in step S3 are data augmentation in the following ways:
[0016] N1: Perform rigid transformation on the images: Use random cropping, random rotation, random horizontal flipping, Mosaic augmentation, Cutout, Mixup, CutMix on the pedestrian images;
[0017] N2: Perform transformation on the pixel values of the images: Use random brightness transformation, random saturation transformation, add random environmental noise, random sharpening, random blurring, random grayscaling on the pedestrian images;
[0018] N3: Perform data augmentation on the difficult-to-classify samples in the images: Use PuzzleMix, oversampling, Copy-Paste, instance-level imbalance on the pedestrian images;
[0019] Use the above data augmentation as the training dataset of the detection model to enhance the generalization performance and robustness of the detection model.
[0020] In a preferred embodiment, the improvement method of the YOLOv7 model in step S4 is as follows:
[0021] U1: In the backbone feature extraction network of YOLOv7, the PC-ELAN module is used to replace the original ELAN structure. PC convolution and PW convolution are used as convolution operators, and lightweight neural network modules are utilized to improve the detection speed and ensure real-time detection.
[0022] U2: Without significantly increasing the network parameters, the PC-ELAN module is integrated into the Coordinate Attention mechanism and the Decoupled Fully Connected long-term spatial attention mechanism to improve the feature extraction ability of the network model for feature maps.
[0023] U3: The downsampling method of the backbone adopts S2D downsampling to achieve lightweight network design.
[0024] U4: A Global Context Block is added to the shallow layer of the backbone to capture global dependencies.
[0025] U5: The feature extraction layer uses the Recusive FPN structure and four groups of weighted bidirectional feature pyramid network structures (BiFPN) to replace the original PAN structure, enabling the detection model to pay more attention to important levels and achieving a fast and efficient multi-scale fusion method.
[0026] U6: In the Head part of the YOLOv7 detection model, three Heads are used to predict the head, visible area, and full body area of pedestrians respectively, improving the detection accuracy of the model and conforming to the visual structure of the human eye.
[0027] In a preferred embodiment, the model lightweight processing method in step S5 is as follows: Through structural re-parameterization, the convolutional layer and BN layer in the detection model described in step S4 are merged to reduce the inference time of the detection model; then, the model is further compressed through pruning, quantization, and distillation methods, enabling the detection model to be deployed on lightweight terminals.
[0028] In a preferred embodiment, step S6 specifically includes:
[0029] Step S61: Use Qt to build the upper computer of the detection system to achieve the pedestrian detection function.
[0030] Step S62: Deploy the model after lightweight processing in step S5 to the NVIDIA Jetson Xavier NX 8G terminal through TensorRT.
[0031] Step S63: Use a Sony HD-X12MP-AF camera to collect images.
[0032] Compared with the prior art, the present invention has the following beneficial effects: The improved lightweight YOLOv7 object detection model of the present invention is applied to the pedestrian detection task, which can still ensure high accuracy and real-time performance in a dense crowd, and at the same time meet the pedestrian detection tasks in different application scenarios. It has certain application and research value for the dense pedestrian detection technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is the PC-ELAN structure diagram of the preferred embodiment of the present invention;
[0034] Figure 2 It is the structure diagram of the Attention-FasterNet Block structure of the preferred embodiment of the present invention;
[0035] Figure 3 It is the structure diagram of the coordinate attention mechanism of the preferred embodiment of the present invention;
[0036] Figure 4 It is the structure diagram of the long-term spatial attention mechanism of the preferred embodiment of the present invention;
[0037] Figure 5 It is the S2D downsampling structure diagram of the preferred embodiment of the present invention;
[0038] Figure 6 It is the structure diagram of the Global Context Block of the preferred embodiment of the present invention;
[0039] Figure 7 It is the Neck structure diagram of the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0040] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0041] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which the present application belongs.
[0042] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0043] The present invention provides as Figures 1 - 7A dense pedestrian detection system based on an improved YOLOv7 network is shown. The system design method includes the following steps:
[0044] Step 1: Researchers collect pedestrian images in the following ways:
[0045] 1. Collect pedestrian image information through surveillance cameras in public places;
[0046] 2. Take aerial photos of public places through a drone camera to collect pedestrian image information;
[0047] 3. Collect pedestrian image information through in-vehicle cameras of public transportation vehicles;
[0048] 4. Researchers use a camera in public places to collect pedestrian image information;
[0049] The collected pedestrian images should meet the following requirements:
[0050] 1. The pedestrian forms are diverse, that is, there are human instances with different postures, such as walking postures, standing postures, sitting postures, etc., to solve the problem of handling intra-class variations of pedestrians in the dense pedestrian detection system;
[0051] 2. The pedestrian clothing is diverse, that is, pedestrians wear different styles and colors of clothing to solve the problem of handling intra-class variations of pedestrians in the dense pedestrian detection system;
[0052] 3. The scenes where pedestrians are located are diverse. To meet the application of dense pedestrian detection in different scenes, the scenes where pedestrians are located collected by researchers should include complex scenes such as traffic intersections with large traffic flow, stations, internal scenes of public transportation vehicles, supermarkets, shopping malls, entertainment venues, and tourist attractions, and at the same time have different weather conditions. The robustness and generalization performance of the dense pedestrian detection system are improved through diverse scene images;
[0053] 4. The crowd density of pedestrians is evenly distributed, that is, there are dense pedestrian images and sparse pedestrian images, and there should be various occlusion relationships between pedestrians to ensure the detection performance of the dense pedestrian detection system at different crowd densities and to handle the occlusion problem of dense crowds;
[0054] 5. The scale sizes of pedestrians are evenly distributed, that is, there are large-scale pedestrian targets and small-scale pedestrian targets to ensure the detection performance of the dense pedestrian detection system for small-scale targets and multi-scale targets.
[0055] Step 2: Use DarkLabel software to annotate the location information of the pedestrian images collected in Step 1 in the coco format. The main annotation information includes: the location information of pedestrians in the image, the category information of pedestrians (real people, dummies, people in the mirror, etc.), the ignored areas in the image, the density information of pedestrians in the image, and the scale information of pedestrians in the image. Specifically, use DarkLabel software to annotate the collected pedestrian images, and separately annotate the location information of pedestrians in the image, the density information of pedestrians in the image, and the scale information of pedestrians in the image (that is, the pixel size occupied by pedestrians in the image. Pedestrian targets with a ratio of the pixels occupied by pedestrians to the image pixels less than 1% are regarded as small-scale targets). According to the pedestrian density information of the image and the pedestrian scale information of the image, classify the pedestrian images. Images with a higher pedestrian density or smaller pedestrian scale are classified as the difficult-to-classify first level.
[0056] Step 3: Perform data augmentation operations on the collected pedestrian images. The augmentation methods are as follows:
[0057] 1. Perform rigid transformation on the image: Use random cropping, random rotation, random horizontal flipping, Mosaic augmentation, Cutout, Mixup, CutMix on the pedestrian images;
[0058] 2. Perform transformation on the pixel values of the image: Use random brightness transformation, random saturation transformation, add random environmental noise, random sharpening, random blurring, random grayscaling on the pedestrian images;
[0059] 3. Perform data augmentation on difficult-to-detect samples in the image: Use PuzzleMix (to make the dense pedestrian detection model focus on the salient regions), oversampling (to make the dense pedestrian detection model focus on difficult-to-classify examples), Copy-Paste (to copy pedestrian instances to different positions in the image to increase the density and occlusion degree of the pedestrian image), instance-level imbalance (to increase the proportion of small-scale pedestrians in the image to enhance the detection ability of the dense pedestrian detection system for small targets);
[0060] Use the above data augmentation as the training dataset for the dense pedestrian detection system to enhance the generalization performance and robustness of the dense pedestrian detection model, so that the trained dense pedestrian detection model can be applied to complex scenarios and still maintain a high accuracy in dense pedestrian scenarios.
[0061] Step 4: Select the YOLOv7 detection model as the basic framework of the pedestrian detection model, use Pytorch to build the YOLOv7 detection model, and optimize the network model on the basis of the original model. Perform targeted structural adjustments according to the detection task:
[0062] Replace the ELAN structure of YOLOv7 with the PC-ELAN structure. The PC-ELAN is shown as Figure 1 and at the same time, integrate the Coordinate Attention mechanism and the Long-Term Spatial Attention mechanism (DFC) into each Attention-FasterNet Block. The structure of the Attention-FasterNet Block is shown as Figure 2 . Without significantly increasing the network parameters, it improves the feature extraction ability of the backbone network for feature maps. Among them, the Coordinate Attention mechanism, as shown in Figure 3 , can not only implement the coordinate attention mechanism but also the channel attention mechanism; the Long-Term Spatial Attention mechanism, as shown in Figure 4 , implements an efficient self-attention mechanism, which can capture long-range spatial information and enhance the feature extraction ability of the backbone network.
[0063] Furthermore, the downsampling method of the backbone adopts S2D downsampling, as shown in Figure 5 . S2D is a parameter-free downsampling method that realizes the downsampling function by converting the spatial size of the feature map into the depth of the feature map, avoiding the information loss caused by traditional convolutional downsampling or pooling downsampling, and at the same time can preserve more feature map information.
[0064] Furthermore, add a Global Context Block to the shallow layer of the backbone. The structure is shown as Figure 6 . It realizes a lightweight Non-local network, avoiding the problems of large computational complexity and inapplicability to lightweight models of traditional Non-local, and at the same time has the function of capturing long-range dependence relationships.
[0065] Furthermore, use the Recusive FPN structure (recursive FPN) and four groups of BiFPN structures (weighted bidirectional feature pyramid network) to replace the original PAN structure. The structure is shown as Figure 7 . It makes the detection model pay more attention to important levels and realizes a fast and efficient multi-scale fusion method. Among them, Recusive FPN circulates with the downsampling and the convolutional neural network of the backbone to improve the utilization rate of the backbone; BiFPN uses downsampling and upsampling for information fusion, fusing the small object information, large object information, detail information, and high-level semantic information extracted by the backbone to enhance the network's perception ability.
[0066] Step 5: Through structural reparameterization of the constructed improved YOLOv7 detection model, the convolutional layer and BN layer in the backbone network of the model are merged. During model training, it appears as a convolutional layer and a BN layer, and during model inference, it appears as a convolutional layer, reducing the number of model parameters and computational complexity. Further compress the constructed YOLOX detection model through pruning and quantization so that the detection model can be deployed on lightweight terminals.
[0067] Step 6: Deploy the lightweight processed improved lightweight YOLOv7 detection model to NVIDIA Jetson Xavier NX 8G through TensorRT. This terminal is designed specifically for AI and has more powerful performance compared to embedded devices such as Raspberry Pi and single-chip microcontrollers, supporting all popular AI frameworks. This brings new possibilities to embedded edge computing devices that need to improve performance to support AI workloads while being limited by size, weight, power consumption, or cost. Jetson Xavier NX can provide 14 TOPS at 10 watts of power and 21 TOPS at 15 watts of power, making it very suitable for systems limited in terms of size and power. With 384 CUDA cores, 48 Tensor Cores, and 2 NVDLA engines, it can run multiple modern neural network models in parallel and simultaneously process high-resolution data from multiple sensors.
[0068] In actual applications, collect pedestrian images through a Sony HD-X12MP-AF camera, process the data of the pedestrian images collected by the camera through the Jetson Xavier NX terminal, and use QT to build the upper computer of the detection system to display the pedestrian detection results in real time. Specific functions of the dense pedestrian detection system:
[0069] Automatically save the image information recorded by the camera in real time;
[0070] Display and save the position information of pedestrians in the image information captured by the current camera in real time;
[0071] Display and save the total number of pedestrians and crowd density in the image information captured by the current camera in real time;
[0072] When the dense pedestrian detection system detects that the crowd density reaches the safety threshold, issue a safety warning;
[0073] The recognition area of the image can be specified, and pedestrians in the corresponding area can be detected as needed; at the same time, the ignored area of the image can be specified to meet different pedestrian detection requirements.
Claims
1. A dense pedestrian detection system based on the improved lightweight YOLOv7, characterized in that: It includes the following steps: Step S1: Collect pedestrian images in different scenarios, including images with a high density of pedestrians; Step S2: Perform information annotation on the pedestrian images collected in Step S1. According to the position of pedestrians in the image and the pixel size occupied by the image, annotate the pedestrian position information, image pedestrian density information, and image pedestrian scale information; Step S3: Perform data augmentation on the pedestrian images collected in Step S1; Step S4: Construct an improved YOLOv7 dense pedestrian detection model to achieve pedestrian detection in both dense pedestrian scenarios and non-dense pedestrian scenarios; Step S5: Perform lightweight processing on the improved YOLOv7 dense pedestrian detection model in Step S4; Step S6: Deploy the model after lightweight processing in Step S5 to a terminal device and construct a dense pedestrian detection system; Step S7: Pass the preprocessed pedestrian images in Step S3 through the dense pedestrian detection system constructed in Step S6 to perform real-time detection of pedestrians in the image and output the detection results; The improvement method of the YOLOv7 model in Step S4 is as follows: U1: The backbone feature extraction network backbone of YOLOv7 uses the PC-ELAN module to replace the original ELAN structure, uses PC convolution and PW convolution as convolution operators, and uses a lightweight neural network module to improve the detection speed and ensure detection real-time performance; U2: While not significantly increasing the network parameters, integrate the PC-ELAN module into the Coordinate Attention mechanism and the Decoupled Fully Connected long-term spatial attention mechanism to improve the feature extraction ability of the network model for the feature map; U3: The downsampling method of the backbone uses S2D downsampling to achieve lightweight network design; U4: Add a Global Context Block to the shallow layer of the backbone to capture global dependencies; U5: The feature extraction layer uses a Recusive FPN structure and a four-group weighted bidirectional feature pyramid network structure BiFPN to replace the original PAN structure, making the detection model pay more attention to important levels and achieving a fast and efficient multi-scale fusion method; U6: In the Head part of the YOLOv7 detection model, use three Heads to predict the head, visible area, and whole body area of pedestrians respectively, improving the detection accuracy of the model and conforming to the visual structure of the human eye.
2. A dense pedestrian detection system based on the improved lightweight YOLOv7 according to claim 1, characterized in that, The acquisition methods of the pedestrian images in Step S1 include collecting pedestrian image information through surveillance cameras in public places, collecting pedestrian image information by shooting public places from above with a drone camera, collecting pedestrian image information through in-vehicle cameras of public transportation vehicles, or collecting pedestrian image information by researchers using cameras in public places.
3. A dense pedestrian detection system based on the improved lightweight YOLOv7 according to claim 1, characterized in that, the annotation method in step S2 is as follows: annotate the collected pedestrian images through DarkLabel software, and respectively annotate the position information of pedestrians in the images, the density information of pedestrians in the images, and the scale information of pedestrians in the images, that is, the pixel size occupied by pedestrians in the images. Pedestrian targets with a ratio of the pixels occupied by pedestrians to the image pixels less than 1% are regarded as small-scale targets; according to the pedestrian density information of the images and the pedestrian scale information of the images, classify the pedestrian images, and classify the images with a higher pedestrian density or images with a smaller pedestrian scale into the difficult-to-classify first level.
4. A dense pedestrian detection system based on the improved lightweight YOLOv7 according to claim 1, characterized in that, the data augmentation operation in step S3 is to perform data augmentation in the following ways: N1: Perform rigid transformation on the images: use random cropping, random rotation, random horizontal flipping, Mosaic augmentation, Cutout, Mixup, CutMix on the pedestrian images; N2: Transform the pixel values of the images: use random brightness transformation, random saturation transformation, add random environmental noise, random sharpening, random blurring, random grayscaling on the pedestrian images; N3: Perform data augmentation on the difficult-to-classify samples in the images: use PuzzleMix, oversampling, Copy-Paste, instance-level imbalance on the pedestrian images.
5. A dense pedestrian detection system based on the improved lightweight YOLOv7 according to claim 1, characterized in that, the model lightweight processing method in step S5 is as follows: through structural reparameterization, merge the convolutional layer and the BN layer in the detection model described in step S4 to reduce the inference time of the detection model; then further compress the model through pruning, quantization, and distillation methods to deploy the detection model on lightweight terminals.
6. A dense pedestrian detection system based on the improved lightweight YOLOv7 according to claim 1, characterized in that, the specific steps of step S6 are as follows: Step S61: Use Qt to build the upper computer of the detection system to implement the pedestrian detection function; Step S62: Deploy the model after lightweight processing in step S5 to the NVIDIA Jetson Xavier NX 8G terminal through TensorRT; Step S63: Use the Sony HD-X12MP-AF camera to collect images.
Citation Information
Patent Citations
Improved YOLOv5 lightweight community scene pedestrian detection method
CN115862066A
Kitchen behavior real-time monitoring method and system based on improved target detection model
CN116012789A