Lightweight pedestrian detection algorithm based on improved YOLOv7-tiny

By improving the YOLOv7-tiny network structure and introducing the ECA attention mechanism, the missed detection and missed detection problems in complex environments in pedestrian detection are solved, and the speed and accuracy of detection are improved.

CN119942586APending Publication Date: 2025-05-06HUBEI UNIV OF ECONOMICS

Patent Information

Application Number
CN202411862033.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing pedestrian detection methods cannot meet the real-time operation of the algorithm in scenarios with limited computing power, and it is difficult to maintain the accuracy of detection in complex environments.

Method used

The lightweight pedestrian detection algorithm based on improved YOLOv7-tiny is adopted, and the ELAN module that replaces the YOLOv7-tiny network structure is used to perform feature extraction, and the ECA attention mechanism is used to perform feature enhancement in the feature fusion subnet.

Benefits of technology

The detection accuracy of pedestrians in complex environments is improved, the parameters and calculation amount of the model are reduced, and the real-time and accuracy of the detection are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942586A_ABST
    Figure CN119942586A_ABST
Patent Text Reader

Abstract

The invention relates to a lightweight pedestrian detection algorithm based on improved YOLOv7-tiny, and the algorithm comprises the steps: replacing an ELAN module with an ELCG module on the basis of an original YOLOv7-tiny network to carry out the feature extraction of an input image, and carrying out the feature enhancement in a feature fusion sub-network through employing an ECA attention mechanism; comprising the following steps: obtaining a pedestrian detection picture, and shortening the picture processing time and constructing a data set D1 by using a data enhancement and adaptive size scaling method; an improved YOLOv7-tiny model is constructed; the data set D1 is input into the improved YOLOv7-tiny model to be trained; and pedestrian detection is carried out on the trained improved YOLOv7-tiny model, loss function calculation is carried out on a detection result and a sample label, and parameters in the model are optimized through a gradient descent algorithm and a back propagation algorithm. According to the lightweight pedestrian detection algorithm based on the improved YOLOv7-tiny, an ECA attention mechanism is introduced, and the network structure is improved, so that the detection precision of pedestrians in a complex environment is improved, and the accuracy of capturing and identifying features in pedestrian detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a lightweight pedestrian detection algorithm based on improved YOLOv7-tiny. Background Art

[0002] Pedestrian detection is to find the position and size of pedestrians in images, which has important application value in the fields of intelligent driving, security monitoring, etc. However, the existing detection methods are restricted by hardware conditions in actual application scenarios. In scenarios with limited computing power, they cannot meet the real-time operation of the algorithm and maintain the accuracy of detection.

[0003] In the existing technology, the ShuffleNetV2 lightweight network is used as the backbone feature extraction network of the model. Although it can improve the detection speed, it reduces the accuracy of target detection. The Dense network framework is introduced to improve the sharpness of the input original image during training and strengthen the feature extraction ability of the model, which can improve the accuracy of detection, but it is difficult to maintain the real-time performance of detection. Summary of the invention

[0004] The technical problem of the present invention is: the present invention proposes a lightweight pedestrian detection algorithm based on improved YOLOv7-tiny, which solves the problems of missed detection and false detection faced in pedestrian detection in an environment with dense pedestrians and mutual occlusion, while accelerating the detection speed of pedestrians and improving the accuracy of detection.

[0005] The technical solution of the present invention is a lightweight pedestrian detection algorithm based on improved YOLOv7-tiny: The ELCG module replaces the ELAN module of the YOLOv7-tiny network structure to extract features from the input image, and the ECA attention mechanism is used in the feature fusion subnetwork to enhance features; The following steps are involved: S1: Obtain pedestrian detection images and use data enhancement and adaptive size scaling methods to shorten image processing time and build dataset D1; S2: Build an improved YOLOv7-tiny model; S3: Input the dataset D1 into the improved YOLOv7-tiny model for training; S4: Use the trained improved YOLOv7-tiny model to detect pedestrians, and calculate the loss function between the detection results and the sample labels, and then optimize the parameters in the model through the gradient descent algorithm and back propagation algorithm.

[0006] Preferably, step S1 includes the following sub-steps: 1) Adjust the brightness and contrast of the input image, and improve the image enhancement effect by color jittering and adding noise. Use the Mosaic data enhancement method to randomly scale the image size, randomly crop and splice the image, and set the size to generate the image for training the model; 2) During the adaptive resizing process, the image is scaled proportionally to the set size, and the minimum black border width to be filled is calculated based on the downsampling multiple.

[0007] Preferably, in step S2, the YOLOv7-tiny model is improved, including a backbone feature subnetwork, a multi-scale feature fusion network and a decoupled detection head; the backbone feature subnetwork extracts features from the input image to obtain a primary feature layer; the multi-scale feature fusion subnetwork extracts higher-level features from top to bottom, and then fuses features of different levels from bottom to top to output fused features; the decoupled detection head passes the fused features into the anchor-free frame to output pedestrian detection results; Furthermore, the backbone feature sub-network includes a CBS module, an MP module and an ELCG module connected in sequence. The CBS module includes ordinary convolution Conv, a BN layer and a SiLu activation function. The ELCG module includes Ghost-Conv convolution and a CBS module. The MP module includes a maximum pooling layer and a CBS module.

[0008] Furthermore, the feature map is divided into two branches through the ELCG module. The first branch extracts features through the Ghost-Conv convolution layer, and the second branch extracts features through the Ghost-Conv convolution layer and two CBS convolution modules. The two branches are fused through the Concat layer, and then the Ghost-Conv convolution layer is used to extract features.

[0009] Preferably, the multi-scale feature fusion network includes a CBS module, an ELCG module, an SPPCSPC module, an upsampling module and an ECA attention module, and the SPPCSPC module includes multiple pooling layers for reducing the size of the feature map and reducing the number of channels.

[0010] Preferably, the ECA attention mechanism generates a global channel dependency weight by calculating the one-dimensional local autocorrelation function of each channel, and inputs the weight into each channel of the feature map, including the following sub-steps: 1) Perform global average pooling on the input feature map to generate a feature vector. The calculation formula is:

[0011] In the formula, Represents the result of global average pooling of feature maps, Represents the input of the feature map, C represents the number of channels of the feature map, i andj Indicates the counting unit, W and H indicate i and j The maximum value of 2) Perform convolution operation on the feature vector to capture the local dependencies between channels. The calculation formula is:

[0012] In the formula, k represents the channel number mapping convolution kernel size, γ and b are hyperparameters, and odd represents an odd function; 3) Normalize through Sigmoid activation function, and weight the obtained channel weight with the original feature map to obtain the enhanced feature map.

[0013] Preferably, in step S4, the three high-quality fusion feature layers are passed to the decoupled detection head without anchor frames to obtain the detection results, and the loss function is calculated with the assigned sample labels. The loss function is the coordinate loss L CIoU , target confidence loss L dfl and classification loss L cls , the calculation formula is: ; In the formula, K represents the number of output results of different sizes, α balance represents the weight coefficient of each size result, S represents the length of the grid of the image, B represents the category, α box Denotes the loss function L CIoU The weight coefficient, α dfl Denotes the loss function L dfl The weight coefficient, α cls Denotes the loss function L cls The weight coefficient, i and j represent the calculation units; Loss function L CIoU The calculation formula is:

[0014] Where, d o Indicates the distance between the center points of the two boxes, d c Indicates the distance between the diagonals of the two boxes, w, h, w gt and h gt Represent the width and height of the predicted box and the width and height of the real box respectively, b represents the detection box, and v represents the measurement unit, which is used to measure the similarity of aspect ratio; Loss function L dfl The calculation formula is:

[0015] In the formula, y and y i+1Represents two adjacent discrete values ​​of the predicted bounding box position, y represents the position of the real bounding box, S i and S i+1 Indicates that it corresponds to y i and i+1 The probability of Loss function L cls The calculation formula is:

[0016] In the formula, c p and c gt They represent the probability of the predicted category and the probability of the true category respectively.

[0017] Compared with the prior art, the beneficial effects of the present invention include: 1) This paper improves the lightweight pedestrian detection algorithm based on YOLOv7-tiny and introduces the ECA attention mechanism. By improving the network structure, the detection accuracy of pedestrians in complex environments is improved, and the accuracy of capturing and identifying features in pedestrian detection is improved.

[0018] 2) The present invention is based on an improved YOLOv7-tiny lightweight pedestrian detection algorithm. By introducing a lightweight convolution module ELCG and utilizing the redundancy between feature maps, more feature maps are generated at a lower cost, the number of required convolution kernels is reduced, the number of model parameters and the amount of calculation are reduced, and the practicability of the pedestrian detection algorithm is improved.

[0019] 3) The present invention is based on an improved YOLOv7-tiny lightweight pedestrian detection algorithm, adopts an Anchor-Free based decoupled detection head, does not rely on preset anchor frames, avoids the computational overhead required to generate a large number of anchor frames, and improves the efficiency of pedestrian detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0021] Figure 1 A schematic diagram of the training and detection process of improving the YOLOv7-tiny network in an embodiment of the present invention; Figure 2 A schematic diagram of the structure of an improved YOLOv7-tiny network according to an embodiment of the present invention; Figure 3 A schematic diagram of the ELCG module structure of an improved YOLOv7-tiny network according to an embodiment of the present invention; Figure 4 A schematic diagram of the structure of the Ghost-Conv module of the improved YOLOv7-tiny network according to an embodiment of the present invention; Figure 5The figure is a schematic diagram of the ECA module structure of the improved YOLOv7-tiny network according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] The experimental platform of the present invention adopts the Ubuntu 20.04 operating system, the processor is Intel (R) i7-10700K CPU @ 3.80GHz, the graphics card is NVIDIA GTX 3070 (8GB), and the Pytorch deep learning framework is adopted. The dependent environment is: CUDA11.8, Python3.8, Pytorch2.1.0. Two mainstream pedestrian detection public datasets, CrowdHuman and WiderPerson, are used. The CrowdHuman dataset occupies an important position in the field of pedestrian detection due to its unique crowding and occlusion characteristics. The dataset contains a total of 15,000 training images, 4,370 verification images, and 5,000 test images, and each image contains an average of 22.6 humans. The WiderPerson dataset is known for its high crowding and rich scene changes. The dataset contains 236,073 human instances, of which 8,000 images are used for training, and each image contains an average of 29.51 instances.

[0023] like Figure 1 As shown in the figure, a lightweight pedestrian detection algorithm based on improved YOLOv7-tiny: The ELCG module replaces the ELAN module of the YOLOv7-tiny network structure to extract features from the input image, and the ECA attention mechanism is used in the feature fusion subnetwork to enhance features; The following steps are involved: S1: Obtain pedestrian detection images and use data enhancement and adaptive size scaling methods to shorten image processing time and build dataset D1; Step S1 includes the following sub-steps: 1) Adjust the brightness and contrast of the input image, and improve the image enhancement effect by color jittering and adding noise. Use the Mosaic data enhancement method to randomly scale the image size, randomly crop and splice the image, and set the size to generate the image for training the model; 2) During the adaptive resizing process, the image is scaled proportionally to the set size, and the minimum black border width to be filled is calculated based on the downsampling multiple.

[0024] S2: Build an improved YOLOv7-tiny model; like Figure 2As shown, in step S2, the improved YOLOv7-tiny model includes a backbone feature subnetwork, a multi-scale feature fusion network and a decoupled detection head; the backbone feature subnetwork extracts features from the input image to obtain a primary feature layer; the multi-scale feature fusion subnetwork extracts higher-level features from top to bottom, and then fuses features of different levels from bottom to top to output fused features; the decoupled detection head passes the fused features into the anchor-free frame to output pedestrian detection results; The backbone feature sub-network includes the CBS module, MP module and ELCG module connected in sequence. The CBS module includes ordinary convolution Conv, BN layer and SiLu activation function, the ELCG module includes Ghost-Conv convolution and CBS module, and the MP module includes the maximum pooling layer and CBS module.

[0025] like Figure 3 and Figure 4 As shown in the figure, the feature map of the ELCG module is divided into two branches through the ELCG module. The first branch extracts features through the Ghost-Conv convolution layer, and the second branch extracts features through the Ghost-Conv convolution layer and two CBS convolution modules. The features of the two branches are fused through the Concat layer, and then the features are extracted using the Ghost-Conv convolution layer.

[0026] The multi-scale feature fusion network includes CBS module, ELCG module, SPPCSPC module, upsampling module and ECA attention. The SPPCSPC module contains multiple pooling layers to reduce the size of the feature map and the number of channels.

[0027] like Figure 5 As shown in the figure, the ECA attention mechanism generates a global channel dependency weight by calculating the one-dimensional local autocorrelation function of each channel and inputs the weight into each channel of the feature map, including the following sub-steps: 1) Perform global average pooling on the input feature map to generate a feature vector. The calculation formula is:

[0028] In the formula, Represents the result of global average pooling of feature maps, Represents the input of the feature map, C represents the number of channels of the feature map, i and j Indicates the counting unit, W and H indicate i and j The maximum value of 2) Perform convolution operation on the feature vector to capture the local dependencies between channels. The calculation formula is:

[0029] In the formula, k represents the channel number mapping convolution kernel size, γ and b are hyperparameters, and odd represents an odd function; 3) Normalize through Sigmoid activation function, and weight the obtained channel weight with the original feature map to obtain the enhanced feature map.

[0030] S3: Input the dataset D1 into the improved YOLOv7-tiny model for training; The dataset D1 is divided into training set, test set and validation set according to the ratio, and the validation set is input into the improved YOLOv7-tiny model for training.

[0031] S4: Use the trained improved YOLOv7-tiny model to detect pedestrians, and calculate the loss function between the detection results and the sample labels, and then optimize the parameters in the model through the gradient descent algorithm and back propagation algorithm.

[0032] Preferably, in step S4, the three high-quality fusion feature layers are passed to the decoupled detection head without anchor frames to obtain the detection results, and the loss function is calculated with the assigned sample labels. The loss function is the coordinate loss L CIoU , target confidence loss L dfl and classification loss L cls , the calculation formula is:

[0033] In the formula, K represents the number of output results of different sizes, α balance represents the weight coefficient of each size result, S represents the length of the grid of the image, B represents the category, α box Denotes the loss function L CIoU The weight coefficient, α dfl Denotes the loss function L dfl The weight coefficient, α cls Denotes the loss function L cls The weight coefficient of i and j Indicates the unit of calculation; Loss function L CIoU The calculation formula is:

[0034] Where, d o Indicates the distance between the center points of the two boxes, d c Indicates the distance between the diagonals of the two boxes, w, h, w gt and h gt Represent the width and height of the predicted box and the width and height of the real box respectively. b represents the detection box, v Represents the unit of measurement used to measure the similarity of aspect ratio; Loss function L dfl The calculation formula is:

[0035] In the formula, y and i+1 Represents two adjacent discrete values ​​of the predicted bounding box position, y represents the position of the real bounding box, S i and S i+1 Indicates that it corresponds to y i and i+1 The probability of Loss function L cls The calculation formula is:

[0036] In the formula, c p and c gt They represent the probability of the predicted category and the probability of the true category respectively.

[0037] When performing detection, the detection results use the non-maximum suppression method NMS to filter out duplicate detection frames, delete the detection frames whose IOU with the detection frame with the highest confidence is greater than the set threshold, and finally generate the detection frame, category and confidence of the target.

[0038] In order to verify the effectiveness of this application in pedestrian detection tasks, the present invention and other mainstream YOLO lightweight algorithms are trained and tested on the CrowdHuman and WiderPerson datasets, and the performance of each algorithm is compared. The optimization effect of this application is demonstrated by comparing the accuracy, recall, average precision MAP50 and MAP95, and FPS. The experimental results on the CrowdHuman and WiderPerson datasets are shown in Tables 1 and 2 respectively: Table 1

[0039] Table 2

[0040] On the CrowdHuman and WiderPerson datasets, this application achieved the highest mAP95 compared with other mainstream YOLO lightweight algorithms. Compared with the original YOLOv7-tiny algorithm, the number of parameters in this application was reduced by 14.47%, and the mAP95 on the CrowdHuman dataset was increased by 2.27%; on the WiderPerson dataset, the mAP95 was increased by 2.54%. This shows that the present invention has shown good performance in pedestrian detection tasks, and can achieve high detection accuracy for different datasets, indicating that this application can achieve more accurate target detection.

[0041] In order to verify the effectiveness of the improved method proposed in the invention, ablation experiments were conducted on the CrowdHuman and WiderPerson datasets, and the experimental results are shown in Tables 3 and 4. ECA represents the addition of the attention mechanism, ELCG represents the use of the lightweight convolution module ELCG based on GhostConv, and AFDH represents the use of a decoupled detection head without an anchor frame.

[0042] Table 3

[0043] Table 4

[0044] Experimental results show that the ECA attention mechanism can effectively improve the accuracy of the algorithm, the ELCG module can effectively reduce the amount of parameters while the accuracy of the model is slightly increased, the ADHead detection head can significantly improve the accuracy of the algorithm but the amount of parameters increases significantly. When the ELCG module and the ADHead detection head are used at the same time, the model accuracy is further improved, and the amount of parameters is significantly smaller than using the ADHead detection head alone. When the three improvements are used at the same time, the present invention achieves the highest accuracy, which proves the effectiveness of the improved method of the present invention.

[0045] The above embodiments are only preferred technical solutions of the present invention and should not be regarded as limitations of the present invention. The protection scope of the present invention shall be the technical solutions recorded in the claims, including equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, equivalent replacement improvements within this scope are also within the protection scope of the present invention.

Claims

1. A lightweight pedestrian detection algorithm based on improved YOLOv7-tiny, characterized in that: Based on the original YOLOv7-tiny network, the ELCG module replaces the ELAN module to extract features of the input image, and the ECA attention mechanism is used in the feature fusion subnetwork for feature enhancement; The following steps are involved: S1: Obtain pedestrian detection images and use data enhancement and adaptive size scaling methods to shorten image processing time and build dataset D1; S2: Build an improved YOLOv7-tiny model; S3: Input the dataset D1 into the improved YOLOv7-tiny model for training; S4: Use the trained improved YOLOv7-tiny model to detect pedestrians, and calculate the loss function between the detection results and the sample labels, and then optimize the parameters in the model through the gradient descent algorithm and back propagation algorithm.

2. According to claim 1, a lightweight pedestrian detection algorithm based on improved YOLOv7-tiny is characterized in that: The step S1 includes the following sub-steps: 1) Adjust the brightness and contrast of the input pedestrian image, and improve the image enhancement effect by color jittering and adding noise. Use the Mosaic data enhancement method to randomly scale the image size, randomly crop and splice the image, and set the size to generate the image for training the model; 2) During the adaptive resizing process, the image is scaled proportionally to the set size, and the minimum black border width to be filled is calculated based on the downsampling multiple.

3. According to claim 1, a lightweight pedestrian detection algorithm based on improved YOLOv7-tiny is characterized in that: In step S2, the improved YOLOv7-tiny model includes a backbone feature subnetwork, a multi-scale feature fusion network and a decoupled detection head; the backbone feature subnetwork extracts features from the input image to obtain a primary feature layer; the multi-scale feature fusion subnetwork extracts higher-level features from top to bottom, and then fuses features of different levels from bottom to top to output fused features; the decoupled detection head passes the fused features into the anchor-free frame to output pedestrian detection results.

4. According to claim 3, a lightweight pedestrian detection algorithm based on improved YOLOv7-tiny is characterized in that: The backbone feature sub-network includes a CBS module, an MP module and an ELCG module connected in sequence. The CBS module includes a common convolution Conv, a BN layer and a SiLu activation function. The ELCG module includes a Ghost-Conv convolution and a CBS module. The MP module includes a maximum pooling layer and a CBS module.

5. According to claim 4, a lightweight pedestrian detection algorithm based on improved YOLOv7-tiny is characterized in that: The ELCG module, the feature map is divided into two branches through the ELCG module, the first branch extracts features through the Ghost-Conv convolution layer, the second branch extracts features through the Ghost-Conv convolution layer and two CBS convolution modules, and the two branches are feature-fused through the Concat layer, and then the Ghost-Conv convolution layer is used to extract features.

6. According to claim 3, a lightweight pedestrian detection algorithm based on improved YOLOv7-tiny is characterized in that: The multi-scale feature fusion network includes a CBS module, an ELCG module, an SPPCSPC module, an upsampling module and an ECA attention mechanism. The SPPCSPC module includes multiple pooling layers for reducing the size of the feature map and the number of channels.

7. A lightweight pedestrian detection algorithm based on improved YOLOv7-tiny according to claim 6, characterized in that: The ECA attention mechanism generates a global channel dependency weight by calculating the one-dimensional local autocorrelation function of each channel and inputs the weight into each channel of the feature map, including the following sub-steps: 1) Perform global average pooling on the input feature map to generate a feature vector. The calculation formula is: ; In the formula, Represents the result of global average pooling of feature maps, Represents the input of the feature map, C represents the number of channels of the feature map, i and j Indicates the counting unit, W and H indicate i and j The maximum value of 2) Perform convolution operation on the feature vector to capture the local dependencies between channels. The calculation formula is: ; In the formula, k represents the channel number mapping convolution kernel size, γ and b are hyperparameters, and odd represents an odd function; 3) Normalize through Sigmoid activation function, and weight the obtained channel weight with the original feature map to obtain the enhanced feature map.

8. The lightweight pedestrian detection algorithm based on improved YOLOv7-tiny according to claim 1, characterized in that: In step S4, the three high-quality fusion feature layers are passed into the decoupled detection head without anchor frames to obtain the detection results, and the loss function is calculated with the assigned sample labels. The loss function is the coordinate loss L CIoU , target confidence loss L dfl and classification loss L cls , the calculation formula is: In the formula, K represents the number of output results of different sizes, α balance represents the weight coefficient of each size result, S represents the length of the grid of the image, B represents the category, α box Denotes the loss function L CIoU The weight coefficient, α dfl Denotes the loss function L dfl The weight coefficient, α cls Denotes the loss function L cls The weight coefficient of i and j Indicates the unit of calculation; Loss function L CIoU The calculation formula is: Where, d o Indicates the distance between the center points of the two boxes, d c Indicates the distance between the diagonals of the two boxes, w, h, w gt and h gt Represent the width and height of the predicted box and the width and height of the real box respectively. b represents the detection box, v Represents the unit of measurement used to measure the similarity of aspect ratio; Loss function L dfl The calculation formula is: In the formula, y and y i+1 Represents two adjacent discrete values ​​of the predicted bounding box position, y represents the position of the real bounding box, S i and S i+1 Indicates that it corresponds to y i and i+1 probability; Loss function L cls The calculation formula is: In the formula, c p and c gt They represent the probability of the predicted category and the probability of the true category respectively.

Citation Information

Patent Citations

  • Real-time pedestrian detection method fused with attention mechanism

    CN112733749A

  • Pedestrian occlusion detection method based on improved YOLOX algorithm

    CN115082855A

  • YOLOv7-tiny-based efficient real-time target detection method

    CN118015253A

  • Vehicle and pedestrian detection method based on improved YOLOv7 and storage medium

    CN118053143A

  • Multi-scale aware pedestrian detection method based on improved full convolutional network

    US20210056351A1

Cited By

  • Pedestrian flow monitoring method and system for dense sitting posture scene

    CN120877213A

  • A method and system for crowd supervision in dense sitting scenarios

    CN120877213B