Urban streetscape vehicle detection method based on YOLOv7-coupling
By introducing the C2f module, VoVGSCSP module and GSConv module into the YOLOv7 algorithm and adding detection heads, the problems of high missed detection rate and low detection accuracy in urban street scene environments are solved, and higher detection accuracy and lower missed detection rate are achieved.
Patent Information
- Application Number
- CN202510034490.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-06-10
AI Technical Summary
In urban street scene environments, the YOLOv7 algorithm has problems with high missed detection rates and low detection accuracy in vehicle detection, especially in night or low-light environments.
The urban street view vehicle detection method based on YOLOv7-coupling is adopted, and the original network structure is replaced by introducing the C2f module, VoVGSCSP module and GSConv module, and a detection head with a feature map size of 10×10 is added to obtain richer gradient flow information.
It improves the accuracy of key point detection, reduces the complexity of the model and missed detection rate, and improves the accuracy of small object detection.
Smart Images

Figure CN120125976A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle detection, and specifically relates to a method for detecting urban street scene vehicles based on YOLOv7-coupling. Background Art
[0002] With the rise of intelligent transportation technology, driverless technology has gradually become an important research topic.
[0003] Vehicle detection is an important supporting technology for driverless, which can improve the traffic safety of driverless and enhance traffic efficiency.
[0004] Vehicle detection mainly relies on object detection. Currently, object detection is mainly divided into two-stage object detection and one-stage object detection. Among them, two-stage object detection first generates region proposals, then classifies samples through a convolutional neural network, and finally realizes object detection through classification and regression, including SPPNet and Faster R-CNN, etc. One-stage object detection directly predicts the category and location of objects. First, it detects through anchor boxes, and then uses non-maximum suppression to obtain the final prediction result to achieve object detection, including SSD and YOLO, etc. Although two-stage object detection has high accuracy, its detection speed is slow and it is not suitable for real-time detection. Although one-stage object detection generally performs average in terms of accuracy, its detection speed is fast and it is suitable for real-time detection. And the real-time detection of vehicles is important for the detection accuracy of vehicles during the driverless process. Therefore, one-stage object detection is used more frequently.
[0005] YOLO is one of the algorithms commonly used in one-stage object detection. YOLOv7 is the seventh-generation algorithm of YOLO. It gradually optimizes the disadvantages of YOLO and achieves a good balance between object detection speed and object detection accuracy. However, there are still deficiencies in its application in complex scenarios such as urban street scenes:
[0006] (1) High missed detection rate: In the urban street scene environment, due to occlusion and overlap of vehicles, the accuracy of the bounding box is not high, and there are a large number of missed detection phenomena.
[0007] (2) Low detection accuracy: The appearance features of vehicles become blurred at night or in low-light environments, resulting in a decrease in the accuracy of vehicle bounding boxes.
[0008] In view of this, a method for detecting urban street scene vehicles based on YOLOv7-coupling is designed to solve the above problems. Summary of the Invention
[0009] To solve the problems raised in the above background art, the present invention provides a method for detecting vehicles in urban street scenes based on YOLOv7-coupling, which has the characteristics of being able to obtain richer gradient flow information, improving the detection accuracy of key points, effectively reducing the model complexity and missed detection rate, and enhancing the detection accuracy of small targets.
[0010] To achieve the above object, the present invention provides the following technical solutions: A method for detecting vehicles in urban street scenes based on YOLOv7-coupling includes the following steps:
[0011] S1: Obtain historical urban street scene images containing vehicles;
[0012] S2: Construct a YOLOv7-coupling model, which includes a Backbone network, a Neck network, and a Head network connected in sequence;
[0013] In the Backbone network, introduce a C2f module to replace the E-ELAN module;
[0014] In the Neck network, introduce a VoVGSCSP module to replace the Multi_Concat_Block module, and at the same time introduce a GSConv module to replace the Conv module of the first channel on the left;
[0015] In the Head network, add a detection head with a feature map size of 10×10;
[0016] S3: Train the YOLOv7-coupling model based on the historical urban street scene images containing vehicles until convergence;
[0017] S4: Detect vehicles in urban street scenes based on the YOLOv7-coupling model.
[0018] Further, in the step S2, the C2f module includes two CBS modules and N Bottleneck modules. The first CBS module is connected to the first Bottleneck module, adjacent Bottleneck modules are connected, and after convolutional splicing, it is connected to the second CBS module. The convolutional kernel size of the CBS module is 1×1, and the Bottleneck module includes two connected CBS modules with a convolutional kernel size of 3×3.
[0019] Further, in the step S2, the VoVGSCSP module includes three Conv modules and one GSBottleneck module. The first Conv module is connected to the GSBottleneck module and the second Conv module. After the outputs of the GSBottleneck module and the second Conv module are aggregated, they are connected to the third Conv module. Among them, the GSBottleneck module includes two GSConv modules and one DWConv module. The first GSConv module is connected to the second GSConv module, and the outputs of the DWConv module and the second GSConv module are added together.
[0020] Further, in the step S2, the GSConv module includes one Conv module, one DWConv module and one shuffle module. The Conv module is connected to the DWConv module, and after the outputs are concatenated, they are connected to the shuffle module.
[0021] Further, in the step S3, the history of training the YOLOv7-coupling model includes urban street view images of vehicles with a size of 640 pixels × 640 pixels.
[0022] Further, in the step S4, the vehicle images detected by the YOLOv7-coupling model in urban street views have a size of 640 pixels × 640 pixels ×.
[0023] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0024] The present invention introduces the C2f module to replace the E-ELAN module, introduces the VoVGSCSP module to replace the Multi_Concat_Block module, introduces the GSConv module to replace the first Conv module in the leftmost channel, and at the same time adds a detection head with a feature map size of 10×10. Compared with the prior art, it can obtain richer gradient flow information, improve the key point detection accuracy, effectively reduce the model complexity and the missed detection rate, and improve the small target detection accuracy. Description of the Drawings
[0025] Figure 1 It is the structural diagram of the YOLOv7-coupling model of the present invention;
[0026] Figure 2 It is the structural diagram of the C2f module of the present invention;
[0027] Figure 3 It is the structural diagram of the GSConv module of the present invention;
[0028] Figure 4This is the structural diagram of the VoVGSCSP module of the present invention;
[0029] Figure 5 This is the structural diagram of the P6 layer detection head of the present invention;
[0030] Figure 6 This is the ablation experiment result diagram of the YOLOv7-coupling model of the present invention;
[0031] Figure 7 This is the result diagram of the training and test loss values of each model of the present invention;
[0032] Figure 8 This is the comparison diagram of the detection results of each model of the present invention;
[0033] Figure 9 This is the comparison diagram of TP+FP and missed detection rate of each model of the present invention;
[0034] Figure 10 This is the comparison diagram of the detection results of each model of the present invention. Detailed implementation manners
[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0036] The present invention provides the following technical solutions: A method for detecting vehicles in urban street scenes based on YOLOv7-coupling, including the following steps:
[0037] S1: Obtain historical urban street scene images containing vehicles;
[0038] S2: Construct a YOLOv7-coupling model, where the YOLOv7-coupling model includes a Backbone network, a Neck network, and a Head network connected in sequence;
[0039] In the Backbone network, introduce a C2f module to replace the E-ELAN module;
[0040] In the Neck network, introduce a VoVGSCSP module to replace the Multi_Concat_Block module, and at the same time introduce a GSConv module to replace the Conv module in the first channel on the left;
[0041] In the Head network, add a detection head with a feature map size of 10×10;
[0042] The YOLOv7-coupling model structure is as shown in the appendix Figure 1 as follows;
[0043] S3: Train the YOLOv7-coupling model based on historical urban street view images containing vehicles until convergence;
[0044] S4: Detect vehicles in urban street views based on the YOLOv7-coupling model.
[0045] Specifically, in step S2, the C2f module includes two CBS modules and N Bottleneck modules. The first CBS module is connected to the first Bottleneck module, adjacent Bottleneck modules are connected, and after convolutional splicing, it is connected to the second CBS module. The convolutional kernel size of the CBS module is 1×1. The Bottleneck module includes two connected CBS modules, and the convolutional kernel size of the CBS module is 3×3;
[0046] The C2f module structure is as shown in the appendix Figure 2 as follows. First, perform a convolutional operation on the input feature map through the first 1×1 CBS module. And in order to reduce the computational amount and the number of parameters while retaining sufficient feature information, divide the result into two parts according to the number of channels. Then use N Bottleneck modules with different scales and numbers of channels to perform 3×3 convolutional operations on the input feature map. After that, splice all the elements after convolution along the channel dimension, and perform a convolutional operation through the second 1×1 CBS module to obtain the final output.
[0047] Specifically, in step S2, the VoVGSCSP module includes three Conv modules and one GSBottleneck module. The first Conv module is connected to the GSBottleneck module and the second Conv module. After the outputs of the GSBottleneck module and the second Conv module are aggregated, they are connected to the third Conv module. Among them, the GSBottleneck module includes two GSConv modules and one DWConv module. The first GSConv module is connected to the second GSConv module, and the outputs of the DWConv module and the second GSConv module are added;
[0048] The VoVGSCSP module structure is as shown in the appendix Figure 3 as follows. First, perform a convolutional operation on the input feature map through the first Conv module, divide the result into two parts. One part enters the second Conv module for convolutional operation, and the other part enters the GSBottleneck module for convolutional operation. Then splice the output results of these two parts, and then enter the third Conv module for convolutional operation;
[0049] The structure of the GSBottleneck module is as shown in the appendix Figure 4 As shown, a part enters two GSConv modules for convolution operations, and another part enters the DWConv module for depthwise separable convolution operations, and then the output results of these two parts are added together.
[0050] Specifically, in step S2, the GSConv module includes a Conv module, a DWConv module, and a shuffle module. The Conv module is connected to the DWConv module, and the output after splicing is connected to the shuffle module;
[0051] The structure of the GSConv module is as shown in the appendix Figure 5 As shown, the Conv module receives graph-structured data containing node features and an adjacency matrix for convolution operations. The DWConv module samples from the neighboring values of each node in the graph based on a predetermined sampling strategy to form a subgraph, aiming to capture the local information around the nodes while reducing the computational complexity. The shuffle module aggregates the features within the subgraph and integrates the feature information within the subgraph to obtain more comprehensive node features. The aggregated features will be used as the new features of the nodes for the next layer of calculation, and then these updated node features are output.
[0052] Specifically, in step S3, the historical data for training the YOLOv7-coupling model consists of urban street view images containing vehicles with a size of 640×640.
[0053] Specifically, in step S4, the YOLOv7-coupling model detects vehicle images in urban street views with a size of 640 pixels×640 pixels.
[0054] Experimental environment and parameter configuration
[0055] Experiments were conducted using the PyTorch deep learning framework. Windows 11 was selected as the operating system, and a GPU of the Geforce RTX3050 model with a video memory size of 4G was equipped. The processor selected was the Intel core i5-11400H;
[0056] The initial learning rate was set to 0.0001, the weight decay was set to 0.0005, and to optimize the training process of the model, the momentum value was set to 0.9.
[0057] Experimental results and performance analysis
[0058] The training and testing are respectively carried out using the self-built vehicle dataset and the public COCO dataset. Among them, the self-built vehicle dataset is obtained by shooting videos on urban roads using the built-in camera of the IQOO Z1 mobile phone. Due to problems such as a single background and a fixed shooting angle during the shooting process, data augmentation technology and web crawler technology are used to expand the image data;
[0059] The urban street view vehicle image data is manually labeled to construct a standard vehicle dataset, and the resolution of each image in the dataset is 1080 pixels × 1920 pixels;
[0060] The dataset is divided into a training set and a testing set according to a ratio of 4:1 for the experiment;
[0061] The training and testing results among different models use the mean average precision (mAP) of the category as the main evaluation index, and the precision, recall, function loss value, number of parameters, and parameter calculation amount are used as auxiliary reference indexes;
[0062] The ablation experiment results of the YOLOv7-coupling model are as attached Figure 6 , and the test results of the precision indexes among different models are shown in the following table:
[0063] Table 1 Test results of the precision indexes of five models on the self-made dataset
[0064]
[0065] As can be seen from the above figures and tables:
[0066] The mean average precision of the YOLOv7-coupling model is 95.2%, and the precision and recall are 98.5% and 98% respectively. Compared with the YOLOv7 model, there is an improvement, indicating that the C2f module reconstructs the YOLOv7 backbone network, which can effectively obtain richer gradient flow information and has a certain improvement in the key point detection accuracy;
[0067] The mean average precision of the YOLOv7-coupling model is 95.2%, and the precision and recall are 98.5% and 98% respectively. Compared with the YOLOv7-efficientViT M0 model and the YOLOv7-EfficientFormerV2 model, there is an improvement, indicating that the improvements of the C2f module, VoVGSCSP module, and GSConv module can improve the expression ability of the model, making the output of the convolution calculation as close to SC as possible, and taking into account the lightweight and precision of the model;
[0068] The average precision of the YOLOv7-coupling model is 95.2%, and the precision and recall rates are 98.5% and 98% respectively, showing an improvement compared to the YOLOv7-tscode model, indicating that the improvements in the C2f module, VoVGSCSP module, GSConv module, and detection head can enhance the detection performance of the model;
[0069] The training and test loss values and detection results among different models are shown in Appendix Figure 7 and Appendix Figure 8 as follows:
[0070] As can be seen from the above figures:
[0071] The improvements in the C2f module, VoVGSCSP module, GSConv module, and detection head of the YOLOv7-coupling model can effectively detect small targets in the figure, and the confidence level has also increased compared to before, and each loss value has decreased to the minimum state, indicating that the improved model can effectively increase the non-linear ability of the network, can better capture the details and context information of the target, and thus improve the accuracy of target recognition and positioning;
[0072] The test results of the quantization performance indicators among different models are shown in the following table:
[0073] Table 2 Test Results of Quantization Performance Indicators of Five Models on the Self-made Dataset
[0074]
[0075] As can be seen from the above table:
[0076] The YOLOv7-coupling model has the fewest parameters but the highest average precision, indicating that the introduction of the VoVGSCSP module and GSConv module reduces the model complexity while taking into account the lightweight and precision of the model;
[0077] The YOLOv7-coupling model has the least amount of computation, indicating that its floating-point operations per second are the fewest, indicating that using the C2f module to reconstruct the YOLOv7 backbone network can ensure lightweight while obtaining more abundant gradient flow information;
[0078] The YOLOv7-coupling model has a shorter training time but the highest precision, indicating that the improvements in the C2f module, VoVGSCSP module, GSConv module, and detection head can ensure precision while taking into account the lightweight of the model;
[0079] The comparison results of true positives TP, false positives FP, and missed detection rates of different models are shown in Appendix Figure 9 and the following table:
[0080] Comparison Results of TP, FP and Miss Detection Rates of Five Models in the Self-made Dataset
[0081]
[0082] As can be seen from the above figures and tables:
[0083] Compared with the other four models, the miss detection rate of the YOLOv7-coupling model is only 3.5%, which is the lowest. The miss detection rate has decreased by about 10%, and the TP value is also the highest. This indicates that reconstructing the YOLOv7 backbone network using the C2f module can obtain richer gradient flow information, improve the key point detection accuracy. Adding a smaller detection head to detect smaller-scale objects can improve the detection accuracy of small targets and reduce the model's miss detection rate. Replacing the multi-branch stacked convolution module and conventional convolution at the connection between the feature enhancement network and the detection head with the VoVGSCSP module and GSConv module can reduce the model complexity while improving the accuracy;
[0084] The model was trained using the COCO dataset and validated using the test set in COCO. This dataset covers 80 common object categories and includes different object sizes, shapes, poses, and backgrounds. In addition, the COCO dataset also provides detailed annotation information, including object bounding boxes, object categories, object segmentation masks, and object key points, etc.;
[0085] Under the same experimental configuration environment and dataset, experiments were conducted on different YOLO series models, and the experimental results are shown in the following table:
[0086] Table 4 Comparison with Models in the Same Series on the COCO Dataset
[0087]
[0088] As can be seen from the above table:
[0089] The YOLOv7-coupling model is the smallest model with the lowest number of parameters and computational complexity. Its mAP is 51.7%, which is the highest. The YOLOv7-coupling model has a significant reduction in both the number of parameters and computational complexity compared to the YOLOv7 model, but it has achieved a relatively high improvement, indicating that the YOLOv7-coupling model can maintain a high detection accuracy while reducing the model size and computational complexity;
[0090] The detection effect diagrams of different models on the selected part of the COCO test dataset are shown in the appendix Figure 10 as follows:
[0091] As can be seen from the above figures:
[0092] The improvements to the C2f module, VoVGSCSP module, GSConv module, and detection head of the YOLOv7-coupling model are more sensitive to the detection of small targets, can well adapt to small targets, and improve the detection performance while maintaining real-time performance.
[0093] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. The urban street scene vehicle detection method based on YOLOv7-coupling is characterized by: The following steps are involved: S1: Obtaining city street view images containing vehicles in history; S2: Build a YOLOv7-coupling model. The YOLOv7-coupling model includes a Backbone network, a Neck network, and a Head network connected in sequence. In the Backbone network, the C2f module is introduced to replace the E-ELAN module; In the Neck network, the VoVGSCSP module is introduced to replace the Multi_Concat_Block module, and the GSConv module is introduced to replace the Conv module of the first channel on the left; In the Head network, add a detection head with a feature map size of 10×10; S3: Train the YOLOv7-coupling model based on historical city street scene images containing vehicles until convergence; S4: Detecting vehicles in urban street scenes based on the YOLOv7-coupling model.
2. The urban street scene vehicle detection method based on YOLOv7-coupling according to claim 1, characterized in that: In step S2, the C2f module includes two CBS modules and N Bottleneck modules, the first CBS module is connected to the first Bottleneck module, adjacent Bottleneck modules are connected, and are connected to the second CBS module after convolution splicing. The convolution kernel size of the CBS module is 1×1, and the Bottleneck module includes two connected CBS modules, and the convolution kernel size of the CBS module is 3×3.
3. The urban street scene vehicle detection method based on YOLOv7-coupling according to claim 1, characterized in that: In step S2, the VoVGSCSP module includes three Conv modules and one GSBottleneck module, the first Conv module is connected to the GSBottleneck module and the second Conv module, the outputs of the GSBottleneck module and the second Conv module are aggregated and connected to the third Conv module, wherein the GSBottleneck module includes two GSConv modules and one DWConv module, the first GSConv module is connected to the second GSConv module, and the outputs of the DWConv module and the second GSConv module are added.
4. The urban street scene vehicle detection method based on YOLOv7-coupling according to claim 1, characterized in that: In step S2, the GSConv module includes a Conv module, a DWConv module and a shuffle module. The Conv module is connected to the DWConv module, and the outputs are connected to the shuffle module after splicing.
5. The urban street scene vehicle detection method based on YOLOv7-coupling according to claim 1, characterized in that: In step S3, the history of training the YOLOv7-coupling model includes city street scene images with a size of 640×640 containing vehicles.
6. The urban street scene vehicle detection method based on YOLOv7-coupling according to claim 1, characterized in that: In step S4, the YOLOv7-coupling model detects that the vehicle image in the urban street scene has a size of 640×640.