An autonomous driving target detection method based on improved YOLOv8 network
By improving the Backbone and Neck modules of the YOLOv8 network, the image edge and spatial information extraction capabilities are enhanced, and the problem of insufficient detection capabilities of the YOLOv8 model for small objects is solved, achieving higher detection accuracy.
Patent Information
- Application Number
- CN202510649873.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The basic YOLOv8 model still has room for improvement in target detection accuracy, especially the detection ability of small targets is weak.
By designing the C2f-SCC module to replace the Bottleneck module in the Backbone module of the YOLOv8 network, and using the SOAPN structure to replace the PAFPN structure of the Neck module, we will enhance the image edge and spatial information extraction capabilities and learn feature representations from global to local.
It significantly improves the accuracy of target detection in autonomous driving scenarios, especially the detection performance of small targets.
Smart Images

Figure CN120182727B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target detection technology in the field of computer technology, and in particular to an autonomous driving target detection method based on an improved YOLOv8 network. Background Art
[0002] Autonomous driving has been a hot research area in recent years. Its goal is to enable cars to operate safely on the road, accurately detecting and identifying various objects such as pedestrians, non-motorized vehicles, traffic signs, and lane markings in real time in complex road scenarios. This allows for autonomous driving control. Currently, my country attaches great importance to the development of the autonomous driving industry. Addressing the most critical safety issues of autonomous driving technology, increasing consumer trust in autonomous vehicles, and thereby boosting sales are major challenges for the development of the industry. Object detection systems are a crucial component of autonomous driving systems, and their detection performance is crucial to determining the safety of autonomous driving.
[0003] Traditional object detection algorithms use sliding windows of varying sizes to frame a portion of an image as a candidate region to extract relevant visual features and then use classifiers for recognition, such as the Viola-Jones detector, histogram of oriented gradients, and deformable component models. Because sliding window-based region selection strategies are non-targeted, time-intensive, and suffer from window redundancy, and because manually extracted features lack robustness to diverse variations, they are not well suited for object detection in autonomous driving scenarios. Deep learning-based detection methods utilize DNNs to autonomously learn image features, eliminating the need for manual feature extraction. These methods possess deeper architectures and are capable of learning more complex feature representations than shallower architectures. They also unify feature extraction and classifier learning within a single framework, enabling end-to-end learning. YOLO (You Only Look Once), a representative single-stage detection method within deep learning-based detection, is widely used in autonomous driving object detection systems due to its advantages in detection speed and accuracy.
[0004] However, the basic YOLOv8 model still has room for improvement in target detection accuracy, and its ability to detect small targets is slightly weak. Summary of the Invention
[0005] In view of the above problems, the purpose of the present invention is to provide an autonomous driving target detection method based on an improved YOLOv8 network, which is used to effectively learn feature representations from global to local, thereby improving the detection performance of small targets and overcoming the shortcomings of the above-mentioned prior art.
[0006] The present invention provides an autonomous driving target detection method based on an improved YOLOv8 network, which specifically includes the following steps:
[0007] Step 1: Build a dataset, download the KITTI autonomous driving dataset, select the self-labeled dataset in the KITTI autonomous driving dataset for preprocessing, and obtain the dataset for training the model;
[0008] Step 2: Design a C2f-SCC module. The C2f-SCC module is a spliced fusion Sobel convolution and basic convolution module. The SCC module replaces the Bottleneck module in the C2f module in the Backbone module of the YOLOv8 network model. The edge features of the original image are extracted by the SobelConv branch module, and then the features of the original image are extracted by the Conv branch module. The edge features extracted by the SobelConv branch module and the features extracted by the Conv branch module are then spliced together.
[0009] Step 3: Design a SOAPN structure, which is a small target enhancement pyramid network structure; use the SOAPN structure to replace the PAFPN structure in the Neck module of the YOLOv8 network model, which is a path aggregation feature pyramid network structure;
[0010] Step 4: Put the training set in the dataset constructed in step 1 into the SS-YOLO network model obtained after the improvement in steps 2 and 3 for training to obtain the optimal autonomous driving target detection model after training, wherein the SS-YOLO network model is a Sobel convolution and basic convolution small target enhanced pyramid network model;
[0011] Step 5: Put the test set in the dataset constructed in step 1 into the optimal autonomous driving target detection model for detection, and obtain the target position and target category in the image of the test set, that is, the final detection result.
[0012] As the preferred embodiment of the present invention, step 1 also includes the following steps:
[0013] Step 101: Preprocess the dataset by randomly flipping, randomly cropping, scaling, horizontally mirroring, and enhancing the images in the self-labeled dataset; then divide the processed dataset into a training set and a test set according to the ratio.
[0014] As the preferred embodiment of the present invention, the following steps are also included in step 2:
[0015] Step 201: Using the C2f-SCC module, input the feature map X of the original image through the SobelConv branch module, and calculate the edge information in the horizontal and vertical directions through the Sobel operator of the SobelConv branch module; let the horizontal and vertical directions of the Sobel operator be and , then the edge features The calculation formula is as follows:
[0016] ;
[0017] in, represents the convolution operation;
[0018] Step 202: Use the Conv branch module to extract the features of the original image through standard 3x3 convolution , recorded as:
[0019] ;
[0020] in, Represents a 3x3 convolution operation;
[0021] Step 203: Splice the branch outputs of the SobelConv branch module and the Conv branch module in the channel dimension to obtain the spliced features , recorded as:
[0022] ;
[0023] Then, channel compression is performed through 1×1 convolution for information integration to obtain the integrated features. , recorded as:
[0024] ;
[0025] in, Represents 1×1 convolution;
[0026] Step 204: Through 1×1 convolution again, the enhanced feature map X' is finally obtained:
[0027] ;
[0028] Among them, the feature map X of the input original image is added to the feature map through the residual connection , used for information integrity.
[0029] As the preferred embodiment of the present invention, step 3 also includes the following steps:
[0030] Step 301: Using the SOAPN structure, first, the features of the P5 feature layer, P4 feature layer, and P3 feature layer of the Backbone module are transferred layer by layer along a top-down path to enhance semantic information;
[0031] Step 302: The P2 feature layer of the Backbone module is then passed through the SPDConv module to obtain features containing small target information in the P2 feature layer, which are then fused with features from the P5 feature layer, the P4 feature layer, and the P3 feature layer to obtain fused features.
[0032] Step 303: Then, the fusion features in step 302 are integrated using the CSP-OmniKernel module to obtain integrated features. The CSP-OmniKernel module is a cross-stage part and full-core module;
[0033] Step 304: Finally, the integrated features are transferred layer by layer along a bottom-up path to enhance the detail expression.
[0034] The beneficial effects of this invention are as follows: Based on the YOLOv8 network model, the C2f module of Backbone is first improved with the SCC (SobelConv-Conv) module, resulting in the C2f-SCC module. This enhances the network's ability to extract image edge and spatial information. Secondly, the SOAPN (Small Object Augment Pyramid Network) structure is designed to replace Neck's original PAFPN (Path Aggregation Feature Pyramid Network) structure. This effectively learns feature representations from global to local, thereby improving the detection performance of small objects. These improvements result in the final SS-YOLO network model, which can effectively improve the accuracy of object detection (especially for small objects) in autonomous driving scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] By referring to the following description in conjunction with the accompanying drawings, and with a more complete understanding of the present invention, other objects and results of the present invention will become more clear and easy to understand. In the accompanying drawings:
[0036] Figure 1 It is the overall structure diagram of the present invention; wherein, Backbone represents the backbone, Neck represents the neck, Head represents the head, Input represents the input, Conv represents the convolution, C2f-SCC represents the splicing fusion Sobel convolution and basic convolution, SPPF represents the spatial pyramid fast pooling, Upsample represents the downsampling, Concat represents the connection, C2f represents the splicing fusion, SPDConv represents the spatial depth conversion convolution, and CSP-Omnikernel represents the cross-stage part and the full core;
[0037] Figure 2 The C2f-SCC module structure diagram of the present invention; wherein Input represents input, SobelConv represents Sobel convolution, Conv represents convolution, Concat represents connection, and add represents superposition;
[0038] Figure 3 This is a structural diagram of the SobelConv branch module of the present invention; wherein Input represents input, Sobel-x represents the edge information in the horizontal direction of the Sobel operator, and Sobel-y represents the edge information in the vertical direction of the Sobel operator;
[0039] Figure 4 The SOAPN structure diagram of the present invention; wherein Upsample represents downsampling, Concat represents connection, C2f represents splicing fusion, SPDConv represents spatial depth conversion convolution, and CSP-Omnikernel represents cross-stage part and full core;
[0040] Figure 5 This is a structural diagram of the CSP-OmniKernel module of the present invention; wherein Input represents input, Conv represents convolution, Split represents segmentation, Omnikernel represents full kernel, and Concat represents connection;
[0041] Figure 6 This is the structural diagram of the OmniKernel module of the present invention; wherein, Conv represents convolution, Local represents local branch, Large represents large branch, Global represents global branch, DConv represents depthwise convolution, DCAM represents dual-domain channel attention module, FSAM represents frequency-based spatial attention module, and add represents superposition. DETAILED DESCRIPTION
[0042] Example 1
[0043] See Figure 1-6 The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0044] This embodiment provides an autonomous driving target detection method based on an improved YOLOv8 network, including the following steps:
[0045] Step 1: Build a dataset. Download the KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute) autonomous driving dataset, jointly created by the Karlsruhe Institute of Technology and Toyota Technological Institute of America. Preprocess the self-labeled dataset from KITTI to obtain the dataset for model training.
[0046] Step 101: Preprocess the dataset by randomly flipping, randomly cropping, scaling, horizontally mirroring, and enhancing the 7,480 images in the self-labeled dataset; the image scaling size is 640×640, and the processed dataset is then divided into a training set and a test set in a ratio of 8:2.
[0047] Step 2: Design a C2f-SCC module. The C2f-SCC module is a spliced and fused Sobel convolution and basic convolution module. The SCC (SobelConv-Conv) module replaces the Bottleneck module in the C2f (splicing and fusion) module in the Backbone module of the YOLOv8 network model. The SCC module is an efficient front-end module for image recognition tasks. It extracts edge features of the original image through the SobelConv branch module, and then extracts features of the original image through the Conv branch module to retain rich spatial details. The features extracted by the SobelConv branch module and the Conv branch module are then spliced together. This allows the learned feature representation to contain both rich edge information and spatial information, which can more comprehensively depict the image content, thereby improving the accuracy of image recognition.
[0048] The specific workflow of the C2f-SCC module is as follows:
[0049] Step 201: Using the C2f-SCC module, input the feature map X of the original image through the SobelConv branch module, and calculate the edge information in the horizontal and vertical directions through the Sobel operator of the SobelConv branch module; let the horizontal and vertical directions of the Sobel operator be and respectively. , then the edge features The calculation formula is as follows:
[0050] ;
[0051] in, represents the convolution operation;
[0052] Step 202: Use the Conv branch module to extract the features of the original image through standard 3x3 convolution , recorded as:
[0053] ;
[0054] in, Represents a 3x3 convolution operation;
[0055] Step 203: Splice the branch outputs of the SobelConv branch module and the Conv branch module in the channel dimension to obtain the spliced features , recorded as:
[0056] ;
[0057] Then, channel compression is performed through 1×1 convolution for information integration to obtain the integrated features. , recorded as:
[0058] ;
[0059] in, Represents 1×1 convolution;
[0060] Step 204: Through 1×1 convolution again, the enhanced feature map X' is finally obtained:
[0061] ;
[0062] Among them, the feature map X of the input original image is added to the feature map through the residual connection , used for information integrity.
[0063] Step 3: Design the SOAPN (Small Object Augment Pyramid Network) structure, which is a small object augmentation pyramid network structure. Use the SOAPN structure to replace the PAFPN (Path Aggregation Feature Pyramid Network) structure in the Neck module of the YOLOv8 network model. The PAFPN structure is a path aggregation feature pyramid network structure. The SOAPN structure can effectively learn feature representations from global to local, thereby improving the detection performance of small objects.
[0064] The specific workflow of the SOAPN structure is as follows:
[0065] Step 301: Using the SOAPN structure, first, the features of the P5 feature layer, P4 feature layer, and P3 feature layer of the Backbone module are transferred layer by layer along a top-down path to enhance semantic information;
[0066] Step 302: The P2 feature layer of the Backbone module is then passed through the SPDConv (Spatial Depth Convolution) module for low-resolution images and small objects, so that the P2 feature layer obtains features containing small object information, which are then fused with the features of the P5 feature layer, the P4 feature layer, and the P3 feature layer to obtain fused features.
[0067] Step 303: Then, the fused features in step 302 are integrated using the CSP-OmniKernel module obtained by combining the CSP (Cross Stage Partial) module and the OmniKernel module to obtain integrated features. The CSP-OmniKernel module is a cross-stage partial and full-core module. Among them, the OmniKernel (full-core) module of the CSP-OmniKernel module consists of a Local branch, a Large branch, and a Global branch. It can effectively learn feature representations from global to local, thereby improving the detection performance of small objects;
[0068] Step 304: Finally, the integrated features are transferred layer by layer along a bottom-up path to enhance the detail expression.
[0069] Step 4: Put the training set in the dataset constructed in step 1 into the SS-YOLO network model obtained after the improvement in steps 2 and 3 for training to obtain the optimal autonomous driving target detection model after training, wherein the SS-YOLO network model is a Sobel convolution and basic convolution small target enhanced pyramid network model;
[0070] Step 5: Put the test set in the dataset constructed in step 1 into the optimal autonomous driving target detection model for detection, and obtain the target position and target category in the image of the test set, that is, the final detection result.
[0071] Example 2
[0072] This example builds a training model on a Linux server, and the virtual environment used is described as follows:
[0073] Create a conda virtual environment with Python 3.8 installed, then install the PyTorch neural network architecture (torch 2.2.2 + cu121). The experiment uses an NVIDIA GeForce RTX 4090 GPU. After setting the initial training parameters, apply the dataset to the improved SS-YOLO network model for object detection training. After training, the optimal autonomous driving object detection model is obtained. The test set is then applied to the model for testing to verify its detection performance.
[0074] The model is evaluated using the following metrics: precision, recall, and average precision (mAP).
[0075] Precision refers to the ratio of correctly predicted positives (TP) to all positive predictions (TP+FP), and is calculated as follows:
[0076] ;
[0077] Recall refers to the ratio of correctly predicted positive (TP) to actual positive (TP+FN), and is calculated as follows:
[0078] ;
[0079] mAP, or average precision across all categories, is a commonly used metric for evaluating the performance of object detection algorithms. It comprehensively considers the model's precision and recall across different categories and is a comprehensive evaluation metric for the model's overall performance. The calculation formula is as follows:
[0080] ;
[0081] in, represents the number of categories, Represents the average precision of each category.
[0082] In order to further understand the gain effect of each module on the YOLOv8 network structure, multiple groups of ablation experiments were set up. The ablation experiments are shown in Table 1:
[0083] Table 1: Ablation experiment results
[0084] Model P(Precision) / % R(Recall) / % mAP@0.5 / % mAP@.5:.95 / % YOLOv8s 88.9 82.7 89.2 65.1 YOLOv8s+SCC 92.5 80.8 90.2 65.5 YOLOv8s+SCC+SOAPN 93.2 82.7 91.6 68.6
[0085] As shown in Table 1, the final optimized model achieves a 4.3% improvement in precision (P), 2.4% improvement in mAP@0.5, and 3.5% improvement in mAP@.5:.95 compared to the initial model. This demonstrates that the method in this example can effectively improve the model's detection performance.
[0086] In summary, this example uses YOLOv8 to build an autonomous driving object detection model, replacing the Bottleneck module with the SCC module and the PAFPN architecture with the SOAPN architecture. This effectively improves the YOLO model's object detection accuracy in autonomous driving scenarios, particularly for small objects.
[0087] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A target detection method for autonomous driving based on an improved YOLOv8 network, characterized in that: The following steps are involved: Step 1: Build a dataset, download the KITTI autonomous driving dataset, select the self-labeled dataset in the KITTI autonomous driving dataset for preprocessing, and obtain the dataset for training the model; Step 2: Design a C2f-SCC module. The C2f-SCC module is a spliced fusion Sobel convolution and basic convolution module. The SCC module replaces the Bottleneck module in the C2f module in the Backbone module of the YOLOv8 network model. The edge features of the original image are extracted by the SobelConv branch module, and then the features of the original image are extracted by the Conv branch module. The edge features extracted by the SobelConv branch module and the features extracted by the Conv branch module are then spliced together. Step 3: Design a SOAPN structure, which is a small target enhancement pyramid network structure; use the SOAPN structure to replace the PAFPN structure in the Neck module of the YOLOv8 network model, which is a path aggregation feature pyramid network structure; Step 4: Put the training set in the dataset constructed in step 1 into the SS-YOLO network model obtained after the improvement in steps 2 and 3 for training to obtain the optimal autonomous driving target detection model after training, wherein the SS-YOLO network model is a Sobel convolution and basic convolution small target enhanced pyramid network model; Step 5: Put the test set in the dataset constructed in step 1 into the optimal autonomous driving target detection model for detection, and obtain the target position and target category in the image of the test set, that is, the final detection result.
2. The automatic driving target detection method based on the improved YOLOv8 network according to claim 1 is characterized in that: Step 1 also includes the following steps: Step 101: Preprocess the dataset by randomly flipping, randomly cropping, scaling, horizontally mirroring, and enhancing the images in the self-labeled dataset; then divide the processed dataset into a training set and a test set according to the ratio.
3. The automatic driving target detection method based on the improved YOLOv8 network according to claim 1 is characterized in that: Step 2 also includes the following steps: Step 201: Using the C2f-SCC module, input the feature map X of the original image through the SobelConv branch module, and calculate the edge information in the horizontal and vertical directions through the Sobel operator of the SobelConv branch module; let the horizontal and vertical directions of the Sobel operator be and , then the edge features The calculation formula is as follows: ; in, represents the convolution operation; Step 202: Use the Conv branch module to extract the features of the feature map X of the original image through standard 3x3 convolution , recorded as: ; in, Represents a 3x3 convolution operation; Step 203: Splice the branch outputs of the SobelConv branch module and the Conv branch module in channel dimension to obtain the spliced features , recorded as: ; Then, channel compression is performed through 1×1 convolution for information integration to obtain the integrated features. , recorded as: ; in, Represents 1×1 convolution; Step 204: Through 1×1 convolution again, the enhanced feature map X' is finally obtained: ; Among them, the feature map X of the input original image is added to the feature map through the residual connection , used for information integrity.
4. The automatic driving target detection method based on the improved YOLOv8 network according to claim 1 is characterized in that: Step 3 also includes the following steps: Step 301: Using the SOAPN structure, first, the features of the P5 feature layer, P4 feature layer, and P3 feature layer of the Backbone module are transferred layer by layer along a top-down path to enhance semantic information; Step 302: The P2 feature layer of the Backbone module is then passed through the SPDConv module to obtain features containing small target information in the P2 feature layer, which are then fused with features from the P5 feature layer, the P4 feature layer, and the P3 feature layer to obtain fused features. Step 303: Then, the fusion features in step 302 are integrated using the CSP-OmniKernel module to obtain integrated features. The CSP-OmniKernel module is a cross-stage part and full-core module; Step 304: Finally, the integrated features are transferred layer by layer along a bottom-up path to enhance the detail expression.
Citation Information
Patent Citations
GIS infrared image recognition system and method based on improved YOLOv5
CN117975040A
Aerial image small target detection method and system based on improved YOLOv8 algorithm, medium and processor
CN119625269A