Automatic driving target detection method based on improved YOLOv8
By improving the backbone, neck, and head parts of the YOLOv8 model and optimizing its network structure, the balance between accuracy and real-time performance in autonomous driving target detection is resolved, and the performance and computational efficiency of the detection algorithm in complex scenarios are improved.
Patent Information
- Application Number
- CN202510745104.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-19
AI Technical Summary
Existing autonomous driving target detection algorithms have difficulty balancing accuracy and real-time performance. High-precision algorithms have large computational complexity and are unable to meet real-time requirements, while lightweight algorithms have reduced detection accuracy in complex scenarios.
The YOLOv8 model is improved by replacing the C2f module in the backbone with the C2f_DCNv4 module. The self-attention mechanism module COT is introduced in the neck part. The head part uses the LSDCH lightweight shared convolutional detection head. The network structure is optimized to improve computational efficiency and detection accuracy.
By optimizing the YOLOv8 model, the detection accuracy and real-time performance in complex scenarios are improved, adapting to the deployment requirements of edge devices, and enhancing the detection capabilities of small and obscured targets.
Smart Images

Figure CN120673362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image detection and autonomous driving, and specifically to an autonomous driving target detection method based on improved YOLOv8. Background Art
[0002] Object recognition in autonomous driving is a core technology for ensuring driving safety. It primarily uses sensors and algorithms to perceive and classify road objects, such as pedestrians, vehicles, and traffic signs. At the algorithmic level, deep learning-based convolutional neural networks (CNNs) dominate, with network structures such as Faster R-CNN, YOLO, and SSD enabling rapid and accurate object detection. Furthermore, network design is continuously optimized to enhance performance for small object detection and occluded object recognition. In terms of data acquisition, cameras capture 2D images for object detection using CNNs. LiDAR also collects 3D point cloud data, extracting features using technologies such as PointNet to achieve object localization and recognition in 3D space. Furthermore, the integration of information from multiple sensors, including LiDAR, cameras, and millimeter-wave radar, significantly improves detection accuracy and environmental adaptability, enhancing system reliability in complex weather conditions and providing solid support for autonomous driving decision-making.
[0003] In autonomous driving object detection, balancing accuracy and real-time performance is a challenging issue. High-precision object detection algorithms, such as Faster R-CNN, utilize complex network structures and multi-stage processing to extract rich target features and accurately identify various targets. However, the extensive convolutional computations and region proposal processes lead to a significant increase in computational complexity, making it difficult to meet the real-time requirements of autonomous driving systems. Lightweight algorithms, such as YOLO, directly output detection results through a single forward propagation, significantly improving detection speed. However, in complex scenarios, their simplified network structure struggles to fully capture target details, significantly reducing detection accuracy. Summary of the Invention
[0004] To address the above issues, the present invention proposes an autonomous driving target detection method based on an improved YOLOv8. This method constructs an autonomous driving target detection dataset and a basic YOLOv8n model, and then improves upon this model to address the issues raised in the above background technology. The present invention provides the following technical solutions:
[0005] An autonomous driving target detection method based on improved YOLOv8 includes the following steps:
[0006] Step 1: Collect autonomous driving scene images, annotate the objects, generate bounding boxes and category labels, randomly divide the dataset into training set, validation set, and test set in an 8:1:1 ratio, and perform data augmentation on the dataset;
[0007] Step 2: Improve the basic YOLOv8n model, train the improved YOLOv8n model using the training set, and adjust the parameters using the validation set;
[0008] Step 3: Use the test set to evaluate the detection effect of the trained YOLOv8n model. If the evaluation passes, the YOLOv8n model for target detection is obtained.
[0009] The improvements to the YOLOv8n model are as follows: the C2f module in the backbone is replaced with the C2f_DCNv4 module; the third and fourth C2f modules are replaced by the self-attention mechanism module COT in the neck part; and the detection head that comes with YOLOv8n is replaced by the LSDCH lightweight shared convolutional detection head in the head part.
[0010] Preferably, the C2f module in the backbone part is replaced with a C2f_DCNv4 module, specifically replacing the Bottleneck between Split and Concat in the C2f module with Bottleneck-DCNv4, where the Bottleneck-DCNv4 module is formed by stacking two DCNv4 modules.
[0011] Preferably, the structure of the self-attention mechanism module COT is as follows: the input feature X is used to derive the Key Map, Query and Value Map, wherein the Key Map is subjected to static context encoding by K:k×k convolution to capture the static association between keys; the Query is processed by Q:1×1 and then concatenated with the encoded KeyMap, and then the dynamic multi-head attention matrix is learned by δ:1×1 and combined with the content of the Value Map processed by V:1×1 to generate a dynamic context representation; finally, the static and dynamic contexts are fused through the Fusion module to output the feature Y.
[0012] Preferably, the LSDCH lightweight shared convolution detection head structure is as follows: the LSDCH detection head input is P3, P4 and P5 level features, each feature is first normalized by 1×1 Conv_GN to promote information exchange in the channel dimension; then the features are merged and passed through two 3×3 shared convolutions to further aggregate spatial and semantic information; the output of the shared convolution is sent to the classification and regression branches respectively, where the regression branch introduces the Scale layer to maintain multi-scale features through feature scaling, thereby improving the detection capability of targets of different scales.
[0013] Preferably, group normalization is used in the LSDCH lightweight shared convolution detection head instead of traditional batch normalization. The specific steps are as follows: Input feature map X∈R n×c×h×w , where n is the batch size, c is the number of channels, h and w are the spatial dimensions; c channels are divided into g groups, each containing channels; for each group, the mean is calculated over the spatial dimensions h×w and variance in is the total number of elements in the group; the calculated mean and variance are used to normalize the features within the group, that is, ε is a very small constant to prevent division by zero.
[0014] Preferably, the group normalization result is adjusted by the learnable scaling parameter γ and the offset parameter β to obtain the output
[0015] Compared with the prior art, the present invention has the following beneficial effects:
[0016] The backbone adopts the C2f-DCNv4 module. By removing the Softmax normalization in spatial aggregation, the convolution kernel can adaptively adjust the receptive field according to the target shape and posture, optimize the memory access mode, let the thread process the channels of shared sampling offset and aggregation weight, merge memory access instructions, improve computing efficiency, and have a lightweight structure to adapt to the deployment requirements of edge devices.
[0017] The Neck part introduces the self-attention mechanism module COT, which efficiently captures long-distance dependencies by fusing static and dynamic context information, optimizes multi-scale feature fusion, and enhances the detection ability of small targets and occluded targets. It adopts a lightweight design of 1×1 convolution, balances computational efficiency and performance, focuses on key features, improves robustness in complex scenarios, and collaborates with modules in the backbone part to form a complementary mechanism of local detail perception and global context reasoning.
[0018] The head part uses the LSDCH lightweight shared convolution detection head, which reduces redundant calculations through shared convolution, reduces the number of parameters and computational complexity, efficiently reuses features, avoids repeated extraction, and improves information utilization; combined with a multi-scale feature fusion structure, it is independent of batch size and suitable for small batch and dynamic resolution scenarios; the channel groups are normalized through group normalization, which has low computational complexity and is lightweight and friendly; cross-branch feature interaction is achieved through shared convolution, enhancing the global context perception and expression capabilities of the features, and the regression branch introduces the Scale layer to maintain multi-scale features through feature scaling. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0020] Figure 1 It is the main flow chart of the method of the present invention;
[0021] Figure 2 This is a schematic diagram of the network structure of the improved YOLOv8n model of the present invention;
[0022] Figure 3 Schematic diagram of the structure of the C2f_DCNv4 module of the present invention;
[0023] Figure 4 Schematic diagram of the structure of the Bottleneck_DCNv4 unit of the present invention;
[0024] Figure 5 It is a structural diagram of the self-attention mechanism module COT of the present invention;
[0025] Figure 6 It is a structural diagram of the LSDCH lightweight shared convolution detection head of the present invention. DETAILED DESCRIPTION
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0027] In order to make the above-mentioned objects, features and effects of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] Example 1: An autonomous driving target detection method based on improved YOLOv8, comprising the following steps:
[0029] Step 1: Collect autonomous driving scene images, covering different weather, lighting, road conditions, and target types (vehicles, pedestrians, traffic signs, etc.); use Labelimg software to annotate the targets, generate bounding boxes and category labels, and randomly divide the dataset into training set, validation set, and test set in an 8:1:1 ratio, and perform data augmentation on the dataset.
[0030] Step 2: Improve the basic YOLOv8n model, train the improved YOLOv8n model using the training set, and adjust the parameters using the validation set.
[0031] Step 3: Use the test set to evaluate the detection effect of the trained YOLOv8n model. After the evaluation passes, the YOLOv8n model for target detection is obtained.
[0032] The YOLOv8n model is a version of the YOLOv8 model. The improvements of the YOLOv8n model are as follows: the C2f module in the backbone is replaced by the C2f_DCNv4 module. The specific improvement is to replace the Bottleneck between Split and Concat in the C2f module with Bottleneck-DCNv4, and the Bottleneck-DCNv4 module constructed by the deformable convolutional network is composed of two DCNv4 modules stacked together; the Neck part uses the self-attention mechanism module COT to replace the third and fourth C2f modules; the Head part uses the LSDCH lightweight shared convolutional detection head to replace the detection head that comes with YOLOv8n.
[0033] like Figure 2 The figure shows the network structure of the improved YOLOv8n model. The input image is first preprocessed and enters the backbone part for feature extraction. Then it enters the neck part for feature fusion. After that, it enters the head part to generate the bounding box, class probability and confidence score. Finally, after post-processing, the autonomous driving vehicle target detection result is output.
[0034] like Figure 3 As shown in Figure 2, the C2f_DCNv4 module replaces all Bottleneck units in the C2f module with Bottleneck_DCNv4 units. Figure 4 As shown, the Bottleneck_DCNv4 unit is composed of two stacked DCNv4 modules.
[0035] The self-attention mechanism module COT is introduced into the Neck part of YOLOv8n and replaced with the third and fourth C2f modules in Neck. The structure of the self-attention mechanism module COT is as follows: Figure 5 As shown. COT efficiently captures long-distance dependencies and enhances feature expression by fusing static and dynamic context information. It uses a lightweight design of 1×1 convolution to reduce computational overhead and has strong adaptability. The dynamic and static combination mechanism enables it to improve the utilization of global information without significantly increasing the number of parameters, optimize multi-scale feature fusion, enhance the model's adaptability to complex scenarios, and maintain efficient reasoning, which is conducive to actual deployment. Figure 5In [1], the input X is used to derive the Key Map, Query, and Value Map. The Key Map is subjected to static context encoding by K:k×k convolution to capture the static association between keys. The Query is processed by Q:1×1 and then concatenated with the encoded key. Then, a dynamic multi-head attention matrix is learned by δ:1×1 (two consecutive 1×1 convolutions) and combined with the content of the Value Map processed by V:1×1 to generate a dynamic context representation. Finally, the static and dynamic contexts are fused through the Fusion module to output Y(H×W×C). This module enhances the visual representation capability through this combination of static and dynamic. Without significantly increasing the number of parameters and computational complexity, it effectively improves the model performance of tasks such as image classification and target detection, and achieves efficient use of global long-distance dependencies. Among them:
[0036] X is the input feature, with a size of H×W×C (H is height, W is width, and C is the number of channels);
[0037] Q: 1×1: Perform a 1×1 convolution on the branch used to generate the Query in the input X, adjust the number of channels to D, and output H×W×D;
[0038] K: 1×1: Perform k×k convolution on the Key Map branch (the specific parameters depend on the situation) to extract local features;
[0039] V:1×1: Perform 1×1 convolution on the Value Map branch to adjust the number of channels;
[0040] δ:1×1: consists of two consecutive 1×1 convolutions to generate dynamic weights;
[0041] Concat: concatenation operation, fusing features of different branches;
[0042] Fusion: Fusion module, integrating information to output the final feature Y.
[0043] The LSDCH lightweight shared convolution detection head structure used in the head part is as follows Figure 6As shown, the LSDCH detection head inputs features from the P3, P4, and P5 layers. Each feature is first normalized through a 1×1 Conv_GN (fused convolution and group normalization) to facilitate information exchange in the channel dimension and enrich the feature representation. These features are then merged and passed through two 3×3 shared convolutions to further aggregate spatial and semantic information, suppressing redundancy while expanding the receptive field and enhancing feature extraction capabilities. The output of the shared convolution is fed into the classification (Con_Cls) and regression (Con_Reg) branches, respectively. The regression branch introduces a Scale layer, which maintains multi-scale features through feature scaling, improving the detection capability of objects of different scales. The shared convolution mechanism in this structure reduces parameter redundancy, significantly reduces model complexity, achieves lightweightness, and improves parameter utilization. Furthermore, cross-branch feature interaction is achieved through shared convolution, alleviating the "information island" problem of independent branches and enhancing the global context perception and expression capabilities of features. Secondly, the Scale layer explicitly optimizes multi-scale features to enhance the balanced detection capabilities of large and small targets. Finally, group normalization (GN) replaces traditional batch normalization, is independent of batch size, and improves training stability and detection accuracy. It is particularly suitable for small batch or dynamic resolution scenarios, such as real-time inference on edge devices. These designs have enabled LSDCH to achieve effective breakthroughs in lightweightness, feature utilization efficiency, and multi-scale detection performance.
[0044] GN (Group Normalization) process: Input feature map X∈R n×c×h×w (n is the batch size, c is the number of channels, h and w are the spatial dimensions), divide the c channels into g groups, each group contains For each group, the mean is calculated over the spatial dimensions h×w ( is the total number of elements in the group) and variance The calculated mean and variance are used to normalize the features within the group, that is, (ε is a very small constant to prevent division by zero.) The normalized result is adjusted by the learnable parameters γ (scaling) and β (bias) to obtain the final output
[0045] Example 2:
[0046] The computer-readable storage medium of this embodiment stores a computer program thereon, which, when executed by a processor, implements the steps of an autonomous driving target detection method based on improved YOLOv8 in Example 1.
[0047] The computer-readable storage medium of this embodiment may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal; the computer-readable storage medium of this embodiment may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card, a secure digital card, a flash memory card, etc. equipped on the terminal; further, the computer-readable storage medium may also include both an internal storage unit of the terminal and an external storage device.
[0048] The computer-readable storage medium of this embodiment is used to store computer programs and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.
[0049] Example 3:
[0050] The computer device of this embodiment includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the autonomous driving target detection method based on improved YOLOv8 in Example 1 are implemented.
[0051] In this embodiment, the processor can be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The memory can include read-only memory and random access memory, and provide instructions and data to the processor. A part of the memory can also include non-volatile random access memory. For example, the memory can also store information about the device type.
[0052] Those skilled in the art will clearly understand that each embodiment can be implemented by means of software plus a necessary general-purpose hardware platform, or of course, by means of hardware. Based on this understanding, the essence of the above technical solution or the portion that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0053] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. An autonomous driving target detection method based on improved YOLOv8, characterized in that: The following steps are involved: Step 1: Collect autonomous driving scene images, annotate the objects, generate bounding boxes and category labels, randomly divide the dataset into training set, validation set, and test set in an 8:1:1 ratio, and perform data augmentation on the dataset; Step 2: Improve the basic YOLOv8n model, train the improved YOLOv8n model using the training set, and adjust the parameters using the validation set; Step 3: Use the test set to evaluate the detection effect of the trained YOLOv8n model. If the evaluation passes, the YOLOv8n model for target detection is obtained. The improvements to the YOLOv8n model are as follows: the C2f module in the backbone is replaced with the C2f_DCNv4 module; the third and fourth C2f modules are replaced by the self-attention mechanism module COT in the neck part; and the detection head that comes with YOLOv8n is replaced by the LSDCH lightweight shared convolutional detection head in the head part.
2. The automatic driving target detection method based on improved YOLOv8 according to claim 1, characterized in that: The C2f module in the backbone is replaced with the C2f_DCNv4 module. Specifically, the Bottleneck between Split and Concat in the C2f module is replaced with Bottleneck-DCNv4. The Bottleneck-DCNv4 module is composed of two stacked DCNv4 modules.
3. The automatic driving target detection method based on improved YOLOv8 according to claim 1, characterized in that: The structure of the self-attention mechanism module COT is as follows: the input feature X is derived from KeyMap, Query and Value Map, among which the Key Map is statically context-encoded by K:k×k convolution to capture the static association between keys; the Query is processed by Q:1×1 and then spliced with the encoded Key Map, and then the dynamic multi-head attention matrix is learned by δ:1×1 and combined with the content of the Value Map processed by V:1×1 to generate a dynamic context representation; finally, the static and dynamic context are fused through the Fusion module to output the feature Y.
4. The automatic driving target detection method based on improved YOLOv8 according to claim 1, characterized in that: The specific structure of the LSDCH lightweight shared convolutional detection head is as follows: the input of the LSDCH detection head is P3, P4 and P5 level features. Each feature is first normalized by a 1×1 Conv_GN to promote information exchange in the channel dimension; then the features are merged and passed through two 3×3 shared convolutions to further aggregate spatial and semantic information; the output of the shared convolution is sent to the classification and regression branches respectively, among which the regression branch introduces the Scale layer to maintain multi-scale features through feature scaling, thereby improving the detection capability of targets of different scales.
5. The automatic driving target detection method based on improved YOLOv8 according to claim 4, characterized in that: LSDCH lightweight shared convolution detection head uses group normalization instead of traditional batch normalization. The specific steps are: input feature map X∈R n×c×h×w , where n is the batch size, c is the number of channels, h and w are the spatial dimensions; c channels are divided into g groups, each containing channels; for each group, the mean is calculated over the spatial dimensions h×w and variance in is the total number of elements in the group; the calculated mean and variance are used to normalize the features within the group, that is, ε is a very small constant to prevent division by zero.
6. The automatic driving target detection method based on improved YOLOv8 according to claim 5, characterized in that: The group normalization result is adjusted by the learnable scaling parameter γ and offset parameter β to obtain the output 7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the autonomous driving target detection method based on improved YOLOv8 are implemented as described in any one of claims 1 to 6.
8. A computer device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the autonomous driving target detection method based on improved YOLOv8 are implemented as described in any one of claims 1 to 6.