Road target detection method based on improved YOLOv8
Through the improved YOLOv8 detection network model and point cloud data fusion technology, the problem of low target detection accuracy in complex road contexts is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510043816.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the accuracy of road target detection is low, especially in the context of complex roads, and it is difficult to effectively detect dense targets and occlusion conditions.
Using the improved YOLOv8 detection network model, iterative training is carried out through the stacked backbone network, Neck and lightweight detection head GSCD to improve feature extraction and detection accuracy, and combining the point cloud data object detection network and decision-level fusion module to obtain multi-dimensional information to improve detection accuracy.
It significantly improves the accuracy of target detection in complex road contexts, can handle dense targets and occlusion more effectively, and enhances the robustness of road target detection.
Smart Images

Figure CN120071077A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving, and relates to a road target detection method, specifically to a road target detection method based on improved YOLOv8, which can be used in complex driving scenarios under intelligent transportation conditions. Background Art
[0002] With its unique algorithms, computer vision has the ability to deeply understand the content of pictures and videos, and then extract valuable information from them to achieve the purpose of real-time monitoring of various targets in the surrounding environment. And the target detection ability with high robustness will undoubtedly build a solid defense line for travel safety.
[0003] As a key direction in the field of target detection, road target detection plays a crucial role in many important fields such as advanced driver assistance systems and intelligent driving systems. It can accurately judge whether road targets appear in the corresponding scenarios and clarify their relative positions, and then transmit these key information to the driver, thereby assisting the driver to make more reasonable and safe decisions, which has an important positive effect on reducing or even avoiding collision accidents between vehicles and road targets. The coverage of road target detection technology in practical applications is increasingly expanding, and it mainly relies on relevant algorithms based on deep learning to detect and identify road targets in the input images.
[0004] In the patent document with the application number CN202311085132.4, the application publication number CN117037119A, and the name "Road Target Detection Method and System Based on Improved YOLOv8", a road target detection method and system based on improved YOLOv8 are proposed. This method first preprocesses pictures using data augmentation methods to construct a dataset, and divides the training set, validation set, and test set proportionally. Then, an improved YOLOv8 network including a feature extraction module, a feature fusion module, and a detection module is constructed. In the feature extraction module, the original C2f module is replaced by the C2f_FNEMA module. The feature fusion module adds a detection layer on the basis of the original three-scale detection layer and uses the C2f_FN module to replace the original C2f module. The detection module uses the decoupled head Detect module. Subsequently, the training set is input into this network, different levels of features are generated through feature extraction, and then input into the feature fusion module to obtain enhanced features. Finally, the enhanced features are input into the detection module, and its branch decoupling is used to calculate the regression and classification losses and output the positions and categories of the detection boxes, improving the problem of low detection accuracy for dense target occlusion in complex road backgrounds. Due to the small size of road targets and the large scale variation of different types of road targets in this invention, the detection accuracy of this model is still relatively low. Summary of the Invention
[0005] The object of the present invention is to overcome the defects existing in the above-mentioned prior art, and propose a road target detection method based on improved YOLOv8 to solve the technical problem of low detection accuracy existing in the prior art.
[0006] To achieve the above object, the technical solution adopted by the present invention includes the following steps:
[0007] (1) Obtain a training sample set and a test sample set:
[0008] Obtain a training sample set including N RGB images and their labels, and a test sample set including P RGB images and P frame point cloud data;
[0009] (2) Construct an improved YOLOv8 network model:
[0010] Construct an improved YOLOv8 detection network model including a cascaded backbone network backbone C, Neck, and lightweight detection head GSCD O ; where backbone C includes stacked preliminary feature extraction modules, deformable convolution C2f_DCNv2_Dynamic modules, and spatial pyramid SPPF_CSPC modules; GSCD includes cascaded feature fusion modules, multiple parameter-sharing convolutional layers, and feature group normalization layers;
[0011] (3) Iteratively train the improved YOLOv8 network model:
[0012] Iteratively train the improved YOLOv8 network model O with the training sample set to obtain a trained improved YOLOv8 network model O * ;
[0013] (4) Construct a road target detection network model based on improved YOLOv8:
[0014] Construct a road target detection network model W including the trained improved YOLOv8 network model O obtained in step (3) * and a point cloud data target detection network parallel to it, and a decision-level fusion module cascaded with the output ends of these two networks; where the point cloud data target detection network includes sequentially stacked point cloud Filtering layer, Segmentation layer, and Euclidean Clustering layer;
[0015] (5) Obtain road target detection results:
[0016] Use the test sample set as the input of the road target detection network model W for forward propagation to obtain the road target detection results corresponding to the test sample set.
[0017] Compared with the prior art, the present invention has the following advantages:
[0018] 1. During the iterative training process of the improved YOLOv8 target detection network model of the present invention, the initial feature extraction module extracts features from each RGB image; the deformable convolution C2f_DCNv2_Dynamic performs dynamic convolution on the extracted shallow feature maps; the spatial pyramid SPPFCSPC performs pooling at different scales on the feature maps with target key features after dynamic convolution and splices them in the feature dimension to obtain feature maps with more diverse and rich feature information. The lightweight detection head GSCD detects the feature maps output by the Neck, effectively improving the accuracy of target detection in complex road backgrounds.
[0019] 2. The road target detection network model of the present invention can obtain multi-dimensional information of road targets through the trained improved YOLOv8 target detection network model, the point cloud data target detection network, and the decision-level fusion module, further improving the accuracy of road target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flowchart for the implementation of the present invention.
[0021] Figure 2 is a schematic structural diagram of the backbone C in the improved YOLOv8 of the present invention.
[0022] Figure 3 is a schematic structural diagram of the road target detection network model of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0023] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0024] Refer to Figure 1 , the present invention includes the following steps:
[0025] Step 1) Obtain a training sample set and a test sample set:
[0026] Acquire M two-dimensional RGB images including multiple road targets through a camera. Each image frame has a resolution of 512×512 and 3 channels. After annotating the targets in N of the RGB images, form a training sample set with the N RGB images and their labels. Synchronize the time and space of the remaining P = M - N two-dimensional RGB images and the P frames of three-dimensional point cloud data obtained by the lidar. Among them, use the message_fliter package in ROS to match the image data with a time stamp close to the current point cloud data to complete the time synchronization of the RGB image and the point cloud data. Use the camera_calibration tool in ROS and the calibration toolbox in Autoware to obtain the internal and external parameter matrices of the camera and project the point cloud onto the image to achieve spatial synchronization. In this embodiment, M = 1000 and N = 800.
[0027] Step 2) Construct an improved YOLOv8 object detection network model:
[0028] Construct an improved YOLOv8 detection network model including a cascaded backbone network backbone C, Neck, and lightweight detection head GSCD O ; where the structure of backbone C is as Figure 3 , including a stacked preliminary feature extraction module, a deformable convolution C2f_DCNv2_Dynamic module, and a spatial pyramid SPPFCSPC module; GSCD includes a cascaded feature fusion module, multiple parameter-sharing convolutional layers, and a feature group normalization layer, where:
[0029] Backbone network C, where the preliminary feature extraction module consists of stacked first Conv module, second Conv module, first CSP module, third Conv module, second CSP module, fourth Conv module, third CSP module, and fifth Conv module; the deformable convolution C2f_DCNv2_Dynamic module consists of stacked one convolutional layer, one deformable convolutional layer, one modulated deformable convolutional layer, and one MPCA layer; the spatial pyramid SPPFCSPC module includes a first branch composed of multiple convolutional layers and pooling layers arranged in parallel and cascaded, a second branch composed of multiple convolutional layers, and a feature splicing layer cascaded with the two branches; the preliminary feature extraction module in backbone C extracts shallow features containing initial information of targets at different scales through multiple convolutional layers, residual block CSP, and SiLU activation function layer, which is beneficial for the model to capture small target details and initial information of targets at different scales; the deformable convolution module overcomes the limitations of a convolutional kernel with a fixed shape when dealing with different road targets with irregular shapes and large scale changes by adjusting the sampling position and weight of the convolutional kernel, enabling the model to accurately extract features of road targets with variable shapes and scales; the spatial pyramid SPPFCSPC module effectively fuses target information at different scales through multi-scale convolution, pooling, and feature splicing, improving the accuracy of the model for detecting road targets at different scales and overcoming the shortcomings of insufficient feature extraction for small targets and inability to obtain sufficient global information for large targets.
[0030] Neck upsamples the deep feature map output by the backbone network, splices the obtained feature map with an increased scale with the feature map output by the third CSP module in the channel dimension, performs convolution on the spliced feature map to obtain an intermediate layer feature map after feature fusion, downsamples the intermediate layer feature map, splices the feature map with a reduced scale with the feature map output by the second CSP module in the channel dimension, performs convolution on the spliced intermediate layer feature map to obtain a shallow layer feature map after feature fusion, downsamples the shallow layer feature map, splices the feature map with a reduced scale with the intermediate layer feature map in the channel dimension, performs convolution on the spliced feature map, downsamples the convolved feature map, splices it with the deep feature map in the feature dimension, and performs convolution on the spliced feature map to obtain a deep fusion feature map; Neck fuses multi-scale information by splicing feature maps at different scales, improving the feature extraction ability of the model for multi-scale target features and overcoming the shortcoming of difficult feature extraction caused by large scale changes of road targets.
[0031] The GSCD includes a feature fusion module, a first convolutional layer, a first feature group normalization layer, a second convolutional layer, a second feature group normalization layer, a first classification convolutional layer, a first bounding box regression convolutional layer, a second classification convolutional layer, a second bounding box regression convolutional layer, a third classification convolutional layer, and a third bounding box regression convolutional layer; among them, the feature fusion module includes three cascaded branches, each branch contains a convolutional layer and a feature group normalization layer; the output of the feature fusion module is stacked with the first convolutional layer, the first feature group normalization layer, the second convolutional layer, and the second feature group normalization layer in sequence, and the output end of the second feature group normalization layer is cascaded with the input ends of the first classification convolutional layer, the first bounding box regression convolutional layer, the second classification convolutional layer, the second bounding box regression convolutional layer, the third classification convolutional layer, and the third bounding box regression convolutional layer; by fusing different feature information, the GSCD enables the model to obtain target features of different scales, overcomes the shortcoming of being difficult to detect multi-scale targets, and improves the detection accuracy of road targets of different scales.
[0032] Step 3) Iteratively train the improved YOLOv8 network model:
[0033] Iteratively train the improved YOLOv8 object detection network model O with the training sample set to obtain the trained improved YOLOv8 object detection network model O * :
[0034] (3a) Initialize the iteration number as t and the maximum iteration number as T, where T ≥ 300. The weight and bias parameters of the improved YOLOv8 network in the t-th iteration are ω t , θ t , and let t = 1;
[0035] (3b) The improved YOLOv8 object detection network detects each RGB image in the training sample set to obtain the detection result of the two-dimensional detection box of the road target in the RGB image;
[0036] (3c) Use the cross-entropy loss function to calculate the loss value of the improved YOLOv8 network through the detection result of the target two-dimensional detection box obtained in step (3b) and its corresponding label And use the stochastic gradient descent method to update the weight parameter ω t and the weight parameter θ t to obtain the improved YOLOv8 network model O of this iteration t :
[0037] (3c1) Define the loss value of the improved YOLOv8 The calculation formula is:
[0038]
[0039] Among them, The classification represents the classification loss and the bounding box regression loss in the detection result, λ cls and λ reg are hyperparameters, ∑ represents the summation operation, and x n , respectively represent the actual label of the nth RGB image and the classification prediction value in its corresponding detection result, and y n and respectively represent the coordinates of the true bounding box of the nth RGB image and the coordinates of the predicted bounding box in the detection result;
[0040] (3c) Update the weight parameter ω t , the bias parameter θ t , and the update formulas are respectively:
[0041]
[0042]
[0043] Among them, ω t-1 , θ t-1 respectively represent the weights and bias parameters of the improved YOLOv8 at the (t - 1)-th iteration, α represents the learning rate, respectively represent taking the partial derivatives with respect to ω t-1 and θ t-1 ;
[0044] (3d) Judge whether t = T holds. If so, obtain the trained improved YOLOv8 network model O * , otherwise, let t = t + 1, O t = O, and execute step (3b);
[0045] Step 4) Construct a road target detection network model based on the improved YOLOv8:
[0046] Construct a road target detection network model W that includes the trained improved YOLOv8 target detection network model O * obtained in step (3) and a point cloud data target detection network parallel to it, as well as a decision-level fusion module cascaded with the output ends of these two networks. Its structure is as Figure 3 shown; among them, the point cloud data target detection network includes a point cloud Filtering layer, a Segmentation layer, and an Euclidean Clustering layer stacked in sequence:
[0047] (4a) The point cloud Filtering layer processes each frame of point cloud data using a voxelization method to obtain point cloud data with reduced density;
[0048] (4b) The Segmentation layer uses a seed point fitting algorithm to separate the road point cloud below the height threshold from the target point cloud on the road, thereby extracting the road target point cloud data;
[0049] (4c) The Euclidean Clustering layer divides the point clouds with close distances in the road target point cloud data into clusters, and obtains the three-dimensional detection box of each point cloud cluster by calculating the coordinates of the eight extreme points of the cluster;
[0050] Step 5) Obtain the road target detection result:
[0051] Use the test sample set as the input of the road target detection network model W for forward propagation to obtain the road target detection result corresponding to the test sample set:
[0052] (5a) The improved YOLOv8 network and the point cloud data target detection network respectively detect each RGB image and each frame of point cloud data in the test sample set to obtain the road target two-dimensional detection box detection result of the RGB image and the road target three-dimensional detection box detection result of the point cloud data;
[0053] (5b) The decision-level fusion module projects the three-dimensional detection box of the road target detected by the point cloud data target detection network onto the image according to the intrinsic matrix and the extrinsic matrix, calculates the ratio of the intersection area to the union area of the projected road target two-dimensional detection box and the road target two-dimensional detection box detected by the improved YOLOv8 network. If the ratio is higher than the threshold, the results of the two are fused, further improving the accuracy of road target detection.
Claims
1. A road target detection method based on improved YOLOv8, characterized in that: The following steps are involved: (1) Obtain training sample set and test sample set: Obtain a training sample set including N RGB images and their labels, and a test sample set including P RGB images and P frames of point cloud data; (2) Build an improved YOLOv8 network model: Construct an improved YOLOv8 detection network model including the cascaded backbone network backbone C, Neck and the lightweight detection head GSCD O ; The backbone C includes a stacked preliminary feature extraction module, a deformable convolution C2f_DCNv2_Dynamic module, and a spatial pyramid SPPFCSPC module; GSCD includes a cascaded feature fusion module, multiple parameter sharing convolution layers, and a feature group normalization layer; (3) Iterative training of the improved YOLOv8 network model: The improved YOLOv8 network model O is iteratively trained through the training sample set to obtain the trained improved YOLOv8 network model O * ; (4) Build a road object detection network model based on improved YOLOv8: Construct the trained improved YOLOv8 network model O obtained in step (3) * and a point cloud data target detection network parallel to it, and a road target detection network model W of a decision-level fusion module cascaded with the output ends of the two networks; wherein the point cloud data target detection network includes a point cloud filtering layer, a segmentation layer and a Euclidean Clustering layer stacked in sequence; (5) Obtaining road target detection results: The test sample set is used as the input of the road object detection network model W for forward propagation to obtain the road object detection results corresponding to the test sample set.
2. The method according to claim 1, characterized in that The steps for obtaining the training sample set and the test sample set in step (1) are as follows: (1a) Obtain M two-dimensional RGB images including multiple road targets, and after annotating the targets in N of the RGB images, the N RGB images and their labels form a training sample set, where M ≥ 1000. (1b) Preprocess the remaining P = MN two-dimensional RGB images and P frames of three-dimensional point cloud data, and form a test sample set with the preprocessed P two-dimensional RGB images and P frames of three-dimensional point cloud data that are synchronized in time and space.
3. The method according to claim 1, characterized in that The improved YOLOv8 network model described in step (2), wherein: Backbone network backbone C, in which the preliminary feature extraction module includes stacked multiple convolutional layers, residual block CSP and SiLU activation function layer; the deformable convolution C2f_DCNv2_Dynamic module includes stacked convolutional layers, deformable convolutional layers, MPCA attention layers and modulated deformable convolutional layers; the spatial pyramid SPPFCSPC module includes a first branch consisting of multiple convolutional layers and pooling layers arranged in parallel and cascaded, and a second branch including multiple convolutional layers, and a feature splicing layer cascaded with the two branches; The lightweight detection head GSCD includes a cascaded feature fusion module, multiple parameter-sharing convolutional layers and a feature group normalization layer, wherein the feature fusion module is composed of parallel arranged convolutional layers and feature group normalization layers.
4. The method according to claim 1, characterized in that The iterative training of the improved YOLOv8 network model described in step (3) is implemented as follows: (3a) The number of initialization iterations is t, the maximum number of iterations is T, T ≥ 300, and the weights and bias parameters of the improved YOLOv8 network in the tth iteration are ω respectively. t ,θ t , and let t = 1; (3b) The improved YOLOv8 network detects each RGB image in the training sample set and obtains the two-dimensional detection box detection result of the road target in the RGB image; (3c) Using the cross entropy loss function, the loss value of the improved YOLOv8 network is calculated using the target two-dimensional detection box detection results and their corresponding labels obtained in step (3b). And using stochastic gradient descent, For the weight parameter ω t and weight parameter θ t Update to get the improved YOLOv8 network model O for this iteration t ; (3d) Determine whether t=T is true. If so, obtain the trained improved YOLOv8 network model O * , otherwise, let t = t + 1, O t =O, and execute step (3b).
5. The method according to claim 4, characterized in that The loss value of the modified YOLOv8 described in step (3c) The calculation formula is: in, Classification represents the classification loss and bounding box regression loss in the detection results, λ cls and λ reg is a hyperparameter, ∑ represents the summation operation, x n , Represents the actual label of the nth RGB image and its corresponding classification prediction value in the detection result, y n and They represent the coordinates of the true bounding box of the nth RGB image and the coordinates of the predicted bounding box in the detection result, respectively.
6. The method according to claim 4, characterized in that The weight parameter ω described in step (3c) t , bias parameter θ t To update, the update formulas are: Among them, ω t-1 ,θ t-1 They represent the weight and bias parameters of the improved YOLOv8 at the t-1th iteration, α represents the learning rate, Respectively represent the t-1 and θ t-1 Find the partial derivative.
7. The method according to claim 4, characterized in that The improved YOLOv8 network described in step (3b) detects each RGB image in the training sample set, and the implementation steps are as follows: (3b1) The preliminary feature extraction module in the backbone network backbone C extracts features from each RGB image; the deformable convolution C2f_DCNv2_Dynamic performs dynamic convolution on the extracted shallow feature map; the spatial pyramid SPPFCSPC performs pooling of different scales on the feature map Y with target key features after dynamic convolution and splices it in the feature dimension to obtain a feature map Z with richer and more diverse feature information; (3b2) Neck upsamples the feature map Z multiple times, and splices the multiple upsampled feature maps, then downsamples the spliced feature map P3 twice and splices them to obtain the spliced feature maps P4 and Z; (3b3) The feature fusion module in the lightweight detection head GSCD performs convolution and feature group normalization on the feature maps P3, P4 and Z to obtain a feature map H with preliminary feature fusion. The parallel arranged classification convolution and prediction box convolution convolute the feature map H to obtain the target detection result.
8. The method according to claim 1, characterized in that The steps for obtaining the road target detection result described in step (5) are as follows: (5a) The improved YOLOv8 network and the point cloud data target detection network detect each RGB image and each frame of point cloud data in the test sample set, respectively, to obtain the two-dimensional detection frame detection result of the road target of the RGB image and the three-dimensional detection frame detection result of the road target of the point cloud data; (5b) The decision-level fusion module fuses the two-dimensional detection frame detection result of the road target with the three-dimensional detection frame detection result of the road target to obtain the road target detection result.
9. The method according to claim 8, characterized in that The point cloud data target detection network described in step (5a) detects each frame of point cloud data in the test sample set, and the implementation steps are as follows: (5a1) The point cloud filtering layer uses voxelization method to process each frame of point cloud data to obtain point cloud data with reduced density; (5a2) The Segmentation layer uses a seed point fitting algorithm to separate the road point cloud below the height threshold from the target point cloud on the road, thereby extracting the road target point cloud data; (5a3) The Euclidean Clustering layer divides the point clouds with close spacing in the road target point cloud data into a cluster, and obtains the three-dimensional detection box of each point cloud cluster by calculating the coordinates of the eight extreme points of the cluster.
Citation Information
Patent Citations
Road target detection method and system based on improved YOLOv8
CN117037119A