Vehicle and pedestrian detection method based on improved YOLOv8n

By improving the YOLOv8n model, the DilatedReparamBlock, TripletAttention and SlimNeck layers were introduced, and the DySample operator was used to solve the problems of mis-detection, missed detection and insufficient model lightweight in vehicle pedestrian detection, achieving more efficient detection accuracy and speed.

CN119992590APending Publication Date: 2025-05-13BEIJING TECH & BUSINESS UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510061821.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In a target-intensive detection environment, vehicle and pedestrian detection is prone to error detection and missed detection, and the existing models are not lightweight enough, resulting in insufficient detection accuracy and speed.

Method used

Improve the YOLOv8n model, introduce the DWR module modified by DilatedReparamBlock in the Backbone part, add the TripletAttention attention mechanism, replace the Neck layer with the SlimNeck layer, and use the lightweight dynamic upsampling operator DySample to improve the upsampling operation.

Benefits of technology

The accuracy, recognition speed and detection effect of dense small targets are improved, while the number of parameters and calculations of the model is reduced, achieving more efficient feature extraction and detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992590A_ABST
    Figure CN119992590A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle and pedestrian target detection method based on YOLOv8n, relates to the field of automatic driving and computer vision, and is used for solving the problems of wrong detection, missing detection and the like of shielded objects and small target objects in vehicle and pedestrian detection. Firstly, a C2fDWRDRB module and a heavy parameterization large convolution kernel layer are introduced to improve the performance without increasing the calculation cost and improve the detection speed and precision of the model, secondly, a Triplet Attention mechanism is added to improve the performance of the model under a lightweight architecture, in order to further reduce the complexity of the model and maintain the detection precision, GSConv and VoV-GSCSP modules are combined in research, and the detection efficiency of the model is improved. And a Slim-Neck method is introduced, so that more efficient and more accurate vehicle and pedestrian key point detection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a vehicle and pedestrian target detection method based on improved YOLOv8. Background Art

[0002] With the advancement of technology and the continuous increase in the number of cars, traffic problems are becoming increasingly serious. Autonomous driving technology is expected to effectively alleviate this dilemma. In autonomous driving systems, vehicle and pedestrian detection is a basic link, so it is particularly important to improve vehicle and pedestrian detection technology. Object detection is a basic task in the field of computer vision. Its main purpose is to locate and identify one or more target objects in an image or video.

[0003] In recent years, the main trend of autonomous driving vehicles is to apply deep learning-based object detection algorithms to pedestrian and vehicle detection. Such algorithms are divided into two categories: one is the One-Stage method based on regression detection, represented by the YOLO series of algorithms, SSD algorithm and RetinaNet algorithm; the other is the Two-Stage method based on candidate region generation, represented by R-CNN, FasterR-CNN and other algorithms.

[0004] In order for self-driving cars to operate safely and stably on the road, accurate detection, identification and judgment of real-time targets are the basis and core to ensure their operation.

[0005] At present, in the detection environment with dense targets, there are problems in vehicle and pedestrian detection, such as occluded objects and small targets are prone to false detection, missed detection, and the model is not lightweight enough. In previous studies, researchers mainly improved the attention mechanism, loss function, and Neck layer. This paper studies the above problems and improves the YOLOV8n model. Summary of the invention

[0006] Based on the summary and analysis of relevant research results at home and abroad, this paper proposes an improved network structure based on the YOLOv8n model to address the problems and challenges in existing vehicle detection. The structure is used to solve the problems of false detection, missed detection, and insufficient model weight in vehicle and pedestrian detection in a target-dense detection environment.

[0007] A vehicle and pedestrian target detection method based on an improved YOLOv8n network is provided, wherein a vehicle and pedestrian target detection model based on the improved YOLOv8n network is used to detect vehicle and pedestrian targets. The vehicle and pedestrian target detection model based on the improved YOLOv8n network is obtained in advance, and the following steps are included:

[0008] Step 1: Obtain the vehicle and pedestrian detection dataset and preprocess the vehicle and pedestrian dataset using data enhancement technology.

[0009] Step 2: Build a vehicle and pedestrian target detection model based on the improved YOLOv8n network; for the Backbone part of YOLOv8n, introduce the DWR module modified by DilatedReparamBlock in C2f in the Backbone part, which can re-parameterize the large convolution kernel layer, improving performance without increasing cost;

[0010] Step 3: Add a lightweight and efficient TripletAttention mechanism to solve the problem of decreased detection accuracy due to network lightweighting;

[0011] Step 4: For the Neck part of YOLOv8n, replace the Neck layer of the original model with the SlimNeck layer to reduce the number of network parameters and calculations to improve the speed and efficiency of the model;

[0012] Step 5: Use the lightweight dynamic upsampling operator DySample to improve the SlimNeck upsampling operation to better extract local feature information and improve the model's ability to perceive details.

[0013] Beneficial effects of this invention:

[0014] A vehicle target detection method based on YOLOv8n proposed in this invention improves the detection accuracy, recognition speed and detection effect of dense small targets by replacing part of the convolution in the backbone structure, introducing the attention mechanism and improving the feature extraction module. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is the improved YOLOv8n network structure diagram;

[0016] Figure 2 This is the C2f_DWR_DRB module structure diagram;

[0017] Figure 3 This is the GSConv module structure diagram;

[0018] Figure 4 This is the schematic diagram of the TripletAttention module; DETAILED DESCRIPTION

[0019] The accompanying drawings are only used as examples to assist in the description and should not be considered as any limitation or restriction to the content of this patent.

[0020] In the technical field of this invention, some contents of some drawings are omitted for the sake of simplicity and clarity of technical description, which should be understood.

[0021] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0022] The present invention proposes a vehicle target detection method based on YOLOv8n, which includes the following steps:

[0023] Step 1: Select and download the data set required for target detection, expand and divide the data set according to the needs, and enhance the image to obtain a data set suitable for the research object;

[0024] Step 2: Build a vehicle and pedestrian target detection model based on the improved YOLOv8n network; for the Backbone part of YOLOv8n, introduce the DWR module modified by DilatedReparamBlock in C2f in the Backbone part, which can re-parameterize the large convolution kernel layer, improving performance without increasing cost;

[0025] Step 3: Add a lightweight and efficient TripletAttention mechanism to solve the problem of decreased detection accuracy due to network lightweighting;

[0026] Step 4: For the Neck part of YOLOv8n, replace the Neck layer of the original model with the slimneck layer to reduce the number of network parameters and calculations to improve the speed and efficiency of the model;

[0027] Step 5: Use the lightweight dynamic upsampling operator DySample to improve the SlimNeck upsampling operation to better extract local feature information and improve the model's ability to perceive details.

[0028] Furthermore, the step 1 is specifically as follows:

[0029] To verify the effectiveness of the improved model, the experiment uses the KITTI dataset, which is collected on complex roads such as rural areas, highways, and urban areas;

[0030] In the data preprocessing stage, the study first eliminated categories with less data, merged Car, Van, Truck, and Tram into the Car class, merged the Pedestrian class and the Person class into the Person class, kept the original Cyclist class unchanged, and deleted the Dontcare and Misc classes. The merged data set was screened to remove images without labels or with a single target type or a small number of targets. In the processed data set, about 7,500 images were selected for the experiment, and the data was divided into training set, validation set, and test set in a ratio of 8:1:1.

[0031] Furthermore, the step 2 is specifically as follows:

[0032] This module is an improved C2f module after secondary innovation of the Dilation-wiseResidual (DWR) module in DWRSeg using the DilatedReparam-Block in UniRepLKNet. The DilatedReparamBlock module can enhance the non-large convolution kernel layer and consists of a small kernel convolution layer and multiple expanded small kernel convolution layers.

[0033] Its main advantage is that by reparameterizing the large convolution kernel layer, it can effectively reduce the model parameter calculation and model depth, and improve the detection accuracy and running speed. Therefore, DilatedReparamBlock is used to replace the DWR module in the C2f_DWR module, and finally the C2f_DWR_DRB module is obtained.

[0034] C2f_DWR_DRB uses a two-step residual feature extraction method to extract features, placing convolution operations and normalization at the end. This method can greatly improve the model's ability to perceive features and improve the model's nonlinear expression capabilities.

[0035] Furthermore, the step 3 is specifically as follows:

[0036] TripletAttention consists of three parallel branches, two of which are responsible for capturing the cross-dimensional interaction between channel C and space H or W, and the last branch is similar to CBAM and is used to construct SpatialAttention; the outputs of the final three branches are aggregated using average.

[0037] The attention mechanism adds the new features formed by the three lines to obtain the average value, so that the attention mechanism can better grasp the information exchange between space and channels, thereby improving the detection accuracy of the model.

[0038] Furthermore, the step 4 is specifically as follows:

[0039] The Conv convolution module in the neck part of YOLOv8n is replaced with a GSConv convolution module, which uses grouped convolution and depth-separable convolution to reduce the computational cost of the model to achieve lightweight convolution extraction.

[0040] The GSConv module reduces the amount of computation while retaining feature information as much as possible with a lower time complexity. GSConv first performs standard convolution processing on the input features with half the number of channels, and then performs a depth-wise separable convolution on the results. The two features are then concatenated and shuffled, so that local feature information is evenly exchanged on different channels, enhancing nonlinear expression capabilities.

[0041] VoV-GSCSP is a cross-level partial network module using a one-time aggregation method, with the GSBottleneck module composed of GSConv and Conv as the basic unit. The cross-level partial network (GSCSP) module VoV-GSCSP is designed using a one-time aggregation method. The VoV-GSCSP module reduces the complexity of calculation and network structure, but maintains sufficient accuracy. It is a lightweight module with high accuracy.

[0042] The C2f module is replaced with the VoV-GSCSP module to simplify the network structure and reduce the amount of model calculation. The Slim-Neck method is introduced through the GSConv module and the VoV-GSCSP module to reduce the complexity of the model while maintaining accuracy. The key points of vehicles and pedestrians are obtained more efficiently and accurately.

[0043] Furthermore, the step 5 is specifically as follows:

[0044] Dysample is an upsampling method based on point sampling, which avoids dynamic convolution and can improve the feature fusion ability of the network;

[0045] Dysample is an ultra-lightweight dynamic upsampling operator. It uses a dynamic sampling mechanism to generate upsampled feature maps based on the local information of the input feature map. Dysample adopts a self-attention mechanism to enhance the interdependence between different channels in the feature map, thereby better extracting local feature information.

[0046] For the input feature X of size C×H×W, a sampling point generator is used to generate a sampling set S of size 2×H2×W2. The input feature X is sampled bilinearly to obtain an output X′ of size C×H2×W2. The whole process can be defined as:

[0047] X′=grid_sample(X,S)

[0048] The above embodiments are only used to explain the present invention and are intended to help readers understand the principles of the present invention rather than to limit the present invention. Readers may make various modifications or improvements to the present invention based on relevant knowledge in this field. Without violating the spirit and principles of the present invention, all modifications should be included in the scope of the claims of the present invention.

Claims

1. A vehicle and pedestrian target detection method based on YOLOv8n in a target-intensive detection environment, the method comprising the following steps: Step 1: Obtain the vehicle and pedestrian detection dataset and preprocess the vehicle and pedestrian dataset using data enhancement technology. Step 2: Build a vehicle and pedestrian target detection model based on the improved YOLOv8n network. For the Backbone part of YOLOv8n, introduce the DWR module modified by DilatedReparamBlock in C2f in the Backbone part to reparameterize the large convolution kernel layer, which improves the performance without increasing the cost. Step 3: Add a lightweight and efficient TripletAttention mechanism to solve the problem of decreased detection accuracy due to network lightweighting; Step 4: For the Neck part of YOLOv8n, replace the Neck layer of the original model with the SlimNeck layer to reduce the number of network parameters and calculations to improve the speed and efficiency of the model; Step 5: Use the lightweight dynamic upsampling operator DySample to improve the SlimNeck upsampling operation to better extract local feature information and improve the model's ability to perceive details.

2. A vehicle target detection method of YOLOv8n according to claim 1, characterized in that, The specific step 1 is: (1) Select the required images from the KITTI public dataset and organize them into a new dataset; (2) Categories with less data were removed, Car, Van, Truck, and Tram were merged into the Car category, Pedestrian and Person categories were merged into the Person category, the original Cyclist category remained unchanged, and Dontcare and Misc categories were deleted. The merged dataset was screened to remove images without labels or with a single target type or a small number of targets, to obtain a complex dataset suitable for research; (3) In the processed data set, about 7,500 images were selected for the experiment, and the data were divided into training set, validation set and test set in a ratio of 8:1:

1.

3. A vehicle target detection method of YOLOv8n according to claim 1, characterized in that, The specific step 2 is: (1) Improve the backbone network structure of YOLOv8n and introduce the DWR module modified by DilatedReparamBlock in C2f in the Backbone layer to reparameterize the large convolution kernel layer, thereby improving performance without increasing cost. (2) Use DilatedReparamBlock to replace the DWR module in the C2f_DWR module, and finally obtain the C2f_DWR_DRB module; (3) C2f_DWR_DRB uses a two-step residual feature extraction method to extract features, placing convolution operations and normalization at the end. This method can greatly improve the model's ability to perceive features and improve the model's nonlinear expression capabilities.

4. A vehicle target detection method of YOLOv8n according to claim 1, characterized in that, The specific step 3 is: (1) TripletAttention consists of three parallel branches, two of which are responsible for capturing the cross-dimensional interaction between channel C and space H or W. The last branch is similar to CBAM and is used to construct SpatialAttention. Finally, the outputs of the three branches are aggregated using average; (2) The attention mechanism adds the new features formed by the three lines to obtain the average value, so that the attention mechanism can better grasp the information exchange between space and channels, thereby improving the detection accuracy of the model.

5. A vehicle target detection method of YOLOv8n according to claim 1, characterized in that, The specific step 4 is: (1) The Conv convolution module in the neck part of YOLOv8 is replaced with the GSConv convolution module, and the grouped convolution and depth-separable convolution are used to reduce the computational cost of the model to achieve lightweight convolution extraction; (2) GSConv first performs standard convolution processing on the input features with half the number of channels, and then performs depth-wise separable convolution on the results. The two features are then concatenated and shuffled, so that local feature information is evenly exchanged on different channels, enhancing nonlinear expression capabilities. The GSConv module retains feature information as much as possible with a lower time complexity while reducing the amount of computation; (3) The C2f module is replaced by the VoV-GSCSP module to simplify the network structure and reduce the amount of model calculation. The Slim-Neck method is introduced through the GSConv module and the VoV-GSCSP module to reduce the complexity of the model while maintaining accuracy, and to obtain the key points of vehicles and pedestrians more efficiently and accurately.

6. A vehicle target detection method of YOLOv8n according to claim 1, characterized in that, The specific step 5 is: (1) Dysample is an ultra-lightweight dynamic upsampling operator. It uses a dynamic sampling mechanism to generate an upsampled feature map based on the local information of the input feature map. Dysample adopts a self-attention mechanism to enhance the interdependence between different channels in the feature map, thereby better extracting local feature information. (2) Dysample is an upsampling method based on point sampling, which avoids dynamic convolution and can improve the feature fusion capability of the network.

Citation Information

Cited By

  • Pedestrian behavior prediction method for autonomous vehicle

    CN121583001A