Vehicle target detection method and system based on multi-scale feature extraction

By enhancing the YOLOv8 model with MPAB and MSAFDN modules and optimizing the CIOU loss function, the method addresses low accuracy and high miss rates in urban vehicle detection, achieving improved detection performance.

CN120318487APending Publication Date: 2025-07-15ANHUI UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510296526.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing vehicle target detection methods have low detection accuracy and high missed detection rates in complex urban environments, making it difficult to meet the needs of intelligent transportation and autonomous driving.

Method used

The Backbone layer and Head layer of the YOLOv8 model are improved, namely the MPAB module and the MSAFDN module respectively. The loss function of the YOLOv8 model is optimized by combining multi-scale feature extraction, cross-layer connection, channel attention mechanism and improved CIOU loss function.

Benefits of technology

It significantly improves the accuracy and robustness of vehicle target detection, reduces missed detection rates, and enhances the applicability of the model in complex backgrounds and multi-objective scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318487A_ABST
    Figure CN120318487A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target detection, in particular to a vehicle target detection method and system based on multi-scale feature extraction, and the method comprises the steps: firstly, respectively improving the C2f modules of the sixth layer and the eighth layer in the Backbone layer of a YOLOv8 model into MPAB modules; secondly, a Deect module in a Head layer of the YOLOv8 model is improved into an MSAFDN module; thirdly, further optimizing a loss function of the YOLOv8 model by improving a CIOU loss function; and finally, detecting a vehicle target by using the improved YOLOv8 model to obtain a detection result. According to the method, the YOLOv8 model is improved, so that the accuracy of vehicle target detection is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and specifically provides a vehicle object detection method and system based on multi-scale feature extraction. Background Art

[0002] With the rapid development of fields such as intelligent transportation, autonomous driving, and intelligent monitoring, the importance of vehicle object detection methods in practical applications has become increasingly prominent. By accurately detecting the position and category of vehicles in images or videos, various intelligent applications such as traffic flow management, road anomaly detection, and parking management can be realized. Current vehicle object detection methods need to maintain high efficiency in various complex scenarios, such as vehicle occlusion, lighting changes, and dense traffic environments, which pose higher requirements for the robustness and accuracy of detection models.

[0003] With the continuous progress of deep learning and computer vision technologies, many application scenarios have put forward higher requirements for the accuracy and low miss-detection rate of vehicle detection. Therefore, urban vehicle object detection has become a hot research field in academia. Currently, object detection algorithms based on deep learning can be mainly divided into two categories: two-stage detection methods (such as RCNN, etc.) and one-stage detection methods (such as SSD and YOLO, etc.).

[0004] Although the above methods have improved the performance of vehicle detection models to a certain extent, due to the complexity and diversity of vehicles in urban environments, the detection accuracy is often low and the miss-detection rate is high.

[0005] Therefore, a vehicle object detection method and system based on multi-scale feature extraction are proposed. Summary of the Invention

[0006] The purpose of the present invention is to provide a vehicle object detection method and system based on multi-scale feature extraction. The present invention relates to the technical field of object detection, and specifically provides a vehicle object detection method and system based on multi-scale feature extraction, including: First, the 6th and 8th C2f modules in the Backbone layer of the YOLOv8 model are respectively improved to MPAB modules; Second, the Detect module in the Head layer of the YOLOv8 model is improved to MSAFDN module; Then, by improving the CIOU loss function, the loss function of the YOLOv8 model is further optimized; Finally, the improved YOLOv8 model is used to detect vehicle objects to obtain detection results. The present invention effectively improves the accuracy of vehicle object detection by improving the YOLOv8 model.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A vehicle object detection method based on multi-scale feature extraction, including:

[0009] S1. Improve the first C2f module and the second C2f module of the Backbone layer of the YOLOv8 model into MPAB modules. The first C2f module and the second C2f module are the 6th layer C2f module and the 8th layer C2f module of the Backbone respectively. The MPAB module is used to perform multi-scale segmentation, cross-layer connection, feature fusion, channel attention mechanism, and feature weighting operations on the input image to obtain a third comprehensive feature.

[0010] S2. Improve the Detect module of the Head layer of the YOLOv8 model into an MSAFDN module. The MSAFDN module is used to comprehensively process each scale feature map extracted from different scale layers of the feature pyramid to obtain classification and regression results.

[0011] S3. Based on the MPAB module, the MSAFDN module, and the improved CIOU loss function, obtain an improved YOLOv8 model. Improve the CIOU loss function according to the width and height differences between the true box and the predicted box to obtain an improved CIOU loss function. And improve the loss function of the YOLOv8 model through the improved CIOU loss function.

[0012] S4. According to the MPAB module, the MSAFDN module, and the improved CIOU loss function, obtain an improved YOLOv8 model. Detect vehicle targets through the improved YOLOv8 model to obtain detection results.

[0013] Preferably, the specific steps of the MPAB module are as follows:

[0014] S10. Perform multi-scale segmentation on the input image to obtain multiple small-scale feature maps and large-scale feature maps. Input the small-scale feature maps into Path1 and the large-scale feature maps into Path2. Path1 includes axial MHSA and Bottleneck. Path2 includes a convolutional layer and Bottleneck. The small-scale feature maps are processed by Path1 to obtain the first small-scale feature maps. The large-scale feature maps are processed by Path2 to obtain the second large-scale feature maps.

[0015] S20. Introduce cross-layer connections between Path1, Path2, and the convolutional layer, and connect the shallow features of multiple first small-scale feature maps and second large-scale feature maps to the deep features through the cross-layer connections.

[0016] S30. Perform feature fusion on multiple first small-scale feature maps and second large-scale feature maps to obtain a first fusion feature map.

[0017] S40. Assign weight values to each channel in the first fusion feature map through a channel attention mechanism;

[0018] S50. Extract key features from the first fusion feature map according to the weight values to obtain a second key feature map;

[0019] S60. Obtain a third comprehensive feature by performing weighted residual connection on the second key feature map and the input image.

[0020] Preferably, the specific steps of the MSAFDN module are as follows:

[0021] S100. Receive each scale feature map extracted from different scale layers of the feature pyramid. The each scale feature map includes a first scale feature map, a second scale feature map, and a third scale feature map. Perform depthwise separable convolution and group normalization processing on each scale feature map respectively; obtain a first scale processed feature map, a second scale processed feature map, and a third scale processed feature map respectively;

[0022] S200. Introduce a multi-head self-attention mechanism to process the first scale processed feature map, the second scale processed feature map, and the third scale processed feature map respectively, and fuse them with the first scale processed feature map, the second scale processed feature map, and the third scale processed feature map respectively to obtain a first fusion feature map, a second fusion feature map, and a third fusion feature map

[0023] S300. Input the first fusion feature map, the second fusion feature map, and the third fusion feature map into the regression unit respectively. The regression unit includes an SE channel attention mechanism, depth convolution, an activation function, dynamic offset generation, and deformable convolution operations; process the first fusion feature map, the second fusion feature map, and the third fusion feature map through the SE channel attention mechanism, depth convolution, and activation function respectively to obtain a first channel feature map, a second channel feature map, and a third channel feature map; perform deformable convolution operations respectively; and enhance them through mask and offset, and finally obtain a first regression feature map, a second regression feature map, and a third regression feature map; the mask and the offset are generated by performing dynamic offset on the first fusion feature map, the second fusion feature map, and the third fusion feature map;

[0024] S400. Input the first fused feature map, the second fused feature map, and the third fused feature map into the classification unit respectively. The classification unit includes a CBAM spatial attention mechanism, depth convolution, an activation layer, a feature enhancement layer, and a feature map weighting operation. The feature enhancement layer obtains the first feature weight map, the second feature weight map, and the third feature weight map by performing two convolution processes and two activation function processes on the first fused feature map, the second fused feature map, and the third fused feature map. The first fused feature map, the second fused feature map, and the third fused feature map are processed by the CBAM spatial attention mechanism, depth convolution, and an activation function respectively to obtain the first spatial feature map, the second spatial feature map, and the third spatial feature map. The feature map weighting operation is used to multiply the first feature weight map with the first spatial feature map, the second feature weight map with the second spatial feature map, and the third feature weight map with the third spatial feature map respectively to obtain the first classification feature map, the second classification feature map, and the third classification feature map.

[0025] S500. Perform regression operations on the first regression feature map, the second regression feature map, and the third regression feature map respectively to obtain three regression results. Perform classification operations on the first classification feature map, the second classification feature map, and the third classification feature map respectively to obtain three classification results.

[0026] Preferably, the improved CIOU loss function improves the CIOU loss function based on the width and height differences between the ground truth box and the predicted box.

[0027] Preferably, the improved CIOU loss function is:

[0028]

[0029] where represents the improved CIOU loss function; IoU represents the ratio of the intersection to the union of the predicted box and the ground truth box; ρ 2 (b, b gt ) represents the square of the Euclidean distance between the center point of the predicted box and the center point of the ground truth box, which is used to measure the difference in their positions; b and b gt represent the center points of the predicted box and the ground truth box respectively; c represents the diagonal length of the smallest rectangle that can contain both the predicted box and the ground truth box; α represents a tuning parameter; w and h represent the width and height of the predicted box respectively; w gt and h gt represent the width and height of the ground truth box respectively; ε represents a constant term.

[0030] A vehicle target detection system based on multi-scale feature extraction includes:

[0031] The C2f improvement module is used to improve the first C2f module and the second C2f module in the Backbone layer of the YOLOv8 model into the MPAB module. The first C2f module and the second C2f module are the 6th layer C2f module and the 8th layer C2f module of the Backbone respectively. The MPAB module is used to perform multi-scale segmentation, cross-layer connection, feature fusion, channel attention mechanism and feature weighting operations on the input image to obtain the third comprehensive feature;

[0032] The Detect improvement module is used to improve the Detect module in the Head layer of the YOLOv8 model into the MSAFDN module. The MSAFDN module is used to comprehensively process each scale feature map extracted from different scale layers of the feature pyramid to obtain classification and regression results;

[0033] The loss function improvement module is used to improve the CIOU loss function according to the width and height differences between the ground truth box and the predicted box to obtain the improved CIOU loss function; and improve the loss function of the YOLOv8 model through the improved CIOU loss function;

[0034] The object detection module is used to obtain the improved YOLOv8 model according to the MPAB module, the MSAFDN module and the improved CIOU loss function; detect vehicle targets through the improved YOLOv8 model to obtain detection results.

[0035] Preferably, the specific steps of the MPAB module are as follows:

[0036] S10. Perform multi-scale segmentation on the input image to obtain multiple small-scale feature maps and large-scale feature maps. Input the small-scale feature maps into Path1 and the large-scale feature maps into Path2. Path1 includes axial MHSA and Bottleneck; Path2 includes a convolutional layer and Bottleneck. The small-scale feature maps are processed by Path1 to obtain the first small-scale feature maps; the large-scale feature maps are processed by Path2 to obtain the second large-scale feature maps;

[0037] S20. Introduce cross-layer connections between Path1, Path2 and the convolutional layer, and connect the shallow features of multiple first small-scale feature maps and second large-scale feature maps to the deep features through the cross-layer connections;

[0038] S30. Perform feature fusion on multiple first small-scale feature maps and second large-scale feature maps to obtain the first fusion feature map;

[0039] S40. Assign weight values to each channel in the first fusion feature map through the channel attention mechanism;

[0040] S50. Extract key features from the first fusion feature map according to the weight values to obtain a second key feature map;

[0041] S60. Obtain a third comprehensive feature by performing weighted residual connection on the second key feature map and the input image.

[0042] Preferably, the specific steps of the MSAFDN module are as follows: S100. Receive each scale feature map extracted from different scale layers of the feature pyramid. The each scale feature map includes a first scale feature map, a second scale feature map, and a third scale feature map. Perform depthwise separable convolution and group normalization processing on each scale feature map respectively; obtain a first scale processed feature map, a second scale processed feature map, and a third scale processed feature map respectively;

[0043] S200. Introduce a multi-head self-attention mechanism to process the first scale processed feature map, the second scale processed feature map, and the third scale processed feature map respectively, and fuse them with the first scale processed feature map, the second scale processed feature map, and the third scale processed feature map respectively to obtain a first fusion feature map, a second fusion feature map, and a third fusion feature map

[0044] S300. Input the first fusion feature map, the second fusion feature map, and the third fusion feature map into the regression unit respectively. The regression unit includes an SE channel attention mechanism, depth convolution, an activation function, dynamic offset generation, and deformable convolution operations; process the first fusion feature map, the second fusion feature map, and the third fusion feature map through the SE channel attention mechanism, depth convolution, and the activation function respectively to obtain a first channel feature map, a second channel feature map, and a third channel feature map; perform deformable convolution operations respectively; and enhance them through mask and offset. Finally, obtain a first regression feature map, a second regression feature map, and a third regression feature map; the mask and the offset are generated by performing dynamic offset on the first fusion feature map, the second fusion feature map, and the third fusion feature map;

[0045] S400. Input the first fusion feature map, the second fusion feature map, and the third fusion feature map into the classification unit respectively. The classification unit includes a CBAM spatial attention mechanism, depth convolution, an activation layer, a feature enhancement layer, and a feature map weighting operation. The feature enhancement layer obtains the first feature weight map, the second feature weight map, and the third feature weight map by performing two convolutional processes and two activation function processes on the first fusion feature map, the second fusion feature map, and the third fusion feature map respectively. Process the first fusion feature map, the second fusion feature map, and the third fusion feature map through the CBAM spatial attention mechanism, depth convolution, and activation function respectively to obtain the first spatial feature map, the second spatial feature map, and the third spatial feature map. The feature map weighting operation is used to multiply the first feature weight map with the first spatial feature map, the second feature weight map with the second spatial feature map, and the third feature weight map with the third spatial feature map respectively to obtain the first classification feature map, the second classification feature map, and the third classification feature map.

[0046] S500. Perform regression operations on the first regression feature map, the second regression feature map, and the third regression feature map respectively to obtain three regression results. Perform classification operations on the first classification feature map, the second classification feature map, and the third classification feature map respectively to obtain three classification results.

[0047] Preferably, the improved CIOU loss function improves the CIOU loss function based on the width and height differences between the ground truth box and the predicted box.

[0048] Preferably, the improved CIOU loss function is:

[0049]

[0050] Where represents the improved CIOU loss function; IoU represents the ratio of the intersection to the union of the predicted box and the ground truth box; ρ 2 (b, b gt ) represents the square of the Euclidean distance between the center point of the predicted box and the center point of the ground truth box, which is used to measure the difference in their positions; c represents the length of the diagonal of the smallest rectangle that can contain both the predicted box and the ground truth box; b and b gt represent the center points of the predicted box and the ground truth box respectively; α represents a tuning parameter; w and h represent the width and height of the predicted box respectively; w gt and h gt represent the width and height of the ground truth box respectively; ε represents a constant term.

[0051] Compared with the prior art, the beneficial effects of the present invention are:

[0052] 1. The present invention proposes an improved MPAB module. Through the combined innovative design of multi-path shunting, cross-layer connection, feature fusion, channel attention mechanism, and weighted residuals, the diversity of feature extraction and network stability are significantly improved. Multi-scale segmentation divides the input image into feature maps of different scales, and shunts them through Path1 (axial MHSA and Bottleneck) and Path2 (convolutional layer and Bottleneck), improving the comprehensiveness of feature extraction and the efficiency of information flow; cross-layer connection increases feature diversity by connecting shallow and deep features; and by fusing global and local features, the feature expression ability of the model at different scales is enhanced; the channel attention mechanism effectively highlights important features; weighted residuals improve the performance of deep networks and ensure information flow; integrating these designs, the MPAB module can achieve stronger feature extraction ability in vehicle target detection tasks, thus helping to improve the accuracy of vehicle target detection and avoid missed detections.

[0053] 2. The present invention proposes an MSAFDN module. The MSAFDN module can effectively extract spatial and semantic information of different resolutions through multi-scale feature processing and adaptive feature fusion module, enhancing the network's global and local perception ability of vehicle targets and improving the applicability and accuracy of vehicle target detection in complex background and multi-target scenarios; adopting the strategy of feature decoupling to achieve effective classification and regression tasks, and combining operations such as spatial attention mechanism and feature map multiplication to improve classification accuracy; combining operations such as channel attention, dynamic offset generation, and deformable convolution to further strengthen key features and improve regression accuracy; integrating the above designs can effectively improve the accuracy of vehicle target detection and avoid missed detections.

[0054] 3. The present invention proposes an improved CIOU loss function. By introducing the absolute difference of width and height and normalization processing, this improved method solves the optimization bottleneck problem of traditional CIOU loss in the case of the same aspect ratio but large numerical differences. Through the smooth growth characteristic of the absolute difference, the model can more robustly adjust the width and height dimensions of the prediction box, improving the training efficiency; the newly added normalization processing not only avoids the numerical instability when the size of the ground truth box is close to zero, but also strengthens the ability to capture the differences in width and height, improving the sensitivity of the model to vehicle detection targets of different scales and further enhancing the robustness and accuracy of the detection model; in addition, by introducing adjustable weight factor α and smoothing factor ε, it reflects the dynamic balance between the center point distance loss and the width-height difference loss. Integrating the above designs can effectively improve the accuracy of vehicle target detection and avoid missed detections. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a schematic flowchart of a vehicle target detection method based on multi-scale feature extraction provided by an embodiment of the present invention;

[0056] Figure 2 Schematic diagram of an MPAB structure provided by an embodiment of the present invention;

[0057] Figure 3 Schematic diagram of an MSAFDN module structure provided by an embodiment of the present invention;

[0058] Figure 4 Schematic diagram of a vehicle target detection system structure based on multi-scale feature extraction provided by an embodiment of the present invention. Detailed implementation manners

[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0060] Embodiment 1

[0061] In order to improve the detection accuracy of vehicle A, a vehicle target detection method based on multi-scale feature extraction is applied;

[0062] Refer to Figure 1 , which is a schematic diagram of the process of a vehicle target detection method based on multi-scale feature extraction provided by an embodiment of the present invention, including:

[0063] S1. Improve the first C2f module and the second C2f module in the Backbone layer of the YOLOv8 model into MPAB modules. The first C2f module and the second C2f module are the 6th layer C2f module and the 8th layer C2f module of the Backbone respectively. The MPAB module is used to perform multi-scale segmentation, cross-layer connection, feature fusion, channel attention mechanism and feature weighting operations on the input image to obtain a third comprehensive feature;

[0064] Refer to Figure 2 , which is a schematic diagram of an MPAB structure provided by an embodiment of the present invention;

[0065] Furthermore, the specific steps of the MPAB module are:

[0066] S10. Perform multi-scale segmentation on the input image to obtain multiple small-scale feature maps and large-scale feature maps. Input the small-scale feature maps into Path1, and input the large-scale feature maps into Path2. Path1 includes an axial MHSA and a Bottleneck; Path2 includes a convolutional layer and a Bottleneck. The small-scale feature maps are processed by Path1 to obtain the first small-scale feature maps; the large-scale feature maps are processed by Path2 to obtain the second large-scale feature maps.

[0067] S20. Introduce cross-layer connections between Path1, Path2, and the convolutional layer, and connect the shallow features of multiple first small-scale feature maps and second large-scale feature maps to the deep features through the cross-layer connections.

[0068] S30. Perform feature fusion on multiple first small-scale feature maps and second large-scale feature maps to obtain the first fused feature map.

[0069] S40. Assign weight values to each channel in the first fused feature map through the channel attention mechanism.

[0070] S50. Extract key features from the first fused feature map according to the weight values to obtain the second key feature map.

[0071] S60. Obtain the third comprehensive feature through weighted residual connection between the second key feature map and the input image.

[0072] This embodiment proposes an improved MPAB module. Through the combined innovative design of multi-path shunting, cross-layer connection, feature fusion, channel attention mechanism, and weighted residual, it significantly improves the diversity of feature extraction and network stability. Multi-scale segmentation divides the input image into feature maps of different scales, and shunts them through Path1 (axial MHSA and Bottleneck) and Path2 (convolutional layer and Bottleneck) to improve the comprehensiveness of feature extraction and the efficiency of information flow; cross-layer connection increases feature diversity by connecting shallow and deep features; and by fusing global and local features, it enhances the model's feature expression ability at different scales; the channel attention mechanism effectively highlights important features; weighted residual improves the performance of the deep network and ensures information fluidity; integrating these designs, the MPAB module can achieve stronger feature extraction ability in vehicle target detection tasks, which is conducive to improving the accuracy of vehicle target detection and avoiding missed detections.

[0073] S2. Improve the Detect module in the Head layer of the YOLOv8 model to the MSAFDN module; the MSAFDN module is used to comprehensively process each scale feature map extracted from different scale layers of the feature pyramid to obtain classification and regression results.

[0074] Refer to Figure 3 , which is a schematic structural diagram of an MSAFDN module provided by an embodiment of the present invention;

[0075] Furthermore, the specific steps of the MSAFDN module are as follows:

[0076] S100. Receive each scale feature map extracted from different scale layers of the feature pyramid. The each scale feature map includes a first scale feature map, a second scale feature map and a third scale feature map, and perform depthwise separable convolution and group normalization processing on each scale feature map respectively; obtain a first scale processed feature map, a second scale processed feature map and a third scale processed feature map respectively;

[0077] S200. Introduce a multi-head self-attention mechanism to process the first scale processed feature map, the second scale processed feature map and the third scale processed feature map respectively, and fuse them with the first scale processed feature map, the second scale processed feature map and the third scale processed feature map respectively to obtain a first fused feature map, a second fused feature map and a third fused feature map

[0078] S300. Input the first fused feature map, the second fused feature map and the third fused feature map into the regression unit respectively. The regression unit includes an SE channel attention mechanism, depth convolution, an activation function, dynamic offset generation and deformable convolution operations; process the first fused feature map, the second fused feature map and the third fused feature map through the SE channel attention mechanism, depth convolution and the activation function respectively to obtain a first channel feature map, a second channel feature map and a third channel feature map; and perform deformable convolution operations respectively; and enhance them through mask and offset, and finally obtain a first regression feature map, a second regression feature map and a third regression feature map; the mask and the offset are generated by performing dynamic offset on the first fused feature map, the second fused feature map and the third fused feature map;

[0079] S400. Input the first fused feature map, the second fused feature map, and the third fused feature map into the classification unit respectively. The classification unit includes a CBAM spatial attention mechanism, depthwise convolution, an activation layer, a feature enhancement layer, and a feature map weighting operation. The feature enhancement layer obtains the first feature weight map, the second feature weight map, and the third feature weight map by performing two convolutional processes and two activation function processes on the first fused feature map, the second fused feature map, and the third fused feature map. The first fused feature map, the second fused feature map, and the third fused feature map are processed by the CBAM spatial attention mechanism, depthwise convolution, and an activation function respectively to obtain the first spatial feature map, the second spatial feature map, and the third spatial feature map. The feature map weighting operation is used to multiply the first feature weight map with the first spatial feature map, the second feature weight map with the second spatial feature map, and the third feature weight map with the third spatial feature map respectively to obtain the first classification feature map, the second classification feature map, and the third classification feature map.

[0080] S500. Perform regression operations on the first regression feature map, the second regression feature map, and the third regression feature map respectively to obtain three regression results. Perform classification operations on the first classification feature map, the second classification feature map, and the third classification feature map respectively to obtain three classification results.

[0081] This embodiment proposes an MSAFDN module. The MSAFDN module can effectively extract spatial and semantic information of different resolutions through multi-scale feature processing and an adaptive feature fusion module, enhance the network's global and local perception capabilities of vehicle targets, and improve the applicability and accuracy of vehicle target detection in complex backgrounds and multi-target scenarios. It adopts a feature decoupling strategy to achieve effective classification and regression tasks, combines operations such as a spatial attention mechanism and feature map multiplication, and improves classification accuracy. It further strengthens key features by combining channel attention, dynamic offset generation, and deformable convolution, etc., and improves regression accuracy. With the above comprehensive design, it can effectively improve the accuracy of vehicle target detection and avoid missed detections.

[0082] S3. Based on the MPAB module, the MSAFDN module, and the improved CIOU loss function, obtain the improved YOLOv8 model. Improve the CIOU loss function according to the width and height differences between the ground truth box and the predicted box to obtain the improved CIOU loss function. And improve the loss function of the YOLOv8 model through the improved CIOU loss function.

[0083] Furthermore, the improved CIOU loss function improves the CIOU loss function based on the width and height differences between the ground truth box and the predicted box.

[0084] Furthermore, the improved CIOU loss function is as follows:

[0085]

[0086] Wherein, represents the improved CIOU loss function; IoU represents the ratio of the intersection to the union of the predicted box and the ground truth box; ρ 2 (b, b gt ) represents the square of the Euclidean distance between the center point of the predicted box and the center point of the ground truth box, which is used to measure the difference in their positions; b and b gt represent the center points of the predicted box and the ground truth box respectively; c represents the diagonal length of the smallest rectangle that can contain both the predicted box and the ground truth box; α represents the adjustment parameter; w and h represent the width and height of the predicted box respectively; w gt and h gt represent the width and height of the ground truth box respectively; ε represents a constant term.

[0087] This embodiment proposes an improved CIOU loss function. By introducing the absolute difference in width and height and the normalization process, this improved method solves the optimization bottleneck problem of the traditional CIOU loss in the case of the same aspect ratio but large numerical differences. Due to the smooth growth characteristic of the absolute difference, the model can more robustly adjust the width and height dimensions of the predicted box, improving the training efficiency; the newly added normalization process not only avoids the numerical instability when the size of the ground truth box is close to zero, but also strengthens the ability to capture the difference in width and height, enhancing the sensitivity of the model to vehicle detection targets of different scales, and further improving the robustness and accuracy of the detection model; in addition, by introducing the adjustable weight factor α and the smoothing factor ε, it reflects the dynamic balance between the center point distance loss and the width-height difference loss. With the above comprehensive design, it can effectively improve the accuracy of vehicle target detection and avoid missed detections.

[0088] S4. According to the MPAB module, the MSAFDN module, and the improved CIOU loss function, an improved YOLOv8 model is obtained; the vehicle target is detected by the improved YOLOv8 model to obtain the detection result;

[0089] In this embodiment, first, the 6th and 8th C2f modules in the Backbone layer of the YOLOv8 model are respectively improved to MPAB modules; secondly, the Detect module in the Head layer of the YOLOv8 model is improved to MSAFDN module; then, the loss function of the YOLOv8 model is further optimized by the improved CIOU loss function; finally, the improved YOLOv8 model is used to detect the vehicle target to obtain the detection result. The present invention effectively improves the accuracy of vehicle target detection by improving the YOLOv8 model;

[0090] To verify the effectiveness of the improved YOLOv8 model provided in this embodiment, object detection is performed on vehicle A using different models, including Model 1, Model 2, Model 3, and Model 4 respectively; Model 1 is the improved YOLOv8 model provided in this embodiment; Model 2 is without considering the C2f improvement based on Model 1; Model 3 is without considering the Detect improvement based on Model 1; Model 4 is without considering the loss function improvement based on Model 1; The accuracy rates of different models for object detection of vehicle A are compared respectively, and the average value is obtained through multiple experiments. The specific comparison table is shown in Table 1;

[0091] Table 1 Average accuracy rates of different models for object detection of vehicle A

[0092] Model Average accuracy Model 1 96% Model 2 88% Model 3 82% Model 4 86%

[0093] As can be seen from Table 1, the improved YOLOv8 model provided in this embodiment has a certain degree of effectiveness.

[0094] Embodiment 2

[0095] To improve the object detection accuracy rate of vehicle B, a vehicle object detection system based on multi-scale feature extraction is applied;

[0096] Refer to Figure 4 , which is a schematic structural diagram of a vehicle object detection system based on multi-scale feature extraction provided in an embodiment of the present invention, including:

[0097] The C2f improvement module is used to improve the first C2f module and the second C2f module in the Backbone layer of the YOLOv8 model into MPAB modules. The first C2f module and the second C2f module are the 6th layer C2f module and the 8th layer C2f module of the Backbone respectively; The MPAB module is used to perform multi-scale segmentation, cross-layer connection, feature fusion, channel attention mechanism, and feature weighting operations on the input image to obtain the third comprehensive feature;

[0098] Refer to Figure 2 , which is a schematic structural diagram of an MPAB provided in an embodiment of the present invention;

[0099] Further, the specific steps of the MPAB module are as follows:

[0100] S10. Perform multi-scale segmentation on the input image to obtain multiple small-scale feature maps and large-scale feature maps. Input the small-scale feature maps into Path1 and the large-scale feature maps into Path2. Path1 includes an axial MHSA and a Bottleneck; Path2 includes a convolutional layer and a Bottleneck. The small-scale feature maps are processed by Path1 to obtain the first small-scale feature maps; the large-scale feature maps are processed by Path2 to obtain the second large-scale feature maps;

[0101] S20. Introduce cross-layer connections between Path1, Path2, and the convolutional layer, and connect the shallow features of multiple first small-scale feature maps and second large-scale feature maps to the deep features through the cross-layer connections;

[0102] S30. Perform feature fusion on multiple first small-scale feature maps and second large-scale feature maps to obtain the first fusion feature map;

[0103] S40. Assign weight values to each channel in the first fusion feature map through a channel attention mechanism;

[0104] S50. Extract key features from the first fusion feature map according to the weight values to obtain the second key feature map;

[0105] S60. Obtain the third comprehensive feature by performing weighted residual connection on the second key feature map and the input image.

[0106] The Detect improvement module is used to improve the Detect module in the Head layer of the YOLOv8 model to the MSAFDN module; the MSAFDN module is used to comprehensively process each scale feature map extracted from different scale layers of the feature pyramid to obtain classification and regression results;

[0107] Refer to Figure 3 , which is a schematic diagram of the structure of the MSAFDN module provided by an embodiment of the present invention;

[0108] Further, the specific steps of the MSAFDN module are as follows:

[0109] S100. Receive each scale feature map extracted from different scale layers of the feature pyramid. Each scale feature map includes a first scale feature map, a second scale feature map, and a third scale feature map. Perform depthwise separable convolution and group normalization processing on each scale feature map respectively to obtain a first scale processed feature map, a second scale processed feature map, and a third scale processed feature map;

[0110] S200. Introduce the multi-head self-attention mechanism to process the first-scale processed feature map, the second-scale processed feature map, and the third-scale processed feature map respectively, and fuse them with the first-scale processed feature map, the second-scale processed feature map, and the third-scale processed feature map respectively to obtain the first fused feature map, the second fused feature map, and the third fused feature map

[0111] S300. Input the first fused feature map, the second fused feature map, and the third fused feature map into the regression unit respectively. The regression unit includes the SE channel attention mechanism, depth convolution, activation function, dynamic offset generation, and deformable convolution operation; process the first fused feature map, the second fused feature map, and the third fused feature map through the SE channel attention mechanism, depth convolution, and activation function respectively to obtain the first channel feature map, the second channel feature map, and the third channel feature map; and perform deformable convolution operations respectively; and enhance them through mask and offset, and finally obtain the first regression feature map, the second regression feature map, and the third regression feature map; the mask and the offset are generated by performing dynamic offset on the first fused feature map, the second fused feature map, and the third fused feature map

[0112] S400. Input the first fused feature map, the second fused feature map, and the third fused feature map into the classification unit respectively. The classification unit includes the CBAM spatial attention mechanism, depth convolution, activation layer, feature enhancement layer, and feature map weighting operation; the feature enhancement layer processes the first fused feature map, the second fused feature map, and the third fused feature map through two convolution processes and two activation function processes to obtain the first feature weight map, the second feature weight map, and the third feature weight map; process the first fused feature map, the second fused feature map, and the third fused feature map through the CBAM spatial attention mechanism, depth convolution, and activation function respectively to obtain the first spatial feature map, the second spatial feature map, and the third spatial feature map; the feature map weighting operation is used to multiply the first feature weight map with the first spatial feature map, the second feature weight map with the second spatial feature map, and the third feature weight map with the third spatial feature map respectively to obtain the first classification feature map, the second classification feature map, and the third classification feature map

[0113] S500. Perform regression operations on the first regression feature map, the second regression feature map, and the third regression feature map respectively to obtain three regression results; perform classification operations on the first classification feature map, the second classification feature map, and the third classification feature map respectively to obtain three classification results

[0114] The loss function improvement module is used to improve the CIOU loss function according to the width and height differences between the ground truth boxes and the predicted boxes, obtaining an improved CIOU loss function; and improving the loss function of the YOLOv8 model through the improved CIOU loss function;

[0115] Furthermore, the improved CIOU loss function improves the CIOU loss function based on the width and height differences between the ground truth boxes and the predicted boxes.

[0116] Furthermore, the improved CIOU loss function is:

[0117]

[0118] where represents the improved CIOU loss function; IoU represents the ratio of the intersection to the union of the predicted box and the ground truth box; ρ 2 (b, b gt ) represents the square of the Euclidean distance between the center points of the predicted box and the ground truth box, used to measure the difference in their positions; b and b gt represent the center points of the predicted box and the ground truth box respectively; c represents the diagonal length of the smallest rectangle that can contain both the predicted box and the ground truth box; α represents an adjustment parameter; w and h represent the width and height of the predicted box respectively; w gt and h gt represent the width and height of the ground truth box respectively; ε represents a constant term.

[0119] The object detection module is used to obtain an improved YOLOv8 model according to the MPAB module, the MSAFDN module, and the improved CIOU loss function; detect vehicle objects through the improved YOLOv8 model to obtain detection results.

[0120] To verify the effectiveness of the improved YOLOv8 model provided in this embodiment, vehicle B is detected by different models, including Model 1, Model 2, Model 3, and Model 4 respectively; Model 1 is the improved YOLOv8 model provided in this embodiment; Model 2 does not consider the C2f improvement based on Model 1; Model 3 does not consider the Detect improvement based on Model 1; Model 4 does not consider the loss function improvement based on Model 1; the accuracy rates of different models for detecting vehicle B objects are compared respectively, and the average value is obtained through multiple experiments. The specific comparison table is shown in Table 2;

[0121] Table 2 Average accuracy rates of different models for detecting vehicle B objects

[0122] Model Average accuracy Model 1 95% Model 2 86% Model 3 83% Model 4 85%

[0123] As can be seen from Table 2, the improved YOLOv8 model provided in this embodiment has a certain degree of effectiveness.

[0124] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A vehicle target detection method based on multi-scale feature extraction, characterized in that, Including: S1. Improve the first C2f module and the second C2f module of the Backbone layer of the YOLOv8 model into an MPAB module. The first C2f module and the second C2f module are the 6th layer C2f module and the 8th layer C2f module of the Backbone respectively. The MPAB module is used to perform multi-scale segmentation, cross-layer connection, feature fusion, channel attention mechanism and feature weighting operations on the input image to obtain a third comprehensive feature; S2. Improve the Detect module of the Head layer of the YOLOv8 model into an MSAFDN module. The MSAFDN module is used to comprehensively process each scale feature map extracted from different scale layers of the feature pyramid to obtain classification and regression results; S3. Based on the MPAB module, the MSAFDN module and the improved CIOU loss function, obtain an improved YOLOv8 model. Improve the CIOU loss function according to the width and height differences between the ground truth box and the predicted box to obtain an improved CIOU loss function. And improve the loss function of the YOLOv8 model through the improved CIOU loss function; S4. Based on the MPAB module, the MSAFDN module and the improved CIOU loss function, obtain an improved YOLOv8 model. Detect vehicle targets through the improved YOLOv8 model to obtain detection results.

2. The vehicle target detection method based on multi-scale feature extraction according to claim 1, wherein: The specific steps of the MPAB module are as follows: S10. Perform multi-scale segmentation on the input image to obtain multiple small-scale feature maps and large-scale feature maps. Input the small-scale feature maps into Path1 and the large-scale feature maps into Path2. Path1 includes axial MHSA and Bottleneck. Path2 includes a convolutional layer and Bottleneck. The small-scale feature maps are processed by Path1 to obtain the first small-scale feature maps. The large-scale feature maps are processed by Path2 to obtain the second large-scale feature maps; S20. Introduce cross-layer connection between Path1, Path2 and the convolutional layer, and connect the shallow features of multiple first small-scale feature maps and second large-scale feature maps to the deep features through cross-layer connection; S30. Perform feature fusion on multiple first small-scale feature maps and second large-scale feature maps to obtain a first fused feature map; S40. Assign weight values to each channel in the first fused feature map through the channel attention mechanism; S50. Extract key features from the first fused feature map according to the weight values to obtain a second key feature map; S60. Obtain a third comprehensive feature through weighted residual connection between the second key feature map and the input image.

3. The vehicle target detection method based on multi-scale feature extraction according to claim 1, characterized in that: The specific steps of the MSAFDN module are as follows: S100. Receive each scale feature map extracted from different scale layers of the feature pyramid. Each scale feature map includes a first scale feature map, a second scale feature map and a third scale feature map. Perform depthwise separable convolution and group normalization processing on each scale feature map respectively; Obtain the first-scale processed feature map, the second-scale processed feature map, and the third-scale processed feature map respectively; S200. Introduce a multi-head self-attention mechanism to process the first-scale processed feature map, the second-scale processed feature map, and the third-scale processed feature map respectively, and fuse them with the first-scale processed feature map, the second-scale processed feature map, and the third-scale processed feature map respectively to obtain the first fusion feature map, the second fusion feature map, and the third fusion feature map S300. Input the first fusion feature map, the second fusion feature map, and the third fusion feature map into the regression unit respectively. The regression unit includes an SE channel attention mechanism, depth convolution, an activation function, dynamic offset generation, and deformable convolution operations; process the first fusion feature map, the second fusion feature map, and the third fusion feature map through the SE channel attention mechanism, depth convolution, and activation function respectively to obtain the first channel feature map, the second channel feature map, and the third channel feature map; And perform deformable convolution operations respectively; and enhance them through mask and offset, and finally obtain the first regression feature map, the second regression feature map, and the third regression feature map; the mask and the offset are generated by dynamic offset of the first fusion feature map, the second fusion feature map, and the third fusion feature map; S400. Input the first fusion feature map, the second fusion feature map, and the third fusion feature map into the classification unit respectively. The classification unit includes a CBAM spatial attention mechanism, depth convolution, an activation layer, a feature enhancement layer, and feature map weighting operations; the feature enhancement layer performs two convolution processes and two activation function processes on the first fusion feature map, the second fusion feature map, and the third fusion feature map respectively to obtain the first feature weight map, the second feature weight map, and the third feature weight map; Process the first fusion feature map, the second fusion feature map, and the third fusion feature map through the CBAM spatial attention mechanism, depth convolution, and activation function respectively to obtain the first spatial feature map, the second spatial feature map, and the third spatial feature map; the feature map weighting operation is used to multiply the first feature weight map with the first spatial feature map, the second feature weight map with the second spatial feature map, and the third feature weight map with the third spatial feature map respectively to obtain the first classification feature map, the second classification feature map, and the third classification feature map; S500. Perform regression operations on the first regression feature map, the second regression feature map, and the third regression feature map respectively to obtain three regression results; perform classification operations on the first classification feature map, the second classification feature map, and the third classification feature map respectively to obtain three classification results.

4. A vehicle target detection method based on multi-scale feature extraction according to claim 1, characterized in that: The improved CIOU loss function improves the CIOU loss function based on the width and height differences between the ground truth box and the predicted box.

5. A vehicle target detection method based on multi-scale feature extraction according to claim 1, characterized in that: The improved CIOU loss function is: Among them, L CIoUpro represents the improved CIOU loss function; IoU represents the ratio of the intersection to the union of the predicted bounding box and the ground truth bounding box; ρ 2 (b, b gt ) represents the square of the Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box, which is used to measure the difference in their positions; b and b gt represent the center points of the predicted bounding box and the ground truth bounding box respectively; c represents the length of the diagonal of the smallest rectangle that can simultaneously contain the predicted bounding box and the ground truth bounding box; α represents the adjustment parameter; w and h represent the width and height of the predicted bounding box respectively; w gt and h gt represent the width and height of the ground truth bounding box respectively; ε represents a constant term.

6. A vehicle target detection system based on multi-scale feature extraction, characterized in that, Including: The C2f improvement module is used to improve the first C2f module and the second C2f module of the Backbone layer of the YOLOv8 model into MPAB modules. The first C2f module and the second C2f module are the 6th layer C2f module and the 8th layer C2f module of the Backbone respectively. The MPAB module is used to perform multi-scale segmentation, cross-layer connection, feature fusion, channel attention mechanism and feature weighting operations on the input image to obtain a third comprehensive feature; The Detect improvement module is used to improve the Detect module of the Head layer of the YOLOv8 model into an MSAFDN module; The MSAFDN module is used to comprehensively process each scale feature map extracted from different scale layers of the feature pyramid to obtain classification and regression results; The loss function improvement module is used to improve the CIOU loss function according to the width and height differences between the ground truth box and the predicted box to obtain an improved CIOU loss function; and improve the loss function of the YOLOv8 model through the improved CIOU loss function; The object detection module is used to obtain an improved YOLOv8 model according to the MPAB module, the MSAFDN module and the improved CIOU loss function; detect vehicle targets through the improved YOLOv8 model to obtain detection results.

7. The vehicle target detection system based on multi-scale feature extraction according to claim 6, characterized in that: The specific steps of the MPAB module are as follows: S10. Perform multi-scale segmentation on the input image to obtain multiple small-scale feature maps and large-scale feature maps. Input the small-scale feature maps into Path1 and the large-scale feature maps into Path2. Path1 includes axial MHSA and Bottleneck; Path2 includes a convolutional layer and Bottleneck. The small-scale feature maps are processed by Path1 to obtain the first small-scale feature maps; the large-scale feature maps are processed by Path2 to obtain the second large-scale feature maps; S20. Introduce cross-layer connections between Path1, Path2 and the convolutional layer, and connect the shallow features of multiple first small-scale feature maps and second large-scale feature maps to the deep features through cross-layer connections; S30. Perform feature fusion on multiple first small-scale feature maps and second large-scale feature maps to obtain a first fusion feature map; S40. Assign weight values to each channel in the first fusion feature map through a channel attention mechanism; S50. Extract key features from the first fusion feature map according to the weight values to obtain a second key feature map; S60. Obtain a third comprehensive feature through weighted residual connection of the second key feature map and the input image.

8. The vehicle target detection system based on multi-scale feature extraction according to claim 6, characterized in that: The specific steps of the MSAFDN module are as follows: S100. Receive each scale feature map extracted from different scale layers of the feature pyramid. Each scale feature map includes a first scale feature map, a second scale feature map and a third scale feature map, and perform depthwise separable convolution and group normalization processing on each scale feature map respectively; Respectively obtain a first scale processed feature map, a second scale processed feature map and a third scale processed feature map; S200. Introduce the multi-head self-attention mechanism to process the first-scale processed feature map, the second-scale processed feature map, and the third-scale processed feature map respectively, and fuse them with the first-scale processed feature map, the second-scale processed feature map, and the third-scale processed feature map respectively to obtain the first fused feature map, the second fused feature map, and the third fused feature map. S300. Input the first fused feature map, the second fused feature map, and the third fused feature map into the regression unit respectively. The regression unit includes the SE channel attention mechanism, depth convolution, activation function, dynamic offset generation, and deformable convolution operation. Process the first fused feature map, the second fused feature map, and the third fused feature map through the SE channel attention mechanism, depth convolution, and activation function respectively to obtain the first channel feature map, the second channel feature map, and the third channel feature map. And perform deformable convolution operations respectively; and enhance them through mask and offset, and finally obtain the first regression feature map, the second regression feature map, and the third regression feature map; the mask and the offset are generated by dynamic offset of the first fused feature map, the second fused feature map, and the third fused feature map. S400. Input the first fused feature map, the second fused feature map, and the third fused feature map into the classification unit respectively. The classification unit includes the CBAM spatial attention mechanism, depth convolution, activation layer, feature enhancement layer, and feature map weighting operation. The feature enhancement layer processes the first fused feature map, the second fused feature map, and the third fused feature map through two convolution processes and two activation function processes to obtain the first feature weight map, the second feature weight map, and the third feature weight map. Process the first fused feature map, the second fused feature map, and the third fused feature map through the CBAM spatial attention mechanism, depth convolution, and activation function respectively to obtain the first spatial feature map, the second spatial feature map, and the third spatial feature map; the feature map weighting operation is used to multiply the first feature weight map with the first spatial feature map, multiply the second feature weight map with the second spatial feature map, and multiply the third feature weight map with the third spatial feature map respectively to obtain the first classification feature map, the second classification feature map, and the third classification feature map. S500. Perform regression operations on the first regression feature map, the second regression feature map, and the third regression feature map respectively to obtain three regression results; perform classification operations on the first classification feature map, the second classification feature map, and the third classification feature map respectively to obtain three classification results.

9. The vehicle target detection system based on multi-scale feature extraction according to claim 6, characterized in that: The improved CIOU loss function improves the CIOU loss function based on the width and height differences between the ground truth box and the predicted box.

10. A vehicle target detection system based on multi-scale feature extraction according to claim 6, characterized in that: The improved CIOU loss function is: Among them, L CIoUpro represents the improved CIOU loss function; IoU represents the ratio of the intersection to the union of the predicted bounding box and the ground truth bounding box; ρ 2 (b, b gt ) represents the square of the Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box, which is used to measure the difference in their positions; b and b gt represent the center points of the predicted bounding box and the ground truth bounding box respectively; c represents the length of the diagonal of the smallest rectangle that can simultaneously contain the predicted bounding box and the ground truth bounding box; α represents the adjustment parameter; w and h represent the width and height of the predicted bounding box respectively; w gt and h gt represent the width and height of the ground truth bounding box respectively; ε represents a constant term.