Traffic accident detection method based on FFC and GCSA models

By introducing the FFC module and GCSA module into the YOLOv8 network and using the Focal-EIoU loss function to enhance feature expression and detection capabilities, the problems of low efficiency and insufficient accuracy in traffic accident detection in complex scenarios are solved, and more efficient traffic accident detection is achieved.

CN120708171APending Publication Date: 2025-09-26HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510809393.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing traffic accident detection methods perform poorly in complex scenarios, such as sudden changes in illumination, rain and fog interference, target occlusion, and small-scale object detection, resulting in limited detection efficiency and reliability.

Method used

A traffic accident detection method based on the FFC and GCSA models is adopted. By replacing the backbone network module of the YOLOv8 network with the FFC module, introducing the global channel-spatial attention GCSA module, and using the Focal-EIoU loss function, feature expression and detection capabilities are enhanced.

Benefits of technology

The accuracy and robustness of traffic accident detection in complex environments are improved, especially the small target detection performance, which solves the problem of low detection efficiency of existing methods in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708171A_ABST
    Figure CN120708171A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic accident detection method based on FFC and GCSA models. The method is innovatively improved based on a YOLOv8 network model. Firstly, a multi-scale feature fusion module FFC is designed in Backbone to replace an original C2f module, and the capability of fusing the multi-scale features of the network can be further improved on the premise of not increasing the network calculation amount, so that the network can fuse the multi-scale features on the level of finer granularity; then, a GCSA attention mechanism is introduced between a C2f module and a Deect module of a Neck layer, and the recognition capability of the model for key features in a complex environment is remarkably enhanced; and finally, establishing a loss function Focal-EIoU Loss for the improved YOLOv8 network structure, so that the network classification detection capability is improved, and the generalization capability of the model is improved. In a specific implementation process, a traffic accident image data set covering multiple scenes is constructed, and specialized data preprocessing is performed; then, end-to-end training and parameter optimization are carried out on the Traffic-YOLO network model obtained after improvement; and finally, integrating the optimized model to a traffic accident detection system for real-time target detection. Compared with the prior art, the method effectively improves the accuracy and robustness of traffic accident detection in a complex scene, and has important practical significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of traffic management and target detection, and in particular relates to a traffic accident detection method based on FFC and GCSA models. Background Art

[0002] Traffic accident detection, as a key component of intelligent transportation systems, is of undeniable importance. Real-time and reliable traffic accident detection methods can play a vital role in reducing casualties and property damage. With the rapid development of intelligent transportation systems and the continuous advancement of computer vision and object detection technologies, traffic accident detection research based on these technologies has garnered widespread attention. Numerous researchers have made significant progress in this field and proposed a variety of innovative solutions.

[0003] However, due to the complexity and variability of traffic scenes, some existing traffic accident detection methods still face challenges in practical applications. For example, in scenarios such as sudden changes in illumination, rain and fog interference, target occlusion, and detection of small-scale objects (such as vehicle debris and small vehicles), the performance of existing methods is often limited. These complex scenarios affect the efficiency and reliability of existing technologies in practical applications. Therefore, how to improve the accuracy and robustness of traffic accident detection, especially to achieve efficient target detection in complex environments, has become an urgent problem to be solved in current technological development. Developing a traffic accident detection method that can effectively cope with the above challenges is of great significance to improving traffic safety and promoting the further development of intelligent transportation systems. Summary of the Invention

[0004] Purpose of the invention: In response to the problems mentioned in the background technology, the present invention proposes a traffic accident detection method based on FFC and GCSA models, which is suitable for traffic accident detection in complex scenarios, improves the accuracy and robustness of traffic accident detection, and is suitable for integration into traffic accident detection systems to carry out real-time target detection.

[0005] Technical solution: The present invention proposes a traffic accident detection method based on the FFC and GCSA models, comprising the following steps:

[0006] S1. Obtain a traffic accident image dataset, and preprocess and label the dataset;

[0007] S2. Using the YOLOv8 network model as the base model, a Traffic-YOLO network model is constructed. The Traffic-YOLO network model replaces the C2f module of the YOLOv8 backbone network with the FFC module. A global channel-spatial attention (GCSA) module is designed to enhance the representation capability of the input feature map and inserted between the YOLOv8 head network and the YOLOv8 neck network.

[0008] S3. Establish Focal-EIoU Loss loss function;

[0009] S4. Train the Traffic-YOLO network model based on the preprocessed data set to obtain a Traffic-YOLO traffic accident detection model, and perform real-time detection using the Traffic-YOLO traffic accident detection model.

[0010] Furthermore, in step S1, the process of obtaining the data set includes: autonomously collecting traffic accident images, collecting open-source traffic accident detection data sets, and integrating the autonomously collected traffic accident images with the open-source traffic accident detection data sets to obtain a final data set.

[0011] Furthermore, the preprocessing process of the dataset includes: unifying the image size and orientation of the images in the dataset; performing geometric transformations such as rotation, flipping, translation and scaling on the images, and adding random noise to the images to simulate different shooting environments and quality;

[0012] When annotating the dataset, use LabelImg and LabelMe annotation tools to annotate the processed images and convert the annotated data into YOLO training format.

[0013] Furthermore, in step S2, the Traffic-YOLO network model includes three parts: the backbone network, the neck network, and the head network.

[0014] The backbone network includes at least several conventional convolutional layers, several FFC modules, and at least one spatial pyramid pooling module (SPPF). In the backbone network, the input image first passes through several groups of convolutional layers and FFC modules to gradually extract feature maps of different scales. Each group of convolutional layers and FFC modules inputs the extracted features into the next layer for higher-level feature extraction until the output is sent to the SPPF module at the end of the backbone network.

[0015] The neck network Neck includes at least two upsampling modules Upsample, several concatenation modules Concat, several convolution modules C2f, and several global channel-spatial attention modules GCSA. In the neck network, the high-level feature map output by the SPPF module of the backbone network is first upsampled and finally fused with the low-level feature map of the backbone network through the concatenation module Concat. The fused feature map is then passed through the convolution module C2f and the global channel-spatial attention GCSA module for feature extraction and information emphasis.

[0016] The head network Head receives the multi-scale feature maps from the neck network and performs the final target detection task through the detection module Detect.

[0017] Furthermore, the FFC module implements a three-level mechanism of “branch division → cascade fusion → residual reconstruction” to achieve multi-feature interaction at a microscale. The specific implementation process is as follows:

[0018] (1) Feature compression and segmentation stage: The input feature map F is compressed twice by 1×1 convolution channels to generate F1 and F2; F2 is evenly divided into 4 parts according to the channel dimension, represented by x1, x2, x3, and x4 respectively;

[0019] (2) Cascade multi-scale fusion stage:

[0020] If the FFC module is a single cascade: x1 directly participates in subsequent splicing without convolution perturbation, retaining high-resolution details; x2 passes through a 3×3 convolution layer to obtain feature map y2 for local feature extraction; x3 and y2 are added and passed through a 3×3 convolution layer to obtain feature map y3, achieving primary fusion; x4 and y3 are added and passed through a 3×3 convolution layer to obtain feature map y4, completing deep fusion;

[0021] If the FFC module is cascaded multiple times: use the feature map y4 obtained by the previous cascade multi-scale fusion as the input of the feature map F2 in the feature compression and segmentation stage, and continue to recursively perform the single cascade operation of the cascade multi-scale fusion stage;

[0022] (3) Feature reorganization and cross-layer connection stage: The four feature maps x1, y2, y3, and y4 are concatenated according to the channel dimension through Concat, and then passed through a 1×1 convolution layer to obtain the output feature map F3; the feature map F3 is added to the F1 feature map to obtain the feature map F4; the F1 obtained in the first step is passed through a 1×1 convolution layer, and the output feature map is recorded as F5; finally, F4 and F5 are concatenated according to the channel dimension, and then a 1×1 convolution is used to output the final feature F′.

[0023] Furthermore, the backbone network Backbone includes 4 FFC modules, each of which sets a cascade fusion recursion depth n, and the cascade fusion recursion depths n in the 4 FFC modules are 1, 2, 2, and 1 respectively; when n=1, only one cascade multi-scale fusion is performed to generate y2, y3, and y4 for feature reorganization and cross-layer connection; when n>1, y4 generated by the last cascade multi-scale fusion is used as F2 in the feature compression and segmentation stage, and n-1 cascade multi-scale fusion operations are recursively performed. The feature reorganization and cross-layer connection stage operations are performed using y2, y3, and y4 generated by the last cascade multi-scale fusion and x1 evenly divided at the time of the last cascade multi-scale fusion input.

[0024] Furthermore, the GCSA module combines channel attention, channel shuffling, and spatial attention mechanisms to enhance the expressiveness of the input feature map. The specific processing flow is as follows:

[0025] (1) The input feature map is first sent to the channel attention submodule and then to the spatial attention submodule. The initial feature map contains multiple channels, and the spatial size of each channel is H×W;

[0026] (2) In the channel attention submodule, the input feature map is first permuted from C×H×W to W×H×C. Then, the inter-channel dependency is captured by a two-layer multilayer perceptron (MLP). The first layer of MLP reduces the number of channels to 1 / 4 of the original. Then, nonlinearity is introduced through the ReLU activation function, and the second layer of MLP restores the number of channels to the original dimension. Finally, the inverse permutation is performed to restore it to C×H×W, and the channel attention map is generated through the Sigmoid activation function. The input feature map and the channel attention map are multiplied element by element to obtain the enhanced feature map.

[0027] (3) Applying the channel shuffle operation, the enhanced feature map is divided into 4 groups, each group contains C / 4 channels, and the grouped feature map is transposed to disrupt the channel order within each group; then, the shuffled feature map is restored to its original shape C×H×W;

[0028] (4) In the spatial attention submodule, the input feature map passes through a 7×7 convolutional layer, reducing the number of channels to 1 / 4 of the original number. It then undergoes batch normalization and a ReLU activation function for nonlinear transformation. It then passes through a second 7×7 convolutional layer to restore the number of channels to the original dimension C, and then passes through a batch normalization layer. Finally, a sigmoid activation function is used to generate a spatial attention map. The shuffled feature map and the spatial attention map are element-wise multiplied to obtain the final output feature map.

[0029] Furthermore, in step S3, the EIoU Loss of the Focal-EIoU Loss loss function measures the differences between the three geometric factors in the bounding box, namely the overlapping area, the center point and the side length, and calculates the difference between the width and height of the target instead of the aspect ratio, and proposes a Focal Loss regression version, which focuses more on high-quality anchor boxes in the regression process. Finally, by combining EIoU Loss and FocalLoss, Focal-EIoU Loss is finally obtained.

[0030] Furthermore, the Focal-EIoU Loss loss function is implemented as follows:

[0031]

[0032] L Focal-EIoU =IoU γ L EIoU

[0033] Where, L EIoU is the EIOU loss function; L Focal-EIoU is the Focal-EIoU loss function; IOU is the intersection-union ratio between the predicted box and the real box; d c is the distance between the center point of the predicted box and the real box; d d is the diagonal length of the minimum bounding rectangle between the predicted box and the true box; w' is the difference between the width of the penalty predicted box and the width of the true box; w c is the minimum bounding rectangle width; h' is the difference between the height of the penalty prediction box and the height of the real box; h c is the minimum circumscribed rectangle height; γ is the hyperparameter that controls the curvature of the curve.

[0034] Furthermore, in step S4, the Traffic-YOLO network model is trained to obtain the traffic accident detection model Traffic-YOLO, the number of model training rounds is set to 200 rounds, 16 images are input for one training, the input image size is uniformly adjusted to 640×640, the initial learning rate is set to 0.01, the minimum learning rate is set to 0.001, and the optimizer is set to stochastic gradient descent SGD. During the training process, the training log is observed in real time through Wandb, and the training results are saved after the training is completed.

[0035] Beneficial effects:

[0036] 1. The present invention abandons the C2f module in the main part of the original model and replaces it with the FFC module. This can further improve the network's ability to fuse multi-scale features without increasing the network's computational complexity, enabling the network to fuse multi-scale features at a more fine-grained level. This solves the problem of small target object feature loss after multiple convolutions in the target detection network, and improves the detection performance of small targets in traffic accidents.

[0037] 2. This paper introduces a lightweight attention mechanism GCSA in the Neck part. This module combines channel attention, channel shuffling and spatial attention mechanisms, and performs attention from both channel and spatial directions, so as to better capture and represent important features in images in complex environments and improve detection accuracy.

[0038] 3. The present invention uses Focal-EIoU Loss to replace the original loss function and accurately defines the aspect ratio of the prediction box, thereby alleviating the problem of imbalance between positive and negative samples, increasing the network classification and detection capabilities, and improving the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 Flow chart of the method of the present invention;

[0040] Figure 2 This is the Traffic-YOLO network structure diagram of the present invention;

[0041] Figure 3 This is the flow chart of the FFC module of the present invention;

[0042] Figure 4 This is a schematic diagram of the GCSA module of the present invention;

[0043] Figure 5 This is a test effect diagram of the benchmark model of the present invention;

[0044] Figure 6 This is a diagram of the detection effect after the improvement of the model of the present invention. DETAILED DESCRIPTION

[0045] The present invention is further illustrated below with reference to specific examples. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0046] This embodiment proposes a traffic accident detection method based on FFC and GCSA models, see Figure 1 , the specific implementation steps are as follows:

[0047] S1. Obtain a traffic accident image dataset and preprocess and label it:

[0048] In the present invention, step S1 involves autonomously collecting traffic accident images, covering as many diverse scenarios as possible, and assembling an open-source traffic accident detection dataset. Furthermore, the large number of collected traffic accident images is integrated with the open-source traffic accident image dataset to form the dataset of the present invention. The data is enhanced using OpenCV's image processing capabilities to increase the diversity of the dataset. Specifically, geometric transformations such as rotation, flipping, translation, and scaling are performed on the images to help the model learn how objects behave in different positions, scales, and orientations. Random noise is also added to the images to simulate different shooting environments and quality.

[0049] The dataset uses annotation software to automatically label various objects in the images and divides them into training and test sets. There are approximately 70,000 object-labeled bounding boxes in total, with approximately 3,824 images in the training set, 1,765 in the validation set, and 1,824 in the test set.

[0050] S2. Based on the YOLOv8 network model, a Traffic-YOLO network model is constructed, such as Figure 2 The Traffic-YOLO network model abandons the original C2f module in the backbone of the YOLOv8 feature extraction network and proposes a new FFC module. It also designs a global channel-spatial attention (GCSA) module that enhances the expressiveness of the input feature map.

[0051] In the present invention, the improved YOLOv8 network structure constructed in step S2 mainly includes a backbone network Backbone, a neck network Neck, and a head network Head.

[0052] The backbone network Backbone includes at least several conventional convolutional layers, several FFC modules and at least one spatial pyramid pooling module SPPF. In the backbone network, the input image first passes through several groups of convolutional layers and FFC modules to gradually extract feature maps of different scales. Each group of convolutional layers and FFC modules inputs the extracted features into the next layer for higher-level feature extraction until the output is sent to the SPPF module at the end of the backbone network.

[0053] The neck network Neck includes at least two upsampling modules Upsample, several concatenation modules Concat, several convolution modules C2f, and several efficient global channel-spatial attention modules GCSA. In the neck network, the high-level feature map output by the SPPF module of the backbone network is first upsampled, and finally fused with the low-level feature map of the backbone network through the concatenation module Concat. The fused feature map is passed through the convolution module C2f and the efficient global channel-spatial attention GCSA module for feature extraction and information emphasis.

[0054] The head network Head receives the multi-scale feature maps from the neck network and performs the final target detection task through the detection module Detect.

[0055] This paper seeks to improve the multi-scale feature representation capability at a finer level, so it abandons the original C2f module in the backbone of the YOLOv8 feature extraction network and proposes a new FFC module, such as Figure 3 As shown in the figure. The FFC module uses a three-level mechanism of "branch division → cascade fusion → residual recombination" to achieve multi-feature interaction at a micro scale, breaking through the limitation of traditional methods that only fuse between network layers. The specific implementation process is as follows:

[0056] (1) Feature compression and segmentation: The input feature map F undergoes two 1×1 convolutions (channel compression) to generate F1 and F2. F2 is evenly divided into four parts according to the channel dimension, represented by x1, x2, x3, and x4 respectively. This process prepares for subsequent more detailed feature processing.

[0057] (2) Cascaded multi-scale fusion:

[0058] x1 is directly retained to avoid convolutional perturbations to maintain high-resolution details.

[0059] x2 generates y2 through 3×3 convolution to extract local features.

[0060] After x3 and y2 are added, y3 is generated through a 3×3 convolution to achieve primary feature fusion.

[0061] After x4 and y3 are added, y4 is generated through 3×3 convolution to complete deep fusion.

[0062] The cascaded multi-scale fusion can be recursively performed n times (n ≥ 1), and the recursive depth of the cascade fusion is controlled by the parameter n:

[0063] When n=1, the above cascade multi-scale fusion is performed only once to generate y2, y3, and y4;

[0064] When n>1, the feature map y4 generated by the previous cascade multi-scale fusion is used as the feature map F2 input of the feature compression and segmentation stage, and the same chain operation is recursively performed n-1 times to achieve deeper feature interaction.

[0065] There are 4 C2f modules in the YOLOv8 feature extraction network Backbone, which are replaced by FFC modules. As the network downsampling proceeds, the parameters n in the 4 FFC modules are 1, 2, 2, and 1 respectively.

[0066] The design is based on:

[0067] Shallow module (n=1): processes high-resolution feature maps (larger in size), retains small target details, and avoids computational redundancy caused by deep recursion.

[0068] Intermediate module (n=2): responsible for medium-scale target detection (such as most traffic accident scene subjects), and enhances robustness to complex scenes such as occlusion and deformation through recursive fusion.

[0069] Deep module (n=1): processes low-resolution semantic features (background / large targets). Excessive fusion may introduce noise, so a lightweight operation is used.

[0070] The calculation formula of the FFC module (when n=1) is:

[0071]

[0072] When n>1, formula (2) needs to be iterated n-1 times.

[0073] Where: F is the input feature map; It is a 1×1 convolution operation; Split() divides the output feature map into 4 parts; is a 3×3 convolution operation; Cat() is a Concat operation, which means concatenating feature maps according to the channel dimension; F′ is the final output feature map.

[0074] (3) Feature Reorganization and Cross-layer Connection: The four feature maps x1, y2, y3, and y4 are concatenated according to the channel dimension through Concat, and then passed through a 1×1 convolution layer to obtain the output feature map F3; F3 and F1 feature maps are added to obtain feature map F4, avoiding gradient degradation and strengthening feature reuse; F1 obtained in the first step is passed through a 1×1 convolution layer, and the output feature map is recorded as F5; finally, F4 and F5 are concatenated according to the channel dimension, and then a 1×1 convolution is used to output the final feature F'. (Note: If multiple layers of cascaded multi-scale fusion are performed, the y2, y3, y4 generated by the last cascaded multi-scale fusion and the x1 evenly divided when the last cascaded multi-scale fusion is input are used for feature reorganization and cross-layer connection stage operations).

[0075] In this way, it not only helps to stabilize the training process, but also better retains important feature information and can achieve deep fusion of multi-scale features at a finer-grained level, thereby significantly improving the model's ability to capture the features of small target objects.

[0076] The present invention introduces a GCSA attention module between the C2f module of the neck network and the Detect module of the head network, such as Figure 4 As shown. Channel attention, channel shuffling, and spatial attention mechanisms are combined to enhance the expressiveness of the input feature map. The specific processing flow is as follows:

[0077] (1) The input feature map is first fed into the channel attention submodule and then into the spatial attention submodule. The initial feature map contains multiple channels, and the spatial size of each channel is H×W.

[0078] (2) In the channel attention submodule, the input feature map is first permuted from C×H×W to W×H×C. Then, the dependencies between channels are captured by two layers of multi-layer perceptrons (MLPs). The first layer of MLP reduces the number of channels to 1 / 4 of the original, then introduces nonlinearity through the ReLU activation function, and then restores the number of channels to the original dimension through the second layer of MLP. Finally, the inverse permutation is performed to restore it to C×H×W, and the channel attention map is generated through the Sigmoid activation function. The input feature map and the channel attention map are multiplied element by element to obtain the enhanced feature map:

[0079] (3) To further mix and share information, a channel shuffling operation is applied. The enhanced feature maps are divided into four groups, each containing C / 4 channels. The grouped feature maps are transposed to disrupt the order of channels within each group. Subsequently, the shuffled feature maps are restored to their original shape (C × H × W). This approach can better mix feature information and enhance feature expression capabilities.

[0080] (4) In the spatial attention submodule, the input feature map passes through a 7×7 convolution layer, reducing the number of channels to 1 / 4 of the original. Then, it undergoes batch normalization and ReLU activation function for nonlinear transformation. Next, the number of channels is restored to the original dimension C through a second 7×7 convolution layer, and then passed through a batch normalization layer. Finally, the spatial attention map is generated through the Sigmoid activation function. The shuffled feature map and the spatial attention map are element-wise multiplied to obtain the final output feature map.

[0081] S3. Establish the loss function Focal-EIoU Loss for the lightweight YOLOv8 network structure.

[0082] The original YOLOv8 model uses CIoU Loss to calculate the bounding box retrospective loss. However, CIoU Loss has a vague definition of the target aspect ratio when calculating the loss. When poor-quality regression samples have a significant impact on the regression loss, good-quality regression samples are difficult to optimize, leading to an imbalance between positive and negative samples. Therefore, to address these issues, as well as issues such as slow BBR convergence and inaccurate regression results, this paper introduces Focal-EIoU Loss to replace the CIoU Loss in the original YOLOv8n.

[0083] Among them, EIoU Loss measures the differences between three geometric factors in the bounding box, namely the overlapping area, the center point and the side length, and calculates the difference between the width and height of the target instead of the aspect ratio. A Focal Loss regression version is proposed, which focuses more on high-quality anchor boxes in the regression process. Finally, by combining EIoU Loss and Focal Loss, Focal-EIoU Loss is finally obtained.

[0084] The calculation formula of Focal-EIoU Loss is as follows:

[0085]

[0086] L Focal-EIoU =IoU γ L EIoU

[0087] Where, L EIoU is the EIOU loss function; LFocal-EIoU is the Focal-EIoU loss function; IOU is the intersection-union ratio between the predicted box and the real box; d c is the distance between the center point of the predicted box and the real box; d d is the diagonal length of the minimum bounding rectangle of the predicted box and the true box; w' is the difference between the width of the penalty predicted box and the width of the true box; w c is the minimum bounding rectangle width; h' is the difference between the height of the penalty prediction box and the height of the real box; h c is the minimum circumscribed rectangle height; γ is the hyperparameter that controls the curvature of the curve.

[0088] According to the above formula, the higher the IoU value, the greater the loss, which plays a weighted role. The better the regression target, the greater the loss, thereby improving the regression accuracy and balancing the contribution of high-quality samples and low-quality samples to the Loss, that is, increasing the contribution of high-quality (large IoU) samples and reducing the contribution of low-quality (small IoU) samples, thereby improving the performance of the loss function.

[0089] S4. Based on the processed data set, the improved YOLOv8 network structure is trained to obtain the Traffic-YOLO traffic accident detection model, and real-time detection is performed using the Traffic-YOLO traffic accident detection model.

[0090] This paper provides a training example to train the aforementioned improved YOLOv8 neural network to obtain the traffic accident detection model Traffic_YOLO. The training environment provided in this paper is Python 3.9.19;

[0091] We used torch2.3.1+cu121 with CUDA:0 (NVIDIA GeForce RTX 4090, 8192 MiB). We set the model training epochs to 200, with 16 input images per training. We used Wandb to monitor the training logs in real time during training and saved the training results after the training. We set the training parameters to 200 iterations, resized the input images to 640×640, set the initial learning rate to 0.01, the minimum learning rate to 0.001, the batch size to 16, and used stochastic gradient descent (SGD) as the optimizer.

[0092] The present invention uses precision, recall, mean average precision (mAP), parameter count (Params), and computational effort (GFLOPs) as evaluation indicators. The formulas for each evaluation indicator are as follows:

[0093] Precision (P):

[0094] Recall (R):

[0095] Average Mean Accuracy (mAP50):

[0096] Among them, TP refers to the number of samples that are positive and predicted as positive, FP refers to the number of samples that are negative and predicted as positive, FN refers to the number of samples that are positive and predicted as negative, n is the number of target categories detected, AP i is the AP of the i-th target class.

[0097] To demonstrate the effectiveness of the proposed Traffic-YOLO network model, we conducted ablation experiments based on a baseline network on a constructed dataset. Because deep learning experiments are inherently random, this experiment averages the results of multiple experiments to improve the feasibility of the experimental results. The experiment first tested the performance of each module in the baseline network one by one. Then, using a stacking method, we gradually added modules to the baseline network and compared the effectiveness of each module. The experimental results are shown in Table 1.

[0098] Table 1. Ablation experiments

[0099]

[0100] Note: √ indicates that the module is used

[0101] The Traffic-YOLO network model of the present invention shows a significant performance improvement compared to the baseline model. Specifically, after replacing the multi-scale fusion module FFC, the mAP of model A is higher than that of the baseline model YOLOv8n. 50 The improvement was 1.1%, and the number of parameters and floating point numbers were reduced by 0.4M and 0.2G respectively. Then, the GCSA module was introduced based on Model A for improvement. The mAP of Model B was compared with Model A. 50 It increased by 2.5%, and the number of parameters and floating point numbers decreased slightly. Finally, based on model B, this paper introduces a loss function Focal-EIoU Loss to improve the generalization ability of the model, and proposes the final improved model. Compared with the baseline model, the improved model improves P by 4% and mAP by 1. 50 It has been improved by 4.4%, the number of parameters has been reduced by 0.4M, and the number of floating point numbers has been reduced by 0.1G.

[0102] The present invention addresses the problem of low detection efficiency caused by sudden changes in illumination, rain and fog interference, target occlusion, and small-scale objects (such as vehicle debris and small vehicles) in complex traffic scenes. The original model is improved and has better detection effect than the baseline model. Figure 5 and Figure 6As shown, the detection of the model will be improved Figure 6 Detection of the original model Figure 5 In the case of small target detection, the original model 5 misdetected or missed small vehicles, and the damaged vehicles at the traffic accident scene that were obscured were not detected. Figure 6 The improved model has no false detection or missed detection and can accurately detect each target. Figure 5 and Figure 6 ,It can be seen that the improved model has better ,detection performance than the original model in the context of small ,vehicle detection and target occlusion.

[0103] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made in accordance with the spirit of the present invention shall be included in the scope of protection of the present invention.

Claims

1. A traffic accident detection method based on FFC and GCSA models, characterized in that: The steps include: S1. Obtain a traffic accident image dataset, and preprocess and label the dataset; S2. Using the YOLOv8 network model as the base model, a Traffic-YOLO network model is constructed. The Traffic-YOLO network model replaces the C2f module of the YOLOv8 backbone network with the FFC module. A global channel-spatial attention (GCSA) module is designed to enhance the representation capability of the input feature map and inserted between the YOLOv8 head network and the YOLOv8 neck network. S3. Establish Focal-EIoU Loss loss function; S4. Train the Traffic-YOLO network model based on the preprocessed data set to obtain a Traffic-YOLO traffic accident detection model, and perform real-time detection using the Traffic-YOLO traffic accident detection model.

2. The traffic accident detection method based on the FFC and GCSA models according to claim 1, characterized in that: In step S1, the process of obtaining the data set includes: autonomously collecting traffic accident images, collecting open-source traffic accident detection data sets, and integrating the autonomously collected traffic accident images with the open-source traffic accident detection data sets to obtain a final data set.

3. The traffic accident detection method based on the FFC and GCSA models according to claim 1, characterized in that: The preprocessing process of the dataset includes: unifying the image size and orientation of the images in the dataset; performing geometric transformations such as rotation, flipping, translation and scaling on the images, and adding random noise to the images to simulate different shooting environments and quality; When annotating the dataset, use LabelImg and LabelMe annotation tools to annotate the processed images and convert the annotated data into YOLO training format.

4. The traffic accident detection method based on the FFC and GCSA models according to claim 1, characterized in that: In step S2, the Traffic-YOLO network model includes three parts: the backbone network, the neck network, and the head network. The backbone network includes at least several conventional convolutional layers, several FFC modules, and at least one spatial pyramid pooling module (SPPF). In the backbone network, the input image first passes through several groups of convolutional layers and FFC modules to gradually extract feature maps of different scales. Each group of convolutional layers and FFC modules inputs the extracted features into the next layer for higher-level feature extraction until the output is sent to the SPPF module at the end of the backbone network. The neck network Neck includes at least two upsampling modules Upsample, several concatenation modules Concat, several convolution modules C2f, and several global channel-spatial attention modules GCSA. In the neck network, the high-level feature map output by the SPPF module of the backbone network is first upsampled and finally fused with the low-level feature map of the backbone network through the concatenation module Concat. The fused feature map is then passed through the convolution module C2f and the global channel-spatial attention GCSA module for feature extraction and information emphasis. The head network Head receives the multi-scale feature maps from the neck network and performs the final target detection task through the detection module Detect.

5. The traffic accident detection method based on the FFC and GCSA models according to claim 1, characterized in that: The FFC module implements a three-level mechanism of "branch division → cascade fusion → residual recombination" to achieve multi-feature interaction at a microscale. The specific implementation process is as follows: (1) Feature compression and segmentation stage: The input feature map F is compressed twice by 1×1 convolution channels to generate F1 and F2; F2 is evenly divided into 4 parts according to the channel dimension, represented by x1, x2, x3, and x4 respectively; (2) Cascade multi-scale fusion stage: If the FFC module is a single cascade: x1 directly participates in subsequent splicing without convolution perturbation, retaining high-resolution details; x2 passes through a 3×3 convolution layer to obtain feature map y2 for local feature extraction; x3 and y2 are added and passed through a 3×3 convolution layer to obtain feature map y3, achieving primary fusion; x4 and y3 are added and passed through a 3×3 convolution layer to obtain feature map y4, completing deep fusion; If the FFC module is cascaded multiple times: use the feature map y4 obtained by the previous cascade multi-scale fusion as the input of the feature map F2 in the feature compression and segmentation stage, and continue to recursively perform the single cascade operation of the cascade multi-scale fusion stage; (3) Feature reorganization and cross-layer connection stage: The four feature maps x1, y2, y3, and y4 are concatenated according to the channel dimension through Concat, and then passed through a 1×1 convolution layer to obtain the output feature map F3; The feature map F3 and F1 are added to obtain the feature map F4; the F1 obtained in the first step is passed through a 1×1 convolution layer, and the output feature map is recorded as F5; finally, F4 and F5 are spliced ​​according to the channel dimension, and then a 1×1 convolution is performed to output the final feature F′.

6. The traffic accident detection method based on the FFC and GCSA models according to claim 5, characterized in that: The backbone network consists of four FFC modules. Each FFC module sets a recursive depth n of cascade fusion. The recursive depth n of cascade fusion in the four FFC modules is 1, 2, 2, and 1, respectively. When n=1, only one cascade multi-scale fusion is performed to generate y2, y3, and y4 for feature reorganization and cross-layer connection. When n>1, y4 generated by the last cascade multi-scale fusion is used as F2 in the feature compression and segmentation stage, and n-1 cascade multi-scale fusion operations are recursively performed. The feature reorganization and cross-layer connection stage operations are performed using y2, y3, and y4 generated by the last cascade multi-scale fusion and x1 evenly divided at the time of the last cascade multi-scale fusion input.

7. The traffic accident detection method based on the FFC and GCSA models according to claim 1, characterized in that: The GCSA module combines channel attention, channel shuffling, and spatial attention mechanisms to enhance the expressiveness of the input feature map. The specific processing flow is as follows: (1) The input feature map is first sent to the channel attention submodule and then to the spatial attention submodule. The initial feature map contains multiple channels, and the spatial size of each channel is H×W; (2) In the channel attention submodule, the input feature map is first permuted from C×H×W to W×H×C. Then, the inter-channel dependency is captured by a two-layer multilayer perceptron (MLP). The first layer of MLP reduces the number of channels to 1 / 4 of the original. Then, nonlinearity is introduced through the ReLU activation function, and the second layer of MLP restores the number of channels to the original dimension. Finally, the inverse permutation is performed to restore it to C×H×W, and the channel attention map is generated through the Sigmoid activation function. The input feature map and the channel attention map are multiplied element by element to obtain the enhanced feature map. (3) Applying the channel shuffle operation, the enhanced feature map is divided into 4 groups, each group contains C / 4 channels, and the grouped feature map is transposed to disrupt the channel order within each group; then, the shuffled feature map is restored to its original shape C×H×W; (4) In the spatial attention submodule, the input feature map passes through a 7×7 convolutional layer, reducing the number of channels to 1 / 4 of the original number. It then undergoes batch normalization and a ReLU activation function for nonlinear transformation. It then passes through a second 7×7 convolutional layer to restore the number of channels to the original dimension C, and then passes through a batch normalization layer. Finally, a sigmoid activation function is used to generate a spatial attention map. The shuffled feature map and the spatial attention map are element-wise multiplied to obtain the final output feature map.

8. The traffic accident detection method based on the FFC and GCSA models according to claim 1, characterized in that: In step S3, the EIoU Loss of the Focal-EIoU Loss loss function measures the differences between the three geometric factors in the bounding box, namely the overlapping area, the center point and the side length, and calculates the difference between the width and height of the target instead of the aspect ratio. A Focal Loss regression version is proposed, which focuses more on high-quality anchor boxes in the regression process. Finally, by combining EIoU Loss and FocalLoss, Focal-EIoU Loss is finally obtained.

9. The traffic accident detection method based on FFC and GCSA models according to claim 8, characterized in that: The Focal-EIoU Loss loss function implementation formula is as follows: L Focal-EIoU =IoU γ L EIoU Where, L EIoU is the EIOU loss function; L Focal-EIoU is the Focal-EIoU loss function; IOU is the intersection-union ratio between the predicted box and the real box; d c is the distance between the center point of the predicted box and the real box; d d is the diagonal length of the minimum bounding rectangle of the predicted box and the true box; w' is the difference between the width of the penalty predicted box and the width of the true box; w c is the minimum bounding rectangle width; h' is the difference between the height of the penalty prediction box and the height of the real box; h c is the minimum circumscribed rectangle height; γ is the hyperparameter that controls the curvature of the curve.

10. The traffic accident detection method based on the FFC and GCSA models according to claim 1, characterized in that: In step S4, the Traffic-YOLO network model is trained to obtain a traffic accident detection model Traffic-YOLO. The number of model training rounds is set to 200 rounds, 16 images are input for each training, the input image size is uniformly adjusted to 640×640, the initial learning rate is set to 0.01, the minimum learning rate is set to 0.001, and the optimizer is set to stochastic gradient descent (SGD). During the training process, the training log is observed in real time through Wandb, and the training results are saved after the training is completed.

Citation Information

Cited By

  • Optimized YOLOv8-based anesthetic psychotropic drug identification model and training method

    CN121259507A

  • Industrial scene gesture recognition method and system based on improved YOLOv8

    CN121259929A

  • Foggy day road accident scene vehicle detection method based on fusion type backbone and hypergraph calculation

    CN122176649A

  • Foggy road accident scene vehicle detection method fusing backbone and hypergraph computation

    CN122176649B