Complex traffic environment-oriented small target detection algorithm model

By improving the YOLOv5s algorithm, adding four-scale detection, introducing a cascade structure of Ghost module and ECA attention mechanism, and using SIoU loss function, the dual challenges of small target detection accuracy and real-time in complex road traffic environments are solved, achieving higher detection accuracy and speed.

CN120220090APending Publication Date: 2025-06-27NANJING FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510289025.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In complex road traffic environments, it is difficult for the prior art to take into account both identification accuracy and real-time in small-object detection, especially when dealing with scenarios where small-objective characteristics are not significant, environmental variability is versatile, and handling high-demand requirements in real-time.

Method used

Improve the YOLOv5s algorithm, add four-scale detection, introduce the cascade structure of the Ghost module and the ECA attention mechanism, and use the SIoU loss function to improve the detection accuracy and speed of the network.

Benefits of technology

Through ablation experiments and comparison tests, the effectiveness of the improved module is verified, the detection accuracy is improved, and the speed is maintained, which is suitable for target detection in complex road traffic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220090A_ABST
    Figure CN120220090A_ABST
Patent Text Reader

Abstract

The invention discloses a complex traffic environment-oriented small target detection algorithm model in the technical field of intelligent automobile perception, which comprises a YOLOv5 algorithm model and further comprises a multi-scale detection module constructed based on the YOLOv5 algorithm model, and the multi-scale detection module is used for upgrading three-scale detection into four-scale detection so as to enrich feature information of a small target; the Ghost and ECA module cascade structure comprises a Ghost module, the Ghost module is used for solving the problem of feature map redundancy in the network training process, avoiding unnecessary feature repeated extraction and supplementing a new feature map through linear operation, ablation experiments are carried out on two data sets, the influence degree of each module on algorithm improvement precision is analyzed, and the algorithm improvement precision is improved. And the effectiveness of the improved module is verified. Besides, through comparison experiments with other mainstream algorithms and comparison tests in different complex road scenes, the algorithm shows superior performance and is superior to an original YOLOv5s model, and generalization applicability of the algorithm is proved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent vehicle perception, and specifically to a small target detection algorithm model for complex traffic environments. Background Art

[0002] With the acceleration of global urbanization and the sharp increase in the number of motor vehicles, the complex road traffic environment has become a major challenge faced by modern society. In such an environment, the detection of small targets (such as pedestrians, bicycles, and motorcycles, etc.) is of crucial significance for traffic monitoring, the development of autonomous driving systems, and accident prevention. However, complex road traffic scenarios usually include, but are not limited to, busy intersections, multi-lane highways, urban streets, and rural roads. Due to their high dynamics and uncertainty, combined with the characteristics of small targets themselves (such as small size, easy to be occluded, appearance diversity, etc.), it greatly increases the difficulty of detection and poses a potential threat to traffic safety.

[0003] In the field of target detection led by existing deep learning technologies, although the YOLO[1-8] series, SSD[9], and Faster R-CNN

[10] have made remarkable progress in many fields, especially YOLOv5, which has achieved certain breakthroughs in speed and accuracy. However, when dealing with small target detection in complex road traffic scenarios, it still faces the dual dilemmas of insufficient recognition accuracy and difficulty in balancing real-time performance. The root causes of these challenges lie in: the insignificance of small target features, the variability of environmental factors, and the high requirements for real-time processing. Therefore, exploring new methods that can improve detection performance in this specific scenario has become an urgent task.

[0004] Among them, the prior art: Wang et al.

[11] addressed the problem of small target missed detection prone to occur in multi-scale object detection of YOLOv5s; Zhao et al.

[12] proposed a bidirectional feature fusion structure. Through the reconstruction of the neck network structure of YOLOv3, the cross-level connection of multi-scale features was achieved, and the algorithm accuracy was improved by 4.0% compared with YOLOv3; Liu et al.

[13] constructed the TSingNet architecture for traffic sign detection, effectively solved the semantic gap between different scales, and fully integrated rich context features, and finally achieved a significant improvement in detecting small and obstructive traffic signs; Tian et al.

[14] proposed the AIoU loss function, which was used to emphasize difficult objects during model training, thereby promoting bounding box regression and significantly improving the learning ability of the model and the overall accuracy of image feature detection. In addition, during the improvement process, not only the improvement of model accuracy should be considered, but also the real-time performance of the algorithm should be ensured to make it applicable to complex road traffic backgrounds; Bie

[15] et al. proposed the Ghost module in the task of improving the real-time performance of the algorithm to reduce model parameters and improve the detection speed, so as to be applied to mobile terminal devices. To sum up, the key to solving the problems of dense objects and small objects lies in methods such as multi-scale detection and improved loss functions. However, at the current stage, there are still many challenges in object detection under complex backgrounds, and further in-depth research is needed to better meet the actual needs [16-20].

[0005] To solve the above problems, based on the YOLOv5s algorithm, this application improves and proposes a YOLOv5s-FGS (Four Layers+Ghost-ECA+SIOU) algorithm. The specific functions to be achieved are as follows:

[0006] (1) To enrich the feature information of small targets and improve their saliency, it aims to increase the shallow detection layer and upgrade the three-scale detection to four-scale to improve the network's detection ability for small targets.

[0007] (2) To achieve both model accuracy and speed and prevent excessive parameters, the Ghost module is introduced in the Neck layer, and a cascaded structure of the Ghost module and the attention mechanism ECA is constructed to extract key feature information. The original C3 module is replaced with the improved C3Ghost module to ensure a certain accuracy while reducing model parameters and improve the operation speed of the algorithm.

[0008] (3) To solve the problem that dense objects block each other and there are large changes in angles and directions between object detections in road traffic, the SIOU loss function is introduced to replace the original CIOU, thereby considering the angle loss (Anglecost) between regressions and improving the detection accuracy of the network. Summary of the Invention

[0009] The object of the present invention is to: realize the enrichment of the feature information of small targets, achieve the balance between model accuracy and speed, prevent excessive parameters, and solve the problem of mutual occlusion of dense targets. The present invention provides a small target detection algorithm model for complex traffic environments.

[0010] In order to achieve the above object, the present invention specifically adopts the following technical solutions:

[0011] A small target detection algorithm model for complex traffic environments, including the YOLOv5 algorithm model, characterized in that: it further includes a multi-scale detection module constructed based on the YOLOv5 algorithm model, which is used to upgrade the three-scale detection to four scales, so as to enrich the feature information of small targets;

[0012] The Ghost and ECA module cascade structure includes a Ghost module. The Ghost module is used to solve the problem of feature map redundancy during the training of the network, avoid unnecessary features from being repeatedly extracted, and use linear operations to supplement new feature maps, and splice the two groups of feature maps obtained on a specified dimension;

[0013] The function optimization module uses CIoU

[29] to calculate the bounding box regression localization loss, and adopts the following formula:

[0014]

[0015] In the formula: σ is the Euclidean distance, o t , o respectively represent the center points of the true box and the predicted box, and d represents the diagonal length.

[0016] Further, the multi-scale detection module includes a 20×20 feature layer and an introduced 160×160 feature layer. The 160×160 feature layer is obtained by doubling the upsampling of the 80×80 feature layer and fusing it with a newly constructed layer;

[0017] The 40×40 feature layer is obtained by doubling the upsampling of the 80×80 feature layer and fusing it with a newly constructed layer.

[0018] Further, the Ghost module is a phased convolution calculation module, which uses standard convolution to generate partial feature maps, where:

[0019] The operation of standard convolution in any convolutional layer can be described as:

[0020] y∈R n*h′*w′ =X*f + b

[0021] In the formula: X∈R c*h*w, where X is the input image, f is the convolution operation, c, h, and w respectively represent the number of channels, height, and width of the feature map, b is the bias term, n is the number of convolution kernels, h′ is the height of the feature map after standard convolution, and w′ is the width of the feature map after standard convolution;

[0022] Based on the standard convolution, a linear operation is performed to obtain the residual feature map:

[0023]

[0024] In the formula: y n is the i-th feature map after standard convolution, represents the i-th linear operation performed by the convolution kernel.

[0025] Furthermore, the number of channels of the feature map obtained by splicing the two parts of feature mapping is the same as that obtained by standard convolution. However, the convolution of the former greatly reduces the computational amount of convolution. Among them, the convolution kernel is k*k, the linear operation kernel is d*d, and the number of grouped convolutions is s. The computational amount of standard convolution is expressed as:

[0026] S = n * h′ * w′ * c * k * k.

[0027] Furthermore, during the GhostConv process, if m feature maps obtained by standard convolution are obtained by the original method (n = m * s), then its computational amount is:

[0028]

[0029] Furthermore, the ratio of the computational amount of standard convolution to the convolution after improvement by the Ghost module is:

[0030]

[0031] Furthermore, the function optimization module also includes SIoU

[31] , which includes a positioning loss function for bounding box regression. The formula is as follows:

[0032]

[0033] where Δ is the distance loss and Ω is the shape loss.

[0034] Furthermore, the distance loss is calculated using the following formula:

[0035]

[0036] In the formula: x t and x are the abscissas of the centers of the true box and the predicted box, y t and y are the ordinates of the centers of the true box and the predicted box, C w and C hrepresent the width and height of the minimum bounding rectangles of the two frames respectively, and Λ is the angular loss.

[0037] Furthermore, the angular loss is calculated using the following formula:

[0038]

[0039] where the angles between the line connecting the centers of the predicted box and the ground truth box and the horizontal and vertical directions are defined as α and β respectively. If then α is preferentially minimized, otherwise β is preferentially minimized. Different from the definition of the distance loss, C h is the height difference between the two centers, and σ is the distance between the two centers.

[0040] Furthermore, the shape loss is calculated using the following formula:

[0041]

[0042] where w and h are the width and height of the predicted box respectively, w t and h t represent the width and height of the ground truth box respectively.

[0043] Compared with the prior art, the beneficial effects of the present invention are:

[0044] In the present invention, through ablation experiments conducted on two datasets, the influence degree of each module on the accuracy improvement of the algorithm is analyzed, and the effectiveness of the improved module is verified. In addition, through comparative experiments with other mainstream algorithms and comparative tests in different complex road scenarios, the algorithm of the present application shows excellent performance, superior to the original YOLOv5s model, confirming its generalization applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is the multi-scale improvement structure diagram of the YOLOv5 network of the present invention;

[0046] Figure 2 is the comparison diagram of the standard convolution and the convolution of the Ghost module of the present invention;

[0047] Figure 3 is the cascaded structure of the Ghost and ECA modules of the present invention;

[0048] Figure 4 is the C3Ghost structure of the present invention;

[0049] Figure 5 is the diagram of the ground truth box and the predicted box of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0051] A small target detection algorithm model for complex traffic environments provided in this embodiment is mainly used to provide a new algorithm model that realizes rich feature information of small targets, balances model accuracy and speed, prevents excessive parameters, and solves the mutual occlusion of dense targets, and provides the following technical solutions, which will be described in detail below with reference to Figures 1-5 make a detailed description:

[0052] In complex road scenarios, due to the relatively large distances between traffic participants or traffic congestion, the pixel values in the image account for a relatively small proportion or the targets are occluded from each other. The original YOLOv5 algorithm cannot detect these targets well. To solve the potential traffic safety hazards caused by missed detection and false detection of small targets in road scenarios, this application improves based on the original structure, upgrades the three-scale detection to four scales, so as to enrich the feature information of small targets and facilitate the detection of smaller targets.

[0053] Introduce a feature layer of 160×160. This new feature layer is obtained by doubling the upsampling of the 80×80 feature layer and fusing it with the newly constructed 160×160 feature layer, aiming to improve the detection ability for smaller targets. Then, double upsampling is performed again through the 40×40 feature layer and fused with the 80×80 feature layer for the detection of small targets. In the original algorithm, the 20×20 feature layer is used for the detection of large targets, and then double upsampling is performed and fused with the 40×40 feature layer to achieve the detection of medium targets. This multi-scale improvement idea brings higher detection accuracy to the algorithm and makes it more suitable for complex traffic scenarios with large size variations. The improved structure is as Figure 1 shown in.

[0054] Among them, to explain and construct the cascaded structure of the Ghost and ECA modules, although adding shallow layers improves the detection accuracy of the network in practice, the model parameters and floating-point computational volume also increase accordingly, affecting the detection speed of the network. To reduce the model parameters and improve the running speed of the network, lightweight modules are used to optimize the network structure.

[0055] During the standard convolution process of YOLOv5s, the extracted features show a certain degree of repetition. To solve the problem of feature map redundancy during the training of the network and avoid unnecessary features being repeatedly extracted, the Ghost module is introduced.

[0056] The Ghost module is a phased convolution calculation module. First, it uses standard convolution to generate partial feature maps. On this basis, to keep the final number of channels constant, it uses linear operations to supplement new feature maps. Finally, the two sets of feature maps obtained are concatenated in the specified dimension

[26] , as Figure 2 shown

[0057] The operation of performing standard convolution in any convolutional layer can be described as:

[0058] y ∈ R n*h′*w′ = X * f + b

[0059] where: X ∈ R c*h*w , X is the input image, f is the convolution operation, c, h, and w represent the number of channels, height, and width of the feature map respectively, b is the bias term, n is the number of convolutional kernels, h′ is the height of the feature map after standard convolution, and w′ is the width of the feature map after standard convolution;

[0060] Based on the standard convolution, a linear operation is performed to obtain the remaining feature maps:

[0061]

[0062] where: y n is the i-th feature map after standard convolution, represents the j-th linear operation performed by the convolutional kernel. where: y n is the n-th feature map after standard convolution, represents the m-th linear operation performed by the convolutional kernel;

[0063] The number of channels of the feature map obtained by concatenating the two parts of the feature maps is the same as that obtained by standard convolution. However, the convolution of the former greatly reduces the computational complexity of the convolution. Among them, the convolutional kernel is k * k, the kernel of the linear operation is d * d, and the number of grouped convolutions is s. The computational complexity of the standard convolution is expressed as:

[0064] S = n * h′ * w′ * c * k * k;

[0065] Assume that in the GhostConv process, m feature maps obtained by standard convolution are obtained by the original method (n = m * s), then its computational complexity is:

[0066]

[0067] Then the ratio of the computational complexity of the standard convolution to the convolution after the improvement of the Ghost module is:

[0068]

[0069] Therefore, the computational cost of the Ghost module is about 1 / s times that of the standard convolution. As more feature maps are generated by linear operations, the computational cost becomes less, and the acceleration effect is better. However, the detection accuracy will also decrease accordingly, that is, feature maps with the same accuracy as the standard convolution cannot be obtained. This application uses the Ghost module not simply to improve the detection speed, but also to take into account the detection accuracy of the network. Therefore, an ECA attention mechanism layer is added to the Ghost module to construct a cascaded structure [27-28]. First, global average pooling is performed on some of the feature maps after standard convolution to obtain an un-dimension-reduced feature map with a size of 1*1*m. One-dimensional convolution is performed on the obtained feature map to achieve cross-channel information interaction. Subsequently, the Sigmoid function is used to generate the weight ratio of each channel, and the weight parameter is multiplied by the original channel feature map. The network constructed by this module is more likely to extract discriminative features of the image from the channel dimension, effectively screening out the key information that contributes more to the final target prediction result, thereby obtaining high-quality feature maps. Figure 3 It is a cascaded structure of Ghost and ECA modules.

[0070] The improved GhostBottleneck is combined with the C3 module to form the C3Ghost module to replace the original C3 module and placed in the backbone layer (verified by the above experiments, placing it only in the Neck layer has a better detection effect than replacing all C3 modules in the backbone layer and Neck layer of the network). The improved Ghost module and C3Ghost are as Figure 4 shown. Therefore, the improved network structure not only realizes the lightweight of the network, but also can extract important feature information, ensuring the accuracy of the network while improving the detection speed.

[0071] YOLOv5s uses CIoU

[29] to calculate the bounding box regression localization loss, which fully considers distance, overlap, and aspect ratio. The formula is shown in (6).

[0072]

[0073] In the formula: σ is the Euclidean distance, o t , o represent the center points of the ground truth box and the predicted box respectively, and d represents the diagonal length.

[0074] As can be seen from the formula, CIoU introduces the aspect ratio loss, avoiding the situation where the intersection over union (IoU) is the same when the centers of the bounding boxes coincide in DIoU

[30] , making it impossible to distinguish. However, CIoU overly relies on the aggregation of bounding box regression metrics and does not consider the possible mismatched directions between the ground truth boxes and the predicted boxes. This leads to difficulties in convergence during training, thereby reducing the model performance. In complex road scenarios, due to their high dynamics, the angles and directions of the objects to be detected vary greatly. Since the CIoU loss function cannot handle the regression of bounding boxes well in this case, this application introduces SIoU

[31] as the localization loss function for bounding box regression, and its formula is shown below. Where Δ is the distance loss and Ω is the shape loss.

[0075]

[0076] The distance loss Δ is shown in the following formula, and the distance loss is defined in combination with the angle loss Λ. The parameter ρ x , ρ y and γ are as shown in the following formula.

[0077]

[0078] In the formula: x t and x are the abscissas of the centers of the ground truth box and the predicted box, and y t and y are the ordinates of the centers of the ground truth box and the predicted box, C w and C h represent the widths and heights of the minimum enclosing rectangles of the two boxes respectively, and Λ is the angle loss.

[0079] The calculation formula of the angle loss Λ is shown below. SIoU defines the angles between the line connecting the centers of the predicted box and the ground truth box and the horizontal and vertical directions as α and β respectively. If then α is preferentially minimized, otherwise β is preferentially minimized. Different from the definition of the distance loss, C h is the height difference between the two centers, and σ is the distance between the two centers.

[0080]

[0081] The calculation formula of the shape loss Ω is shown below. The calculation methods of the parameters w w and w h are as shown in the following formula. Where w and h are the width and height of the predicted box respectively, and w t and h t represent the width and height of the ground truth box respectively. When the value of the parameter θ is 1, the border shape will be immediately optimized, restricting the free movement of the border.

[0082]

[0083] Based on the research of CIOU, SIoU introduces the angular loss, specifically as Figure 5 shown. For training in road traffic scenarios, this method effectively reduces the degrees of freedom, enabling the prediction box to quickly move to the nearest axis, thus accelerating the training convergence and optimizing the effect of bounding box regression.

[0084] To verify the effectiveness of the improved YOLOv5s algorithm in complex road traffic scenarios, this application uses the publicly available KITTI dataset as the main source for algorithm testing and uses part of the UA-DETRAC public dataset for auxiliary experimental verification.

[0085] Complex road traffic scenarios usually include busy intersections, multi-lane highways, urban streets, and rural roads. In addition, the complexity is not only reflected in the road level but also involves changing environmental conditions, such as poor lighting conditions, bad weather, or complex road traffic situations caused by mutual occlusion among traffic participants. The KITTI dataset used in this application fully collects multiple real and complex traffic images such as urban roads, campuses, and highways. The UA-DETRAC dataset provides further evidence for our verification results. Its dataset includes images under different weather conditions and different occlusion degrees, with rich and complex scenarios, which is suitable for the algorithm verification application of this application.

[0086] To reduce parameters and fit a wider range of road scenarios, this application re-integrates the labels of the two datasets. In the KITTI dataset, it is divided into three categories: "Car", "Pedestrian", and "Cyclist". The target classification results are shown in Table 1. In the UA-DETRAC dataset, "Van" and "Others" are merged into the "Van" category, thus dividing the data labels into three categories: "Car", "Bus", and "Van". Table 2 shows the number of targets after processing.

[0087] Table 1 Number of targets in the KITTI dataset after processing

[0088] Tab.1 Number of targets in the KITTI dataset after processing

[0089]

[0090] Table 2 Number of targets in the UA-DETRAC dataset after processing

[0091] Tab.2 Number of targets in the UA-DETRAC dataset after processing

[0092]

[0093] Through the above improvements, the detection accuracy of the algorithm of this application has been improved by 1.9% and 3.9% on the KITTI and UA-DETRAC datasets respectively, while the detection speeds have reached 70 and 64 FPS respectively. Through ablation experiments conducted on the two datasets, the influence degree of each module on the accuracy improvement of the algorithm was analyzed, and the effectiveness of the improved module was verified. In addition, through comparative experiments with other mainstream algorithms and comparative tests in different complex road scenarios, the algorithm of this application has shown superior performance, better than the original YOLOv5s model, confirming its generalization applicability. Therefore, the algorithm of this application has obtained higher detection accuracy while maintaining the advantages of speed and size. Compared with other mainstream algorithms, the YOLOv5s-FGS algorithm is more suitable for object detection in complex road traffic environments. In the future, with the integration of more innovative technologies and the in-depth exploration of algorithm optimization strategies, the research and application of small object detection in complex road traffic scenarios will continue to advance, which can not only further improve the safety and reliability of autonomous driving technology, but also provide a solid technical foundation for the construction of smart cities.

[0094] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.

Claims

1. A small target detection algorithm model for complex traffic environments, including a YOLOv5 algorithm model, characterized in that: It also includes a multi-scale detection module based on the YOLOv5 algorithm model, which is used to upgrade three-scale detection to four-scale, thereby enriching the feature information of small targets; The Ghost and ECA module cascade structure includes a Ghost module, wherein the Ghost module is used to solve the feature map redundancy phenomenon in the network training process, avoid unnecessary features from being repeatedly extracted, and use linear operations to supplement new feature maps, and concatenate the two sets of feature maps obtained in the specified dimension; The function optimization module uses CIoU[29] to calculate the bounding box regression positioning loss using the following formula: Where: σ is the Euclidean distance, o t,o They represent the center points of the real box and the predicted box respectively, and d represents the length of the diagonal.

2. The small target detection algorithm model for complex traffic environment according to claim 1, characterized in that: The multi-scale detection module includes a 20×20 feature layer and an introduced 160×160 feature layer, wherein the 160×160 feature layer is obtained by doubly upsampling the 80×80 feature layer and fusing it with the new construction; The 40×40 feature layer is obtained by upsampling the 80×80 feature layer by a factor of two and fusing it with the new construction.

3. The small target detection algorithm model for complex traffic environment according to claim 1, characterized in that: The Ghost module is a staged convolution calculation module, which uses standard convolution to generate some feature maps, where: The operation of performing a standard convolution in any convolutional layer can be described as: y∈R n*h′*w′ =X*f+b Where: X∈R c*h*w , X is the input image, f is the convolution operation, c, h, w represent the number of channels, height, and width of the feature map, respectively, b is the bias term, n is the number of convolution kernels, h is the height of the feature map after standard convolution, and w′ is the width of the feature map after standard convolution; Perform linear operations on the basis of standard convolution to obtain the remaining feature map: Where: y n is the th feature map after standard convolution, Represents the linear operation performed by the convolution kernel.

4. The small target detection algorithm model for complex traffic environment according to claim 1, characterized in that: The number of channels of the feature map obtained by concatenating the two feature maps is consistent with that obtained by standard convolution, but the former convolution greatly reduces the amount of convolution calculation. The convolution kernel is k*k, the linear operation kernel is d*d, the number of grouped convolutions is s, and the amount of calculation of standard convolution is expressed as: S = n*h′*w′*c*k*k.

5. The small target detection algorithm model for complex traffic environment according to claim 1 is characterized by: In the GhostConv process, m feature maps obtained by standard convolution are obtained by the original method (n = m*s), and the amount of calculation is:

6. The small target detection algorithm model for complex traffic environment according to claim 1, characterized in that: The ratio of the amount of convolution calculation between the standard convolution and the improved convolution of the Ghost module is:

7. The small target detection algorithm model for complex traffic environment according to claim 1 is characterized by: The function optimization module also includes SIoU [31], which includes a localization loss function as a bounding box regression, and the formula is as follows: Where Δ is the distance loss and Ω is the shape loss.

8. The small target detection algorithm model for complex traffic environment according to claim 1 is characterized in that: The distance loss is calculated using the following formula: Where: x t and x is the horizontal coordinate of the center point of the real box and the predicted box, y t and y are the ordinates of the center points of the real box and the predicted box, C w and C h They represent the width and height of the minimum circumscribed rectangle of the two boxes respectively, and Λ is the angle loss.

9. The small target detection algorithm model for complex traffic environment according to claim 1, characterized in that: The angle loss calculation adopts the following formula: The angles between the line connecting the center points of the predicted box and the real box and the horizontal and vertical directions are defined as α and β respectively. Then α is minimized first, otherwise β is minimized first. Different from the definition of distance loss, C h is the height difference between the two center points, and σ is the distance between the two center points.

10. The small target detection algorithm model for complex traffic environment according to claim 1, characterized in that: The shape loss calculation adopts the following formula: Where w and h are the width and height of the prediction box respectively. t and h t Represent the width and height of the real box respectively.