Road traffic vehicle target detection method based on deep learning
By improving the YOLOv10 network structure, introducing the SwinTransformer module and attention mechanism, and combining it with an innovative loss function, the problem of insufficient detection accuracy of YOLOv10 in complex traffic scenarios is solved, and more efficient vehicle target detection is achieved.
Patent Information
- Application Number
- CN202510809399.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
AI Technical Summary
YOLOv10's vehicle target detection accuracy in complex traffic scenarios is insufficient, especially in the detection of small targets in long-distance camera images. The limited receptive field leads to weak global semantic information modeling capabilities, and there are false detections and missed detections in dense traffic or occluded environments.
The YOLOv10 network structure is improved by introducing the SwinTransformer module to replace the SwinStage, adding the SCEC-Swin model, combining the PSELA attention mechanism and the C2f-BG module, using the DyMLPHead detection head, and introducing the InnerATFLIoU loss function to improve the model's detection performance in complex environments.
The model's detection accuracy and robustness in complex traffic scenarios are significantly improved, its small target detection capability is enhanced, the computational complexity is reduced, and its adaptability is stronger.
Smart Images

Figure CN120708026A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning vehicle detection technology, and in particular to a road traffic vehicle target detection method based on deep learning. Background Art
[0002] With the accelerating pace of urbanization and the rapid development of intelligent transportation systems, the automatic detection and identification of vehicles has become a key research topic in the field of intelligent transportation. Efficient and accurate vehicle target detection technology not only plays a key role in applications such as traffic flow monitoring, intelligent parking management, and autonomous driving, but also provides strong support for building smart cities. However, complex traffic scenarios, such as occlusion, varying lighting conditions, and multi-scale vehicles, pose significant challenges to traditional detection algorithms.
[0003] In recent years, deep learning-based object detection algorithms have been widely used. The YOLO series of algorithms, with their end-to-end architecture, efficient detection speed, and good detection accuracy, has become a mainstream approach. With the release of YOLOv10, it has achieved a better balance between speed and efficiency, becoming a representative lightweight detection framework. However, when faced with multi-scale, dense, and complex vehicle targets in traffic scenarios, YOLOv10 still faces certain detection accuracy bottlenecks. For example, in images captured by long-range cameras, the sizes of vehicle targets vary significantly, and YOLOv10 still pays insufficient attention to small targets during feature extraction. YOLOv10 uses a convolutional neural network as its backbone for feature extraction, but its limited receptive field makes it less capable of modeling global semantic information. In environments with dense traffic or occlusion, the boundaries between targets are blurred, and false detections and missed detections still occur. Summary of the Invention
[0004] Purpose of the invention: In response to the technical problem that YOLOv10 has insufficient detection accuracy in complex traffic scenarios when using n weights for target detection model training, the present invention provides a road traffic vehicle target detection method based on deep learning. Based on the YOLOv10 network foundation, it is proposed to improve the SwinStage module in SwinTransformer, introduce the SCEC-Swin model to replace the SwinStage model, improve the accuracy, reduce the calculation amount of the model, and reduce redundant calculations; and add the attention mechanism PSELA, add the innovative C2f-BG layer to the neck, and propose an innovative DyMLPHead and a new loss function InnerATFLIoU, which can effectively improve the model accuracy, make it more suitable for target detection in complex environments, and effectively solve the above problems.
[0005] Technical solution: The present invention provides a road traffic vehicle target detection method based on deep learning, comprising the following steps:
[0006] Step 1: Select the KITTI dataset and randomly divide it into training set, validation set, and test set according to the proportion;
[0007] Step 2: Replace the YOLOv10 backbone network with the SwinTransformer module, and introduce the SCEC-Swin module to replace the SwinStage module in the SwinTransformer module;
[0008] Step 3: Introduce the PSELA attention mechanism after the SPPF layer of the YOLOv10 backbone network, and use the feature enhancement technology of one-dimensional convolution and group normalization to achieve lightweight implementation;
[0009] Step 4: Improve the neck structure of YOLOv10: Introduce C2f-BG to replace the C2f module. Combining the advantages of AKConv and GAM in the C2f module, it better improves the algorithm's ability to process features while reducing the number of model parameters.
[0010] Step 5: Introduce the detection head DyMLPHead to replace the original YOLOv10 detection head, and introduce the loss function InnerATFLIoU;
[0011] Step 6: Use the divided data set to train the target detection model constructed in S2-S5;
[0012] Step 7: Use the trained object detection model to perform object detection on the test set and evaluate the performance of the model.
[0013] Furthermore, the SCEC-Swin module in step 2 includes an SCEC branch and an STB branch, and the specific process is as follows:
[0014] Step 2.1: The input image first passes through the SCEC branch, which consists of two branches. In the first branch, a global average pooling operation is performed on the image, and a 1D convolution is used to model the local channel relationship in the channel dimension:
[0015] I1=Conv1D(AvgPool(I)) (1)
[0016] I1'=I·σ(I1) (2)
[0017] Among them, Conv1D represents a one-dimensional convolution operation, which is used to model the dependency between channels; AvgPool(I) represents the average pooling operation, σ represents the Sigmoid activation function, I1' represents the output of the first branch, and I represents the input image;
[0018] In the second branch, the feature map is pooled along the X and Y directions, and dilated convolution is used to indirectly enhance the feature representation ability of the facial area by increasing the receptive field. Given an input I, the maximum pooling operation is performed on each channel using the two spatial ranges (H, 1) or (1, W) of the pooling kernel along the horizontal and vertical coordinates, respectively. The horizontal coordinate pooling process is expressed as follows using formula (3):
[0019] I h =MaxPool(I,1) (3)
[0020] Among them, MaxPool(I,1) means performing the maximum pooling operation on all width values at the same height position;
[0021] The vertical coordinate pooling process is expressed as follows using formula (4):
[0022] I w =MaxPool(1,I) (4)
[0023] Among them, MaxPool(1,I) means performing the maximum pooling operation on all height values at the same width position;
[0024] Aggregate features along two spatial directions to generate a pair of direction-aware feature maps I h and I w , then I h and I w Cascade in a specific direction and obtain the intermediate feature map I by expanding the convolution layer f :
[0025] I f =ReLu(BN(DC([I h ,I w ]))) (5)
[0026] Among them, DC represents the expanded convolution operation, BN represents normalization, and ReLu represents the activation function;
[0027] Then I f Split into two independent tensors along the spatial dimension and Two additional unfolded convolution operations transform the two tensors into tensors with the same number of channels:
[0028]
[0029] Where, DC h and DC w Represent the extended convolution operations in the horizontal and vertical directions respectively, and Sigmoid represents the activation function;
[0030] Expand the output separately and As the attention weight, the final output is expressed as follows using formula (8):
[0031]
[0032] Step 2.2: The STB branch consists of four key components, including LN, W-MSA, MLP, and SW-MSA. LN calculates the normalized mean and variance of a specific dimension of a single sample. MLP establishes a nonlinear relationship between input features and output. W-MSA divides the feature map into windows and calculates self-attention within each window. SW-MSA simulates the correlation between regions by shifting windows, which can be expressed as follows using Formula (9):
[0033]
[0034] Where I represents the input image, i represents the current block, and I i , I i 'and I i+1 Represent the outputs of MLP, W-MSA and SW-MSA, STB output Indicates the output of the STB branch;
[0035] Step 2.3: The total output is expressed as follows using formula (10):
[0036] SCTB output =STB output +SCEC output (10)
[0037] Furthermore, in step 3, the PSELA attention mechanism process is as follows:
[0038] The input feature map first passes through a convolutional layer for preliminary feature extraction. The extracted feature map is then fed into the ELA module, where the feature representation capability is enhanced through the local attention mechanism. Within the ELA, the input feature map is average pooled horizontally and vertically to obtain compressed feature maps with sizes of C×H×1 and C×1×W, respectively. 1D convolution is then used for feature extraction, followed by normalization and sigmoid activation to generate spatial attention weights in two directions. The two attention maps are element-wise multiplied with the original feature map to obtain a feature map enhanced with spatial attention.
[0039] Secondly, the feature map enhanced by the ELA module will continue to pass through two convolutional layers to further extract feature information. In order to retain the information of the initial input, a residual connection is used to splice the original input with the current feature map to achieve feature fusion.
[0040] Finally, the fused feature map is integrated and reduced in dimension through a convolutional layer to obtain the final output result.
[0041] Furthermore, in step 4, the C2f-BG module replaces the Bottleneck layer in the C2f module with an AG-Bottleneck module and adds a GAM module after the first Conv in the C2f module. The AG-Bottleneck module replaces the Conv in the Bottleneck layer with an AKConv module with any sampling shape and any number of parameters, and adds a GAM module after the second AKConv, finally obtaining the AG-Bottleneck module.
[0042] Furthermore, in step 5, the detection head DyMLPHead is specifically as follows:
[0043] First, the input features are max-pooled and average-pooled and then merged. The merged features are fed into the Conv3×3 convolutional layer to extract local spatial context information.
[0044] The convolutional features use the GELU activation function to enhance the nonlinear expression ability, and finally the HardSigmoid activation function is used to limit the output value between 0 and 1 to generate the weight of each feature;
[0045] Secondly, the input features are sent to the Index to extract spatial position information, and the Offset and Sigmoid outputs are generated through 3×3 convolution;
[0046] Finally, the features are globally averaged pooled to obtain the channel description vector, which is fed into the multi-layer perceptron MLP. The MLP further models the channel description, including two fully connected layers FC and a GELU activation layer, to obtain four parameters: α 1 , β 1 , α 2 , β 2 ;
[0047] Normalize: Normalize the obtained parameters and add them to the vector [1,0,0,0] to enhance specific dimensions;
[0048] Final output: The feature maps processed by multiple branches are integrated into the final output.
[0049] Furthermore, the InnerATFLIoU loss function is as follows:
[0050] InnerIoU loss function L InnerIoU The specific expression is:
[0051] LInnerIoU =1-IOU inner (11)
[0052]
[0053] Among them, IOU inner It represents the inscribed IOU, which is an improved IoU indicator; inter represents the intersection area of the inscribed area between the predicted box and the true box, and union represents the joint area formed by the predicted box and the true box;
[0054]
[0055] Where, and (x c ,y c ) represent the center coordinates of the auxiliary bounding box and the auxiliary prediction box respectively; w gt ,h gt w and h represent the length and width of the real box and the predicted box respectively; ratio represents the adjustable scale factor, ranging from 0.5 to 1.5. and (b l ,b t ), (b l ,b b ), (b r ,b t ), (b r ,b b ) represent the four vertex coordinates of the auxiliary bounding box and the auxiliary prediction box respectively;
[0056] When obtaining the ATFL loss function, we use the classic cross entropy loss function L BCE Starting from, its expression is:
[0057] L BCE =-(ylog(p)+(1-y)log(1-p)) (18)
[0058] Where p represents the predicted probability and y represents the true label; Formula (18) is simplified to:
[0059] L BCE =-log(p t ) (19)
[0060]
[0061] The focal loss function is obtained by adjusting the factor (1-p t ) γ To solve the problem of sample imbalance, its definition and calculation process are shown in formula (21):
[0062] FL(pt )=-(1-p t ) γ logp t (twenty one)
[0063] Where p t Represents the probability that the model predicts that the sample belongs to the true category; γ represents an adjustable parameter;
[0064] To solve the above problems, the threshold focus loss function TFL is proposed as shown in formula (22):
[0065]
[0066] Where γ is an adjustable parameter and η is a hyperparameter.
[0067] The adaptive threshold focal loss function ATFL adaptively adjusts γ and η.
[0068]
[0069] The final formula of the ATFL loss function is:
[0070]
[0071] The InnerATFLIoU loss function is defined as follows:
[0072] L InnerATFLIoU =L ATFL +L InnerIoU (25)
[0073] The InnerATFLIoU loss function significantly improves detection accuracy and robustness by adaptively focusing on key tasks and accurately optimizing the overlapping area of the target box.
[0074] Beneficial effects:
[0075] Based on the YOLOv10 network structure, this paper proposes an improved SwinStage model in SwinTransformer as the backbone network. By introducing the innovative SCEC-Swin model instead of the SwinStage model, computation is more efficient and local channel dependencies can be captured. To further improve the model's detection performance in complex environments, we introduce an innovative convolutional attention architecture, PSELA, after the SPPF layer. PSELA utilizes feature enhancement techniques such as one-dimensional convolution and group normalization to accurately locate feature regions without dimensionality compression, while achieving a lightweight implementation.
[0076] To improve the model's adaptability and performance at different scales, this paper proposes a new detection head, DyMLPHead, to replace the existing one. DyMLPHead employs feature weighting calculations based on scale, space, and task branches, achieving multi-dimensional enhancement of the input feature map. It also introduces an innovative InnerATFLIoU loss function to replace the traditional CIoU loss, improving the detection of small and easily confused objects and enabling more accurate and robust object localization and classification.
[0077] Based on the neck structure in YOLOv10, the innovative C2f-BG is introduced to replace the C2f module. The C2f module combines the advantages of AKConv and GAM, which better improves the algorithm's ability to process features while reducing the number of model parameters.
[0078] Through the above series of improvements, the target detection method proposed in this paper can effectively improve the detection accuracy of the model in complex traffic scenarios, providing strong support for the practical application of target detection algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 This is a diagram of the improved YOLOv10 network structure in an embodiment of the present invention;
[0080] Figure 2 This is a diagram of the improved SCEC-Swin network structure in an embodiment of the present invention;
[0081] Figure 3 FIG1 is a diagram of an improved PSELA network structure according to an embodiment of the present invention;
[0082] Figure 4 This is a diagram of the AKConv network structure in an embodiment of the present invention;
[0083] Figure 5 This is a diagram of the GAM network structure in an embodiment of the present invention;
[0084] Figure 6 This is a structural diagram of the improved bottleneck-AG module in an embodiment of the present invention;
[0085] Figure 7 FIG2 is a structural diagram of an improved C2f-BG according to an embodiment of the present invention;
[0086] Figure 8 This is a structural diagram of the improved DyMLPHead in an embodiment of the present invention;
[0087] Figure 9 This is the unimproved YOLOv10 training result picture in the embodiment of the present invention.
[0088] Figure 10This is an unimproved YOLOv10 detection dataset image in an embodiment of the present invention;
[0089] Figure 11 This is an improved YOLOv10 training result picture in an embodiment of the present invention. Figure 12 This is an image of the improved YOLOv10 detection dataset in an embodiment of the present invention. DETAILED DESCRIPTION
[0090] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.
[0091] The present invention discloses a road traffic vehicle target detection method based on deep learning, which specifically includes the following steps:
[0092] Step 1: Select the KITTI dataset and randomly divide it into training set, validation set, and test set according to the proportion;
[0093] The KITTI dataset contains 3,712 training images and 3,769 validation images. It serves as the basis for vehicle detection and is used for training the YOLO target detection series.
[0094] Step 2: Improve the backbone network of YOLOv10: Based on the YOLOv10 network foundation, use SwinTransformer to replace the backbone part, and replace the SwinStage module in the SwinTransformer module with the SCEC-Swin module. Through the convolution operation, the ability to obtain global information is enhanced, thereby improving the detection performance of the YOLOv10 network model. Figure 2 As shown, the innovative SCEC-Swin module includes a SCEC branch and a STB branch, specifically;
[0095] The SCEC branch consists of two branches to obtain global and local fine-grained directional features. In the first branch, a global average pooling operation is performed on the image to better preserve the global features. Using 1D convolution to model local channel relationships (cross-channel information fusion) on the channel dimension does not reduce the dimension. This process can be expressed as follows using formulas (1) and (2):
[0096] I1=Conv1D(AvgPool(I)) (1)
[0097] I1'=I·σ(I1) (2)
[0098] Among them, Conv1D represents a one-dimensional convolution operation (acting on the channel dimension) used to model the dependency between channels; σ represents the Sigmoid activation function, and I1' represents the output of the first branch.
[0099] In the second branch, the feature map is max-pooled along the X and Y directions to maintain the directionality of the image. At the same time, dilated convolution is used to indirectly enhance the feature representation of the facial region by increasing the receptive field. Specifically, given the input I, we perform max-pooling on each channel using the two spatial ranges (H, 1) or (1, W) of the pooling kernel along the horizontal and vertical coordinates, respectively. Therefore, the horizontal coordinate pooling process is expressed as follows using formula (3):
[0100] I h =MaxPool(I,1) (3)
[0101] Among them, MaxPool(I,1) means performing the maximum pooling operation on all width values at the same height position.
[0102] Similarly, the vertical coordinate pooling process is expressed as follows using formula (4):
[0103] I w =MaxPool(1,I) (4)
[0104] Among them, MaxPool(I,1) means performing the maximum pooling operation on all height values of the same width position.
[0105] The above two transformations aggregate features along two spatial directions respectively. They generate a pair of direction-aware feature maps I h and I w , to capture the directional characteristics of the image area. Then I h and I w Cascade in a specific direction and obtain the intermediate feature map I by expanding the convolution layer f , the feature map encodes spatial information in the horizontal and vertical directions. This process is expressed as follows using formula (5):
[0106] I f =ReLu(BN(DC([I h ,I w ]))) (5)
[0107] Among them, DC represents the unfolded convolution operation.
[0108] Then I f Split into two independent tensors along the spatial dimension and The other two expanded convolution operations transform the two tensors into tensors with the same number of channels, which are expressed as follows by Equations (6) and (7):
[0109]
[0110] Where, DC h and DC w Represent the dilated convolution operations in the horizontal and vertical directions, respectively.
[0111] Then expand the output separately and As the attention weight. The final output of the module is expressed as follows using formula (8):
[0112]
[0113] The STB branch consists of four key components, including LN, W-MSA, MLP, and SW-MSA. These components work together to collect global facial features and local features. LN calculates the normalized mean and variance of a specific dimension of a single sample. MLP establishes a nonlinear relationship between input features and output. W-MSA divides the feature map into windows and calculates self-attention within each window. SW-MSA simulates the correlation between regions by shifting windows. The above process can be expressed as follows using formula (9):
[0114]
[0115] Where I represents the input image, i represents the current block, and I i , I i 'and I i+1 denote the outputs of MLP, W-MSA and SW-MSA respectively.
[0116] The total output is expressed as follows using formula (10):
[0117] SCTB output =STB output +SCEC output (10)
[0118] Step 3: Add the PSELA attention mechanism after the backbone network SPPF, and use the feature enhancement technology of one-dimensional convolution and group normalization to accurately locate the feature area without dimension compression, while achieving lightweight implementation. Figure 3 As shown, the PSELA attention mechanism process is as follows:
[0119] First, the input feature map passes through a convolutional layer for preliminary feature extraction. Next, the feature map is fed into the Efficient Local Attention (ELA) module, which enhances feature representation through a local attention mechanism. Within the ELA, the input feature map is average pooled in the horizontal (X) and vertical (Y) directions, resulting in compressed feature maps with dimensions of C×H×1 and C×1×W, respectively.
[0120] Then, feature extraction is performed through 1D convolution, followed by normalization and sigmoid activation to generate spatial attention weights in two directions. The two attention maps are element-wise multiplied with the original feature map to obtain a feature map enhanced by spatial attention.
[0121] Next, the enhanced feature map continues to pass through two convolutional layers to further extract higher-level feature information. At the same time, to preserve the initial input information, the model uses residual connections: concatenating the original input with the current feature map to achieve feature fusion, improving expressiveness and stability.
[0122] Finally, the fused feature map is integrated and reduced in dimension through a convolutional layer to obtain the final output result.
[0123] Step 4: In the neck structure of v10, the innovative C2f-BG is introduced to replace the C2f module. The C2f module combines the advantages of AKConv and GAM to better improve the algorithm's ability to process features while reducing the number of model parameters. The specific process of the C2f-BG module is as follows:
[0124] AKConv is a deformable convolution operation that can flexibly allocate the number of parameters to the convolution kernel and support arbitrary sampling shapes, thus breaking through the limitations of traditional convolution in fixed receptive fields and regular sampling. This design makes convolution more adaptable when dealing with different data distributions and target positions. By introducing a learnable offset, AKConv dynamically adjusts the sampling position of the convolution kernel, improving the accuracy of feature extraction. In addition, the number of parameters of AKConv can grow or shrink linearly, which supports both high-performance computing scenarios and lightweight model design. The sampling network used is a regular sampling grid with a rule of 3×3 convolution operation. Let R represent the sampling grid, then R is expressed as follows:
[0125] R={(-1,-1),(-1,0),…,(0,1),(1,1)}#(11)
[0126] However, AKConv requires an irregular sampling network, so it is designed to generate the convolution kernel P n The algorithm for the initial sampling coordinates of . Figure 4As shown in Figure 1, the initial sampling coordinates for convolution of any size are generated. Because irregular convolution rarely has a center point in terms of size, in order to adapt to the convolution size used, (0,0) is set as the sampling origin in the upper left corner of the algorithm. The corresponding convolution operation at position P0 is defined as follows:
[0127] Conv(P0)=∑ω×(P0+P n )#(12)
[0128] In the formula, ω represents the convolution parameter. Through a series of operations, AKConv solves the problem that irregular sampling coordinates cannot match the corresponding size of convolution operations.
[0129] AKConv first calculates an offset through convolution, which is used to dynamically adjust the original sampling position, thereby achieving a deformable convolution kernel shape to adapt to image features. The feature map is resampled based on the updated sampling position, and then undergoes feature reshaping, convolution, and normalization, with the output being activated using the SiLU function. Compared to fixed sampling methods, AKConv can flexibly adjust the sampling structure of each position based on image content, significantly enhancing the model's adaptability and expressiveness.
[0130] The structure of the GAM global attention mechanism is as follows Figure 5 As shown in the figure, F1 is the input feature map, F2 is the intermediate feature map, and F3 is the output feature map. GAM consists of a channel attention submodule (M c ) and the spatial attention submodule (M S ) composition, such as Figure 5 As shown in the figure, the channel module extracts key features through global average pooling and maximum pooling, and combines this with a shared multi-layer perceptron to enhance cross-dimensional spatial information interaction, thereby improving feature expressiveness. The spatial module uses convolution operations to fuse spatial information and generates a weight map to weight the original features, effectively achieving global spatial feature modeling.
[0131] like Figure 5 As shown, in the channel sub-attention module, the definition of the intermediate state is shown in formula (13), where M c is the channel attention map, Represents cascade.
[0132]
[0133] like Figure 5 As shown, in the spatial sub-attention module, the definition of the intermediate state is as shown in formula (14), where M S is the spatial attention map:
[0134]
[0135] Design of C2f-BG module: Design the bottleneck layer in the C2f module to obtain the AG-Bottleneck module, such as Figure 6 As shown in the figure, the Conv in the bottleneck layer is replaced with an AKConv convolution with any sampling shape and any number of parameters, and a GAM module is added after the second AKConv to obtain the AG-Bottleneck module.
[0136] like Figure 7 As shown in Figure 1, the bottleneck layer in the C2f module is replaced with the AG-Bottleneck module, and a GAM module is added after the first Conv in the C2f module to obtain the C2f-BG module. The C2f-BG module improves the network's feature processing capabilities and improves small object detection performance.
[0137] Step 5: Based on the neck structure in YOLOv10, the innovative DyMLPHead is introduced to increase the nonlinear capability of the model, enabling the model to learn more complex and advanced feature representations, such as Figure 8 The specific process is as follows:
[0138] First, the input features are subjected to maximum pooling and average pooling operations, and then merged. The merged features are sent to the Conv3×3 convolutional layer to extract local spatial context information.
[0139] The convolutional features use the GELU activation function to enhance their nonlinear expressiveness. Finally, the Hard Sigmoid activation function limits the output value to between 0 and 1, which is used to generate the weight of each feature. The resulting weight is multiplied element-by-element with the input feature to adjust the feature's importance.
[0140] Secondly, the input features are sent to the Index to extract the spatial position information, and the Offset and Sigmoid outputs are generated through 3×3 convolution.
[0141] Finally, the features are subjected to global average pooling (GAP) to obtain a channel description vector. This is fed into a multi-layer perceptron (MLP), which further models the channel description, including two fully connected layers (FC) and a GELU activation layer. Four parameters are obtained: α 1 , β 1 , α 2 , β 2 .
[0142] Normalize: Normalize the obtained parameters and add them to the vector [1, 0, 0, 0] to enhance specific dimensions.
[0143] Final output: The feature maps processed by multiple branches are integrated into the final output.
[0144] In addition, the innovative InnerATFLIoU loss function is introduced to replace the traditional CIoU loss. Combining ATFL's focus on difficult-to-classify samples with the stability advantage of Inner loss in bounding box regression, it improves the detection ability of small and easily confused targets, and achieves more accurate and robust target positioning and classification.
[0145] The specific expression of the InnerIoU loss function is:
[0146] L InnerIoU =1-IOU inner (15)
[0147] Among them, IOU inner Represents the inscribed IOU, which is an improved IoU indicator.
[0148]
[0149] inter represents the intersection area of the inscribed area between the predicted box and the true box, and union represents the area of the joint area formed by the predicted box and the true box;
[0150]
[0151]
[0152] Where, and (x c ,y c ) represent the center coordinates of the auxiliary bounding box and the auxiliary prediction box respectively; w gt ,h gt w and h represent the length and width of the real box and the predicted box respectively; ratio represents the adjustable scale factor, ranging from 0.5 to 1.5. and (b l ,b t ), (b l ,b b ), (b r ,b t ), (b r ,b b ) represent the four vertex coordinates of the auxiliary bounding box and the auxiliary prediction box respectively.
[0153] ATFL loss function:
[0154] Starting from the classic cross entropy loss function, its expression is:
[0155] L BCE=-(ylog(p)+(1-y)log(1-p)) (22)
[0156] Among them, p represents the predicted probability and y represents the true label.
[0157] Formula (22) can be simplified as:
[0158] L BCE =-log(p t ) (twenty three)
[0159]
[0160] The focal loss function (FL) is obtained by adjusting the factor (1-p t ) γ Solve the problem of sample imbalance, but while reducing the loss of easy samples, it will also reduce the loss of difficult samples, which is not conducive to difficult sample learning. The definition and calculation process of the focal loss function (FL) are shown in formula (25):
[0161] FL(p t )=-(1-p t ) γ logp t (25)
[0162] Where p t Represents the probability that the model predicts that the sample belongs to the true category; γ represents an adjustable parameter;
[0163] To solve the above problems, the threshold focus loss function (TFL) is proposed as shown in formula (26):
[0164]
[0165] Where γ is an adjustable parameter and η is a hyperparameter.
[0166] Adaptive Threshold Focal Loss Function (ATFL) adaptively adjusts γ and η.
[0167]
[0168] The final formula of the ATFL loss function is:
[0169]
[0170] The InnerATFLIoU loss function is defined as follows:
[0171] L InnerATFLIoU =L ATFL +L InnerIoU (29)
[0172] InnerATFLIoU significantly improves detection accuracy and robustness by adaptively focusing on key tasks and accurately optimizing the overlapping areas of target frames.
[0173] Step 6: Use the divided KITTI dataset to train the object detection model constructed in S2-S5.
[0174] Step 7: Use the trained object detection model to perform object detection on the test set and evaluate the performance of the model.
[0175] The original YOLOv10 network and the improved new YOLOv10 network were tested. The test process is as follows:
[0176] The original YOLOv10 network was trained iteratively on the KITTI dataset using the YOLOv10n.pt weights and YOLOv10 network. The runtime environment is CUDA 11.8, PyTorch 11.8, and the GPU is NVIDIA GeForce RTX 4050 16G.
[0177] During the model training process, after 300 iterations of training, the dataset detection pictures are as follows Figure 10 As shown, after 300 iterations, Figure 9 As shown in the figure, the mAP50 accuracy is 87.7%, the mAP50-95 accuracy is 66.5%, and the recall accuracy is 79.0%.
[0178] The present invention improves the YOLOv10 network as follows Figure 1 As shown, the KITTI dataset was iteratively trained using the YOLOv10n.pt weights and the improved YOLOv10 network. The operating environment is: CUDA 11.8; Pytorch 11.8; GPU NVIDIA GeForce RTX 4050 16G. The KITTI dataset was trained for 300 iterations. The dataset detection images are as follows Figure 12 As shown, the model training results are as follows Figure 11 As shown in the figure, the mAP50 accuracy is 91.3%, the mAP50-95 accuracy is 68.2%, and the recall accuracy is 85.4%.
[0179] Among them, the average precision mAP is: (91.3-87.7)×100%=3.6%;
[0180] mAP50-95: (68.2-66.5)×100%=1.7%;
[0181] recall: (85.4-79.0)×100%=6.4%.
[0182] Table 1
[0183]
[0184]
[0185] Experimental results are shown in Table 1. On the KITTI dataset, compared to the original YOLOv10 model, the new network model improves mAP50 accuracy by 3.6% from 87.7% to 91.3%. mAP50-95 accuracy increases by 1.7% from 66.5% to 68.2%. Recall accuracy also increases by 6.4% from 79.0% to 68.2%. These results are significant.
[0186] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made in accordance with the spirit of the present invention shall be covered by the present invention.
Claims
1. A road traffic vehicle target detection method based on deep learning, characterized in that: The following steps are involved: Step 1: Select the KITTI dataset and randomly divide it into training set, validation set, and test set according to the proportion; Step 2: Replace the backbone network of YOLOv10 with the SwinTransformer module, and introduce the SCEC-Swin module to replace the SwinStage module in the SwinTransformer module; Step 3: Introduce the PSELA attention mechanism after the SPPF layer of the YOLOv10 backbone network, and use the feature enhancement technology of one-dimensional convolution and group normalization to achieve lightweight implementation; Step 4: Improve the neck structure of YOLOv10: Introduce C2f-BG to replace the C2f module. Combining the advantages of AKConv and GAM in the C2f module, it better improves the algorithm's ability to process features while reducing the number of model parameters. Step 5: Introduce the detection head DyMLPHead to replace the original YOLOv10 detection head, and introduce the loss function InnerATFLIoU; Step 6: Use the divided data set to train the target detection model constructed in S2-S5; Step 7: Use the trained object detection model to perform object detection on the test set and evaluate the performance of the model.
2. A road traffic vehicle target detection method based on deep learning according to claim 1, characterized in that: In step 2, the SCEC-Swin module includes an SCEC branch and an STB branch, and the specific process is as follows: Step 2.1: The input image first passes through the SCEC branch, which consists of two branches. In the first branch, a global average pooling operation is performed on the image, and a 1D convolution is used to model the local channel relationship in the channel dimension: I1=Conv1D(AvgPool(I)) (1) I'1=I·σ(I1) (2) Among them, Conv1D represents a one-dimensional convolution operation, which is used to model the dependency between channels; AvgPool(I) represents the average pooling operation, σ represents the Sigmoid activation function, I'1 represents the output of the first branch, and I represents the input image; In the second branch, the feature map is pooled along the X and Y directions, and dilated convolution is used to indirectly enhance the feature representation ability of the facial area by increasing the receptive field. Given an input I, the maximum pooling operation is performed on each channel using the two spatial ranges (H, 1) or (1, W) of the pooling kernel along the horizontal and vertical coordinates, respectively. The horizontal coordinate pooling process is expressed as follows using formula (3): I h =MaxPool(I,1) (3) Among them, MaxPool(I,1) means performing the maximum pooling operation on all width values at the same height position; The vertical coordinate pooling process is expressed as follows using formula (4): I w =MaxPool(1,I) (4) Among them, MaxPool(1,I) means performing the maximum pooling operation on all height values at the same width position; Aggregate features along two spatial directions to generate a pair of direction-aware feature maps I h and I w , then I h and I w Cascade in a specific direction and obtain the intermediate feature map I by expanding the convolution layer f : I f =ReLu(BN(DC([I h ,I w ]))) (5) Among them, DC represents the expanded convolution operation, BN represents normalization, and ReLu represents the activation function; Then I f Split into two independent tensors along the spatial dimension and Two additional unfolded convolution operations transform the two tensors into tensors with the same number of channels: Where, DC h and DC w Represent the extended convolution operations in the horizontal and vertical directions respectively, and Sigmoid represents the activation function; Expand the output separately and As the attention weight, the final output is expressed as follows using formula (8): Step 2.2: The STB branch consists of four key components, including LN, W-MSA, MLP, and SW-MSA. LN calculates the normalized mean and variance of a specific dimension of a single sample. MLP establishes a nonlinear relationship between input features and output. W-MSA divides the feature map into windows and calculates self-attention within each window. SW-MSA simulates the correlation between regions by shifting windows, which can be expressed as follows using Formula (9): Where I represents the input image, i represents the current block, and I i , I i 'and I i+1 Represent the outputs of MLP, W-MSA and SW-MSA, STB output Indicates the output of the STB branch; Step 2.3: The total output is expressed as follows using formula (10): SCTB output =STB output +SCEC output (10)。 3. The method for detecting road traffic vehicles based on deep learning according to claim 1, characterized in that: In step 3, the PSELA attention mechanism process is as follows: The input feature map first passes through a convolutional layer for preliminary feature extraction. The extracted feature map is then fed into the ELA module, where the feature representation capability is enhanced through a local attention mechanism. Within the ELA, the input feature map is average pooled horizontally and vertically to obtain compressed feature maps with sizes of C×H×1 and C×1×W, respectively. Then, feature extraction is performed through 1D convolution, followed by normalization and Sigmoid activation to generate spatial attention weights in two directions. The two attention maps are element-wise multiplied with the original feature map to obtain a feature map enhanced by spatial attention. Secondly, the feature map enhanced by the ELA module will continue to pass through two convolutional layers to further extract feature information. In order to retain the information of the initial input, a residual connection is used to splice the original input with the current feature map to achieve feature fusion. Finally, the fused feature map is integrated and reduced in dimension through a convolutional layer to obtain the final output result.
4. The method for detecting road traffic vehicles based on deep learning according to claim 1, wherein: In step 4, the C2f-BG module replaces the Bottleneck layer in the C2f module with an AG-Bottleneck module and adds a GAM module after the first Conv in the C2f module. The AG-Bottleneck module replaces the Conv in the Bottleneck layer with an AKConv module with any sampling shape and any number of parameters, and adds a GAM module after the second AKConv, finally obtaining the AG-Bottleneck module.
5. The method for detecting road traffic vehicles based on deep learning according to claim 1, characterized in that: In step 5, the detection head DyMLPHead is specifically as follows: First, the input features are max-pooled and average-pooled and then merged. The merged features are fed into the Conv3×3 convolutional layer to extract local spatial context information. The convolutional features use the GELU activation function to enhance the nonlinear expression ability, and finally the HardSigmoid activation function is used to limit the output value between 0 and 1 to generate the weight of each feature; Secondly, the input features are sent to the Index to extract spatial position information, and the Offset and Sigmoid outputs are generated through 3×3 convolution; Finally, the features are globally averaged pooled to obtain the channel description vector, which is fed into the multi-layer perceptron MLP. The MLP further models the channel description, including two fully connected layers FC and a GELU activation layer, to obtain four parameters: α 1 , β 1 , α 2 , β 2 ; Normalize: Normalize the obtained parameters and add them to the vector [1,0,0,0] to enhance specific dimensions; Final output: The feature maps processed by multiple branches are integrated into the final output.
6. The method for detecting road traffic vehicles based on deep learning according to claim 1, characterized in that: The InnerATFLIoU loss function is as follows: InnerIoU loss function L InnerIoU The specific expression is: L InnerIoU =1-IOU inner (11) Among them, IOU inner It represents the inscribed IOU, which is an improved IoU indicator; inter represents the intersection area of the inscribed area between the predicted box and the true box, and union represents the joint area formed by the predicted box and the true box; Where, and (x c ,y c ) represent the center coordinates of the auxiliary bounding box and the auxiliary prediction box respectively; w gt ,h gt w and h represent the length and width of the real box and the predicted box respectively; ratio represents the adjustable scale factor, ranging from 0.5 to 1.
5. and (b l ,b t ), (b l ,b b ), (b r ,b t ), (b r ,b b ) represent the four vertex coordinates of the auxiliary bounding box and the auxiliary prediction box respectively; When obtaining the ATFL loss function, we use the classic cross entropy loss function L BCE Starting from, its expression is: L BCE =-(ylog(p)+(1-y)log(1-p)) (18) Where p represents the predicted probability and y represents the true label; Formula (18) is simplified to: L BCE =-log(p t ) (19) The focal loss function is obtained by adjusting the factor (1-p t ) γ To solve the problem of sample imbalance, its definition and calculation process are shown in formula (21): FL(p t )=-(1-p t ) γ logp t (21) Where p t Represents the probability that the model predicts that the sample belongs to the true category; γ represents an adjustable parameter; To solve the above problems, the threshold focus loss function TFL is proposed as shown in formula (22): Where γ is an adjustable parameter and η is a hyperparameter. The adaptive threshold focal loss function ATFL adaptively adjusts γ and η. The final formula of the ATFL loss function is: The InnerATFLIoU loss function is defined as follows: L InnerATFLIoU =L ATFL +L InnerIoU (25) The InnerATFLIoU loss function significantly improves detection accuracy and robustness by adaptively focusing on key tasks and accurately optimizing the overlapping area of the target box.
Citation Information
Cited By
Avionics product part target detection method and system based on improved YOLO
CN121639681A
Traffic target detection method and device based on YOLOv12n model, equipment and medium
CN122435571A
Traffic target detection method and device based on YOLOv12n model, equipment and medium
CN122435571B