Lightweight vehicle detection method based on improved YOLOv10
By improving the YOLOv10 model, introducing the PSASENetV2 module and the Dy_Sample operator, optimizing the feature expression and detection head structure, the detection accuracy and efficiency of the YOLO model in small object detection and complex scenarios is solved, and efficient identification of lightweight vehicle detection is achieved.
Patent Information
- Application Number
- CN202510366420.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-04
AI Technical Summary
The existing YOLO model has limitations in dealing with small object detection and category imbalance problems, and the detection accuracy and efficiency decrease in complex scenarios, resulting in high computing power demand and high identification costs.
By improving the YOLOv10 model, the PSASENetV2 module is introduced to combine channel and spatial attention mechanisms, replace the C2f_MLCA module and Dy_Sample dynamic upsampling operator, optimize the feature expression and detection head structure, and reduce the parameter quantity and calculation cost.
While maintaining high detection accuracy, the parameter quantity and calculation cost of the model are significantly reduced, the accuracy and efficiency of small object detection are improved, and it is suitable for vehicle identification under equipment restricted conditions.
Smart Images

Figure CN120259632A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent traffic detection, and is a lightweight vehicle detection method based on improved YOLOv10. Background Art
[0002] With the rapid development of intelligent transportation systems and autonomous driving technologies, vehicle detection and recognition technologies play a crucial role in multiple fields such as traffic monitoring, vehicle counting, violation detection, autonomous driving, and intelligent parking. The accuracy of vehicle detection and recognition directly affects the performance and reliability of these systems. In recent years, object detection algorithms based on deep learning, especially the YOLO (You Only Look Once) series models, have been widely used in the field of object detection due to their fast, accurate, and easy-to-implement characteristics.
[0003] However, the YOLO model still has certain limitations in dealing with small object detection and class imbalance problems. In addition, with the diversification and complexity of application scenarios, more parameters bring higher computing power requirements and higher recognition costs. When the original model faces complex scenario recognition, the accuracy and efficiency of vehicle detection will also decrease. Therefore, there is an urgent need for a method that can reduce the parameter operation cost and meet the actual detection accuracy requirements to meet the actual vehicle detection needs. Summary of the Invention
[0004] Object of the Invention: To solve the problems mentioned in the background art, the present invention discloses a lightweight vehicle detection method based on improved YOLOv10. By setting and replacing custom modules to improve the YOLOv10 model, while maintaining high detection accuracy, the number of parameters and the model calculation cost are reduced, which is suitable for vehicle recognition under device-limited conditions.
[0005] Technical Solution:
[0006] The present invention discloses a lightweight vehicle detection method based on improved YOLOv10, and the method includes the following steps:
[0007] Step 1: Collect vehicle image data, label the data and then divide it into a training set, a validation set, and a test set;
[0008] Step 2: Build a vehicle detection model based on the improved YOLOv10 network:
[0009] Step 2.1 The traditional SENetV2 combines GAP and GMP to capture channel information, which is input into the fully connected layer to generate channel weights; the spatial attention mechanism is fused to obtain the improved SENetV2 module, and the improved SENetV2 module is introduced to replace the Attention in the original PSA module of the baseline model to construct the PSASENetV2 module;
[0010] Step 2.2 The C2f_MLCA module is introduced to replace the C2f module in the Backbone and Head parts;
[0011] Step 2.3 The dynamic upsampling operator DySample is introduced in the YOLOv10 network structure to replace the original static upsampling structure;
[0012] Step 3 The vehicle detection model is trained and verified based on the training set, validation set, and test set to evaluate its performance in vehicle recognition.
[0013] Furthermore, the specific steps of the improved SENetV2 module combining GAP and GMP are as follows:
[0014] Input features
[0015] Assume the input feature map is X ∈ R H×W×C , where H is the height of the feature map, W is the width, C is the number of channels. To reduce the fluctuation of the eigenvalue range, the input feature map will be normalized and the basic features are generated through the convolutional network;
[0016] Channel attention generation
[0017] Global pooling operations are performed on each channel of the input feature X, and the global average pooling GAP and global maximum pooling GMP are calculated respectively to capture channel information, and the features Figure X norm are converted into two channel description vectors and where and are the average and maximum values of channel c respectively, c ∈ [1, C];
[0018] The two global description vectors are concatenated, and the concatenated channel description is input into a two-layer fully connected network, and the channel weight s is generated through the activation function:
[0019] s = σ(W2 · ReLU(W1 · z + b1) + b2)
[0020] where: and is the weight of the fully connected layer; r is the reduction ratio, usually taking r = 16r; σ(·) is the Sigmoid activation function, and the input features are weighted and adjusted for each channel using the channel weight S. Figure X for each channel.
[0021] Furthermore, the steps of the spatial attention mechanism fusion are as follows:
[0022] Generate spatial attention X'
[0023] Perform global average pooling and max pooling on the channel-weighted features in the channel dimension to obtain two spatial feature maps. Concatenate the two spatial feature maps in the channel dimension and input them into a 7×7 convolutional layer to generate spatial attention weights.
[0024] Finally, combine the channel and spatial attention modules to obtain the enhanced feature map:
[0025] X fused (i,j,c) = M spatial (i,j) · s c · X(i,j,c)
[0026] where M spatial (i,j) is the spatial attention weight at position (i, j), s c is the channel attention weight of channel c, and X(i,j,c) is the value of the original feature map at position (i, j) and channel c. By multiplying the spatial attention weight and the channel attention weight with the original feature map, the enhanced feature map is obtained.
[0027] Furthermore, the steps of constructing the PSASENetV2 module are as follows:
[0028] Input a feature map with a shape of (C1, H, W);
[0029] Feature splitting: Use the convolutional layer cv1 to convert the input feature map into a 2C-channel feature map, where C = C1 × e. Split the feature map into two branches a and b, and the number of channels of each branch is C;
[0030] Attention enhancement: Apply the improved SENetV2 module to branch b to enhance the feature representation of branch b through the channel and spatial attention mechanisms;
[0031] Feed-forward neural network and residual connection: Apply FFN to branch b after attention processing, including two 1×1 convolutional layers, and then add the original branch b to branch b after attention and FFN processing;
[0032] Merge and Output: Concatenate branch a and the processed branch b, and then use the convolutional layer cv2 to restore the number of channels to C1 to obtain the final output feature map.
[0033] Furthermore, the operating structure of the C2f_MLCA module is as follows:
[0034] Initialize the C2f_MLCA module:
[0035] Input parameter setting: The number of input channels is c1, the number of output channels is c2, the number of module repetitions is n, whether to use the shortcut connection shortcut, the number of groups is g, and the expansion coefficient is e;
[0036] Calculate the hidden channel number c: Calculate the hidden channel number according to the formula c1 = int(c2 × e);
[0037] Initialize the first convolutional layer cv1: The number of output channels of this convolutional layer is c1, the output channel is c2, the kernel size k1 = 1, and the stride s1 = 1. Its convolutional operation can be expressed as: for the output feature map x, the output y1 of the convolutional layer satisfies y1 = conv(x, w1, b1), where w1 is the convolutional kernel weight and b1 is the bias. The specific operation of the convolutional kernel is:
[0038]
[0039] where i, j, k are the indices of the feature map, and k1 is the kernel size. l and m are the index variables of the convolutional kernel, used to traverse each position of the convolutional kernel. l represents the index of the convolutional kernel in the height direction, and m represents the index of the convolutional kernel in the width direction.
[0040] Initialize the second convolutional layer cv2: The input channel is (2 + n)c, the output channel is c2, the kernel size k1 = 1, and its convolutional operation for the input feature map x' has the output y2 satisfying y2 = conv(x', w2, b2). The convolutional kernel operation formula is the same as above;
[0041] Create a module list m: It consists of n Bottleneck modules. The number of input channels and output channels of each Bottleneck module is c, the shortcut connection is set to shortcut, the number of groups is g, the kernel size is k = ((3, 3), (3, 3)), and the expansion coefficient is 1.0;
[0042] Perform the forward propagation operation:
[0043] Convolution and splitting operation: Convolve the input data x through the convolutional layer cv1 i to obtain the output y cv1 , and then use the chunk method to split y along the channel dimension cv1Split into two tensors and store them in the list y. The splitting operation can be expressed as: Assume y cv1 has 2c channels, then it is split into two tensors y1 and y2 with c channels each, that is, y cv1 = [y1, y2];
[0044] Loop processing. For the last tensor y in the list y last process it sequentially through each Bottleneck module in the module list m;
[0045] Concatenation and convolution output:
[0046] Concatenation operation: Concatenate all tensors in the list y along the channel dimension. Let the concatenated tensor be y concat , whose number of channels is (2 + n)c. The concatenation formula can be expressed as: If y = [y1, y2, …, y n+2 , then y concat = concat(y1, y2, …, y n+2 );
[0047] Final convolution: Apply a convolution operation to the concatenated result y concat through the cv2 convolution layer to obtain the final output y output , that is, y output = cv2(y concat ), to obtain the output result of the C2f_MLCA module.
[0048] Furthermore, step 2.3 is specifically as follows:
[0049] In the YOLOv10 network structure, a dynamic upsampling operator DySample is introduced to replace the original static upsampling structure. Using dynamic point sampling to improve the model's utilization efficiency of edge details, and reducing the computational burden through grouped upsampling. The Dy_Sample module is applied after the C2f_MLCA module, receiving the output feature map from the C2f_MLCA module as input. Through dynamic point sampling, the Dy_Sample module significantly improves the model's utilization efficiency of edge details. To avoid overlapping of sampling positions, Dy_Sample introduces a range factor to limit the range of offsets, and uses a grid sampling function to reorganize the sampling results. Finally, the Dy_Sample module outputs a feature map with higher resolution, enabling the network to capture more detailed information and optimizing multi-scale feature fusion.
[0050] Furthermore, the vehicle detection model uses an LSCD detection head, adopts a shared GroupNorm convolution layer to replace the two ordinary convolution layers used by the Head detection head, and uses a scale layer to perform scale scaling processing on the features.
[0051] Beneficial effects:
[0052] 1. Through the improved PSASENetV2 module, the present invention combines channel and spatial attention mechanisms to capture key vehicle features while reducing redundant calculations. The fusion strategy of GAP and GMP reduces parameter redundancy compared with the traditional SENet while ensuring the accuracy of channel attention, significantly reducing the hardware cost of model deployment.
[0053] 2. Through dynamic feature splitting and recombination, the C2f_MLCA module of the present invention reduces the number of parameters of the Backbone while maintaining the multi-scale feature fusion ability, improves the network's ability to capture useful features, and also significantly improves the model's accuracy while maintaining computational efficiency by combining channel and spatial attention mechanisms.
[0054] 3. The present invention introduces a dynamic upsampling operator DySample to replace the original static upsampling structure, uses dynamic point sampling to improve the utilization efficiency of the model for edge details, reduces the computational burden through grouped upsampling, optimizes multi-scale feature fusion, effectively improves the upsampling efficiency, and thus further improves the detection accuracy of small defects. Description of the drawings
[0055] Figure 1 is the specific flowchart of the present invention;
[0056] Figure 2 is the overall structure diagram of the model of the present invention;
[0057] Figure 3 is the schematic diagram of the improved SENetV2 module of the present invention;
[0058] Figure 4 is the schematic diagram of the PSASENetV2 module of the present invention;
[0059] Figure 5 is the schematic diagram of the C2f-MLCA and MLCA modules of the present invention;
[0060] Figure 6 is the schematic diagram of the structure diagram module of the dynamic upsampling operator Dysample of the present invention;
[0061] Figure 7 is the schematic diagram of the sampling point generator in the dynamic dynamic upsampling operator of the present invention;
[0062] Figure 8 is the network structure diagram of the LSCD detection head of the present invention;
[0063] Figure 9 is the curve graph of the training process of the improved network model of the present invention;
[0064] Figure 10This is the actual detection effect diagram of the embodiments of the present invention. Specific embodiments
[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention:
[0066] As Figure 1 Figure 2 shown, the present invention discloses a lightweight vehicle detection method based on improved YOLOv10, and the method steps are as follows:
[0067] Step 1: Collect vehicle image data, label the data and then divide it into a training set, a validation set, and a test set.
[0068] Step 2: Construct an optimized YOLOv10 network structure, optimize the channel attention mechanism and the spatial attention mechanism on the basis of the traditional SENetV2 to improve the SENetV2 model. Enable the model to simultaneously pay attention to the feature importance of different channels and different spatial positions, and enhance the feature representation ability. Introduce the improved SENetV2 module to optimize the PSA module in the baseline model, construct the PSASENetV2 module, combine channel and spatial attention, optimize feature expression, and improve recognition accuracy; utilize the multi-scale local channel attention mechanism of the C2f_MLCA module to enhance the perception ability of vehicle feature details; adopt the Dy_Sample module to achieve dynamic sampling, adaptively adjust the sampling strategy, reduce the calculation amount, and accelerate the inference process; use the optimized v10Detect_LSCD detection head combined with local sensitive hashing and channel attention to improve the target localization and class discrimination efficiency.
[0069] Step 2.1, as Figure 3 , Figure 4 shown, the PSASENetV2 module: Optimize the channel attention mechanism and the spatial attention mechanism on the basis of the traditional SENetV2 module to improve this module. The improved SENetV2 module optimizes the PSA module in the baseline model, replaces the Attention in the original PSA module with the improved SENetV2 module, and constructs the PSASENetV2 module. A lightweight attention module that dynamically improves the network's perception ability and expression ability of key features by optimizing the channel attention mechanism and the spatial attention mechanism. Compared with the classic SENetV2 module, the optimized module significantly reduces the computational complexity while improving the performance. Specifically, the improved SENetV2 module is optimized in the following ways on the basis of the original SENetV2 module:
[0070] Improvement of Channel Attention Mechanism: Introduce a combination of GMP and GAP to generate a richer channel description vector.
[0071] Add Spatial Attention Mechanism: Improve the original SENetV2 module which only focuses on channel attention, and enhance the expression ability of spatial attention.
[0072] The specific operations of the improved SENetV2 module are as follows:
[0073] Input Feature
[0074] Assume the input feature map is X ∈ R H×W×C , where H is the height of the feature map, W is the width, and C is the number of channels. To reduce the fluctuation of the eigenvalue range, the input feature map will be normalized and basic features will be generated through a convolutional network.
[0075]
[0076] where μ and σ are the mean and standard deviation of the feature map respectively, and ε is a small number to prevent division by zero.
[0077] Channel Attention Generation
[0078] Perform global pooling operations on each channel of the input feature X, and calculate global average pooling (GAP) and global max pooling (GMP) respectively to capture channel information:
[0079]
[0080] Through the above operations, the feature Figure X norm is converted into two channel description vectors and where and are the average and maximum values of channel c respectively, c ∈ [1, C].
[0081] Concatenate the two global description vectors to obtain a richer feature representation:
[0082] Z' = [Z avg ; Z max ∈ R 2c
[0083] Input the concatenated channel description into a two-layer fully connected network, and generate the channel weight s through an activation function:
[0084] s = σ(W2 · ReLU(W1 · z + b1) + b2)
[0085] where: and is the weight of the fully connected layer; r is the reduction ratio, usually taking r = 16r; σ(·) is the Sigmoid activation function.
[0086] Using the channel weight S for the input features Figure X of each channel is weighted and adjusted:
[0087] X′(i,j,c) = s c ·X(i,j,c)
[0088] where s c is the weight of channel c, and X' is the feature map after channel weighting. The weighted feature Figure X ' highlights the important channels more prominently.
[0089] Spatial attention generates X'
[0090] Performing global average pooling and max pooling on the channel-weighted features in the channel dimension to obtain two spatial feature maps:
[0091]
[0092] Concatenating the two spatial feature maps along the channel dimension and inputting them into a 7×7 convolutional layer to generate spatial attention weights:
[0093] M soatial = σ(f 7×7 ([M avg ; M max ))
[0094] Channel and spatial attention fusion
[0095] Finally, combining the channel and spatial attention modules to obtain the enhanced feature map:
[0096] X fused (i,j,c) = M spatial (i,j)·s c ·X(i,j,c)
[0097] where M spatial (i,j) is the spatial attention weight at position (i, j), s c is the channel attention weight of channel c, and X(i,j,c) is the value of the original feature map at position (i, j) and channel c.
[0098] The specific operation steps of the PSASENetV2 module are as follows:
[0099] Input feature map: with shape (C1, H, W)
[0100] Feature splitting: Use the convolutional layer cv1 (1×1 convolution) to convert the input feature map into a feature map with 2C channels, where C = C1×e. Split the feature map into two branches a and b, and the number of channels for each branch is C.
[0101] Attention enhancement: Apply the improved SENetV2 module to branch b to enhance the feature representation of branch b through channel and spatial attention mechanisms.
[0102] Feed-forward neural network and residual connection: Apply FFN to branch b after attention processing, including two 1×1 convolutional layers. Then add the original branch b to branch b after attention and FFN processing.
[0103] Merge and output: Concatenate branch a and the processed branch b, and then use the convolutional layer cv2 (1×1 convolution) to restore the number of channels to C1 to obtain the final output feature map.
[0104] The SENetV2 module can efficiently improve the model's attention ability to specific target regions and key features through the dual mechanisms of channel attention and spatial attention, significantly enhancing the network's detection performance. In addition, the SENetV2 module optimizes the computational complexity of the original SENet, making it applicable to a wider range of lightweight object detection tasks. Through this fusion mechanism, the SENetV2 module can dynamically improve the ability to capture key features.
[0105] Step 2.2, as Figure 5 shown, the C2f_MLCA module: In the improvement of the YOLOv10 network structure, the C2f_MLCA module is introduced to replace the original C2f module. It replaces the Backbone and Head parts of the original YOLOv10 model, where the C2f_MLCA module is used to enhance the feature extraction ability and improve the detection accuracy. Specifically, the C2f_MLCA module replaces the original C2f module in the Backbone part, responsible for more effective feature extraction and downsampling; in the Head part, it also replaces the original C2f module to optimize the feature fusion and object detection process. Through this replacement, the C2f_MLCA module not only improves the network's ability to capture useful features, but also significantly improves the model's accuracy while maintaining computational efficiency by combining channel and spatial attention mechanisms. The specific operations of the C2f_MLCA module are as follows:
[0106] Initialize the C2f_MLCA module:
[0107] Input parameter settings: The number of input channels is c1, the number of output channels is c2, the number of times the module is repeated is n (default value is 1), whether to use a shortcut connection shortcut (default value is False), the number of groups g (default value is 1), and the expansion factor e (default value is 0.5).
[0108] Calculate the hidden channel number c: Calculate the hidden channel number according to the formula c1 = int(c2 × e).
[0109] Initialize the first convolutional layer cv1: The number of output channels of this convolutional layer is c1, the output channel is c2, the kernel size k1 = 1, and the stride s1 = 1. Its convolutional operation can be expressed as: For the output feature map x, the output y1 of the convolutional layer satisfies y1 = conv(x, w1, b1), where w1 is the convolutional kernel weight and b1 is the bias (here the bias is False). The specific operation of the convolutional kernel is:
[0110]
[0111] (Here, i, j, k are the indices of the feature map). k1 is the kernel size. l and m are index variables of the convolutional kernel, used to traverse each position of the convolutional kernel. l represents the index of the convolutional kernel in the height direction (vertical direction), and m represents the index of the convolutional kernel in the width direction (horizontal direction).
[0112] Initialize the second convolutional layer cv2: The input channel is (2 + n)c, the output channel is c2, the kernel size k1 = 1. Similarly, for the input feature map x', its convolutional operation output y2 satisfies y2 = conv(x', w2, b2), (b2 bias is False), and the convolutional kernel operation formula is the same as above.
[0113] Create a module list m: It is composed of n Bottleneck modules. The number of input channels and output channels of each Bottleneck module is c, the shortcut connection is set to shortcut, the number of groups is g, the kernel size is k = ((3, 3), (3, 3)), and the expansion factor is 1.0.
[0114] Perform the forward propagation operation:
[0115] Convolution and splitting operations: Convolve the input data x through the convolutional layer cv1 to obtain the output y cv1 . Then use the chunk method to split y cv1 into two tensors along the channel dimension and store them in the list y. The splitting operation can be expressed as: Assume the number of channels of y cv1 is 2c, then it is split into two tensors y1 and y2 with the number of channels c, that is, y cv1 = [y1, y2] (in the channel dimension).
[0116] Loop processing: For the last tensor in the list y (assumed to be y last ), it is processed sequentially through each Bottleneck module in the module list m. For the i-th Bottleneck module m i , its input is y last , and the output is y mi . The specific process is as follows: Inside the bottleneck module, first pass through the first convolutional layer cv 1i . The input channel number of this convolutional layer is c, the output channel is c' = int(c × e), and the convolutional kernel size is k 0i = 3, and the stride is 1), to obtain the intermediate result y cv1i , y cv1i = cv 1i (y last ); then pass through the second convolutional layer cv 2i (the input channel is c', the output channel is c, the convolutional kernel size is k 1i = 3, the stride is 1, and the number of groups is g), to obtain y cv2i = cv 2i (y cv1i ); then pass through the MLCA module for processing, to obtain y mi = MLCA(y cv2i ); if shutcut is True and the input and output channel numbers are equal c1 = c2, then the final output is y mi = y last + y mi , otherwise it is y mi , and y mi is added to the list y.
[0117] Concatenation and convolutional output:
[0118] Concatenation operation: Concatenate all the tensors in the list y along the channel dimension. Let the concatenated tensor be y concat , and its channel number is (2 + n)c. The concatenation formula can be expressed as: If y = y1, y2,..., y n+2 , then y concat = concat(y1, y2,..., y n+2 ) (along the channel dimension).
[0119] Final convolution: Apply the cv2 convolutional layer to the concatenation result y concat for convolution operation to obtain the final output y output , that is, y output = cv2(y concat ), thus obtaining the output result of the C2f_MLCA module.
[0120] Step 2.3, Dy_Sample Module: In the YOLOv10 network structure, a dynamic upsampling operator DySample is introduced to replace the original static upsampling structure. Dynamic point sampling is used to improve the utilization efficiency of the model for edge details, and grouped upsampling is used to reduce the computational burden. Specifically, the Dy_Sample module is applied after the C2f_MLCA module and receives the output feature map from the C2f_MLCA module as input. Through dynamic point sampling, the Dy_Sample module significantly improves the utilization efficiency of the model for edge details. This module adopts a grouped upsampling strategy, reducing the computational burden while adaptively setting an offset for the input feature map to determine the sampling position. To avoid overlapping of sampling positions, Dy_Sample also introduces a range factor to limit the range of the offset and uses a grid sampling function to reorganize the sampling results. Finally, the Dy_Sample module outputs a feature map with a higher resolution, enabling the network to capture more detailed information, optimizing multi-scale feature fusion, effectively improving the upsampling efficiency, and further enhancing the detection accuracy of small defects.
[0121] The dynamic upsampler DySample is as Figure 6 shown. It resamples the input feature Figure X with a given size of C×H×W and calculates the new upsampled feature Figure X ' through bilinear interpolation, that is, with a size of C×sH×sW. Network sampling can be defined as:
[0122] X' = gri_ample(X, S)
[0123] where S is the given upsampling scale factor, and the size of the input feature Figure X is C×H×W. Among them, the input channels and output channels of X are C and 2S 2 respectively, and an offset O with a size of 2S 2 ×H×W is generated through a linear layer. This offset O is used to adjust the sampling position of the feature map. And through pixel rearrangement, the original feature map is adjusted to improve the resolution. To achieve this process, the sampling set S is the sum of the offset O and the original sampling network G, because the two jointly determine the upsampling method and accuracy of the feature map, and its basic implementation can be defined as:
[0124] O = linear(X), S = G + O
[0125] where X represents the input feature, X' represents the upsampled feature, O represents the generated offset, and G represents the original network.
[0126] The sampling point generator resamples the input feature through network sampling to generate the sampling set. Figure 7The part shown demonstrates the sampling point generator in DySample, where the sampling set is obtained by adding the generated offset to the initial network position. Figure 7 The upper box represents the static range factor, and the lower box represents the dynamic range factor. The generation of the range factor is used to modulate the offset and simplify the context operation. Finally, the generated upsampled feature Figure X ′ has a size of C×sH×sW.
[0127] Step 2.4, v10Detect_LSCD module: The LSCD detection head uses a shared GroupNorm convolutional layer instead of the two ordinary convolutional layers used by the Head detection head, and uses a scale layer to perform scale scaling on the features to solve the problem of inconsistent target scales detected by each detection head.
[0128] In Figure 8 In the LSCD detection head network structure shown, the GN_Conv 1×1 module represents the group normalization GroupNorm + convolution Conv operation, where 1×1 represents the convolution kernel size of 1×1. The two group normalization convolution modules GN_Conv 3×3 share weights. This structure receives the input features of the P3, P4, and P5 levels. First, it uses a shared convolution with a convolution kernel of 1×1 to process, increasing the information exchange in the channel dimension; secondly, it uses a shared convolution with two 3×3 convolution kernels for information aggregation, reducing redundant information and increasing the learning probability of adjacent feature information; finally, the information extracted by the shared convolution is input into the classification and regression heads, and feature scaling is performed through the Scale layer to enhance the retention ability of multi-scale features.
[0129] By using a shared GroupNorm convolutional layer instead of an ordinary convolutional layer, the complexity and the number of parameters of the model can be significantly reduced while ensuring the model performance, thereby reducing the computational overhead in the inference stage, accelerating the training process, and saving computational resources, making the model more general and adaptable.
[0130] Step 3: Adopt advanced training strategies, including the gradient descent algorithm and backpropagation techniques, to optimize the neural network parameters, overfit and accelerate convergence, while ensuring the generalization ability of the model to maximize the detection accuracy of vehicle recognition.
[0131] During the model training phase, an iterative optimization method is adopted, combining the gradient descent algorithm and backpropagation technology to carefully adjust the network parameters. During the optimization process, special emphasis is placed on introducing the improved SENetV2 module to optimize the PSA module in the baseline model, constructing the PSASENetV2 module, combining channel and spatial attention to optimize feature representation and improve recognition accuracy; using the multi-scale local channel attention mechanism of the C2f_MLCA module to enhance the perception ability of vehicle feature details; adopting the Dy_Sample module to achieve dynamic sampling, adaptively adjusting the sampling strategy, reducing the computational amount, and accelerating the inference process; at the same time, using the optimized v10Detect_LSCD detection head combined with locality-sensitive hashing and channel attention to improve the efficiency of target localization and class discrimination; through the above optimization measures, overfitting is avoided and convergence is accelerated, while ensuring the generalization ability of the model to maximize the detection accuracy of vehicle recognition.
[0132] Step 4: Use the validation set for performance verification. Evaluate the generalization ability of the model through the validation set, focusing on key performance indicators such as accuracy, recall, and mean average precision (mAP), and perform necessary model tuning. Conduct a strict performance evaluation of the trained model on the validation set, focusing on key indicators such as accuracy, recall, and mean average precision (mAP). Through these indicators, the recognition ability of the model for different vehicle types and scales can be comprehensively understood. During the evaluation process, if it is found that the model performs poorly on certain specific scenarios or targets, the model parameters will be fine-tuned to improve its generalization ability and recognition accuracy. The goal of this stage is to ensure that the model can stably provide high-quality detection results in practical applications. Finally, the network model parameters are set as follows: the number of training epochs is 350, the momentum is 0.937, the initial learning rate is 0.01, the minimum learning rate is 0.0001, the weight decay coefficient is 0.00005, the SGD optimizer is used, and the batch-size is 4.
[0133] Step 5: Test the trained and validated model on the test set to evaluate its detection performance for the vehicle recognition system after the improvement of YOLOv10. Specifically: Conduct a comprehensive test on the trained and validated model on the test set, carefully record and analyze the performance indicators, and focus on evaluating the effect of the model in detecting various vehicles, especially small vehicles. The test results will be used to analyze the advantages and disadvantages of the model and provide a basis for subsequent model improvement. If the test results show that there is still room for improvement in certain aspects of the model, the model will be further optimized according to the feedback to ensure that the model can meet the requirements of practical applications. The test results are as Figure 9 shown, and the actual detection effect diagram is as Figure 10As shown, the change curves of various indicators during 350 rounds of model training. After the model training is completed, it is tested with the validation set. The precision of vehicles in the validation set reaches 90.1%, the mAP reaches 90.2%, and the recall rate reaches 82.3%. Compared with the original YOLOv10n, the algorithm in this paper improves the precision of vehicles by 2.5%, the mAP by 1.6%; the recall rate is increased by 2.5%, the number of model parameters is reduced by 29%, and the parameter amount is reduced by 23.8%. Among them, the precision represents the accuracy of the model. The larger this indicator is, the better the model recognition effect; the mAP represents the quality of the model in all categories. The larger this indicator is, the better the model network performance. The comparison table of algorithm indicators is shown in Table 1, which proves the feasibility of the algorithm of the present invention.
[0134] Table 1
[0135]
[0136] The above description of the embodiments enables those skilled in the art to implement or use the present invention. Various modifications to the embodiments will be obvious to those skilled in the art. The general principles of the present invention can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention should not be limited to the embodiments shown herein, but should cover the widest scope that conforms to the principles and novel features disclosed in the present invention.
Claims
1. A lightweight vehicle detection method based on improved YOLOv10, characterized in that, The method includes the following steps: Step 1: Collect vehicle image data, annotate the data and then divide it into a training set, a validation set, and a test set; Step 2: Construct a vehicle detection model based on the improved YOLOv10 network: Step 2.1 Combine the traditional SENetV2 with GAP and GMP to capture channel information, and input it into the fully connected layer to generate channel weights; Fuse the spatial attention mechanism to obtain the improved SENetV2 module, introduce the improved SENetV2 module to replace the Attention in the original PSA module of the baseline model, and construct the PSASENetV2 module; Step 2.2 Introduce the C2f_MLCA module to replace the C2f module in the Backbone and Head parts; Step 2.3 Introduce the dynamic upsampling operator DySample in the YOLOv10 network structure to replace the original static upsampling structure; Step 3 Train and validate the vehicle detection model based on the training set, validation set, and test set, and evaluate its performance in vehicle recognition.
2. The lightweight vehicle detection method based on the improved YOLOv10 according to claim 1, wherein, The specific steps of combining the improved SENetV2 module with GAP and GMP are as follows: Input feature Assume the input feature map is \(X\in\mathbb{R}\) H×W×C , where \(H\) is the height of the feature map, \(W\) is the width, and \(C\) is the number of channels. To reduce the fluctuations in the range of feature values, the input feature map will be normalized and the basic features will be generated through a convolutional network; Channel attention generation Perform global pooling operations on each channel of the input feature X, and calculate the global average pooling GAP and the global maximum pooling GMP respectively to capture channel information, and the feature map X norm is converted into two channel description vectors and where and are the average value and the maximum value of channel c respectively, c ∈ [1, c]; Concatenate the two global description vectors, input the concatenated channel description into a two-layer fully connected network, and generate the channel weight s through the activation function: s = σ(W2·ReLU(W1·z′ + b1)+b2) Wherein: and are the weights of the fully connected layer; r is the reduction ratio, usually taken as r = 16r; σ(·) is the Sigmoid activation function, and use the channel weight S to weightedly adjust each channel of the input feature map X.
3. The lightweight vehicle detection method based on the improved YOLOv10 according to claim 2, characterized in that, The steps of fusing the spatial attention mechanism are as follows: Spatial attention generation X' Perform global average pooling and max pooling on the channel-weighted features in the channel dimension to obtain two spatial feature maps, concatenate the two spatial feature maps in the channel dimension, and input them into a 7×7 convolutional layer to generate spatial attention weights; Finally, combine the channel and spatial attention modules to obtain the enhanced feature map: X fused (i, j, c) = M spatial (i, j) · s c · X(i, j, c) Among them, M spatial (i, j) is the spatial attention weight at position (i, j), s c is the channel attention weight of channel c, and X(i, j, c) is the value of the original feature map at position (i, j) and channel c. By multiplying the spatial attention weight and the channel attention weight with the original feature map, an enhanced feature map is obtained.
4. The lightweight vehicle detection method based on the improved YOLOv10 according to claim 1 or 3, characterized in that, The construction steps of the PSASENetV2 module are as follows: Input a feature map with a shape of (C1, H, W); Feature splitting: Use the convolutional layer cv1 to convert the input feature map into a 2C-channel feature map, where C = C1×e, and split the feature map into two branches a and b, and the number of channels of each branch is C; Attention enhancement: Apply the improved SENetV2 module to branch b, and enhance the feature representation of branch b through the channel and spatial attention mechanisms; Feed-forward neural network and residual connection: Apply FFN to branch b after attention processing, including two 1×1 convolutional layers, and then add the original branch b to branch b after attention and FFN processing; Merge and output: Concatenate branch a and the processed branch b, and then use the convolutional layer cv2 to restore the number of channels to C1 to obtain the final output feature map.
5. The lightweight vehicle detection method based on the improved YOLOv10 according to claim 1, characterized in that The operating structure of the C2f_MLCA module is as follows: Initialize the C2f_MLCA module: Input parameter setting: the number of input channels is c1, the number of output channels is c2, the number of module repetitions is n, whether to use a shortcut connection shortcut, the number of groups g, and the expansion factor e; Calculate the hidden channel number c: Calculate the hidden channel number according to the formula c1 = int(c2 × e); Initialize the first convolutional layer cv1: The output channels of this convolutional layer are c1, the output channels are c2, the kernel size k1 = 1, and the stride s1 = 1. Its convolution operation can be expressed as: for the output feature map x, the output y1 of the convolutional layer satisfies y1 = conv(x, w1, b1), where w1 is the convolutional kernel weight and b1 is the bias. The specific operation of the convolutional kernel is: Among them, i, j, k are the indices of the feature map, k1 is the kernel size, and l and m are the index variables of the convolutional kernel, used to traverse each position of the convolutional kernel. l represents the index of the convolutional kernel in the height direction, and m represents the index of the convolutional kernel in the width direction; Initialize the second convolutional layer cv2: The input channels are (2 + n)c, the output channels are c2, the kernel size k1 = 1. Its convolution operation for the input feature map x' has an output y2 that satisfies y2 = conv(x', w2, b2). The convolutional kernel operation formula is the same as above; Create a module list m: It consists of n Bottleneck modules. The number of input channels and output channels of each Bottleneck module are both c. The shortcut connection is set to shortcut, the group is g, the kernel size is k = ((3, 3), (3, 3)), and the expansion factor is 1.0; Perform the forward propagation operation: Convolution and splitting operations. The input data x is convolved through the cv1 convolutional layer i to obtain the output y cv1 , and then the chunk method is used to split y cv1 into two tensors along the channel dimension and stored in the list y. The splitting operation can be expressed as: assuming that the number of channels of y cv1 is 2c, it is split into two tensors y1 and y2 with the number of channels c, that is, y cv1 = [y1, y2]; Loop processing, for the last tensor y in the list y last Process it sequentially through each Bottleneck module in the module list m; Concatenation and convolution output: Concatenation operation: Concatenate all tensors in list y along the channel dimension. Let the concatenated tensor be y concat , whose number of channels is (2 + n)c. The concatenation formula can be expressed as: If y = y1, y2, …, y n+2 , then y concat = concat(y1, y2, …, y n+2 ); Final Convolution: The concatenation result y concat is convolved through the cv2 convolutional layer to obtain the final output y output , that is, y output = cv2(y concat ), obtaining the output result of the C2f_MLCA module.
6. The lightweight vehicle detection method based on the improved YOLOv10 according to claim 1, wherein Step 2.3 is as follows: In the YOLOv10 network structure, a dynamic upsampling operator DySample is introduced to replace the original static upsampling structure. Dynamic point sampling is used to improve the utilization efficiency of the model for edge details, and the computational burden is reduced through grouped upsampling. The Dy_Sample module is applied after the C2f_MLCA module and receives the output feature map from the C2f_MLCA module as input. Through dynamic point sampling, the Dy_Sample module significantly improves the utilization efficiency of the model for edge details. To avoid overlapping of sampling positions, Dy_Sample introduces a range factor to limit the range of offsets and uses a grid sampling function to reorganize the sampling results. The Dy_Sample module outputs a feature map with a higher resolution, enabling the network to capture more detailed information and optimizing multi-scale feature fusion.
7. The lightweight vehicle detection method based on the improved YOLOv10 according to claim 1, characterized in that, The vehicle detection model uses an LSCD detection head, replaces the two ordinary convolutional layers used by the Head detection head with a shared GroupNorm convolutional layer, and uses a scale layer to perform scale scaling processing on the features.