Night vehicle detection method based on optimized YOLOv11 model
By constructing the night vehicle detection data set and optimizing the feature extraction and fusion module of the YOLOv11s model, a variety of innovative modules are used to improve the accuracy and efficiency of night vehicle detection, solving the small-scale vehicle and occlusion scene problems of the YOLOv11 model in night vehicle detection, and achieving efficient vehicle detection.
Patent Information
- Application Number
- CN202510438757.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-08
AI Technical Summary
The existing YOLOv11 model has problems such as low detection accuracy of small-scale vehicles, reduced performance of occlusion scenes, and contradiction between computing efficiency and accuracy in night vehicle detection, which is difficult to meet the real-time requirements of on-board equipment.
The night vehicle detection data set is constructed, the feature extraction module and feature fusion module of the YOLOv11s model are optimized, and the SPD Conv module, the Monte Carlo attention module with dynamic masking mechanism is adopted, the large core selection module, the parallel patch perception attention module, the hierarchical offset Dysample upsampling operator and the improved bidirectional feature pyramid network are optimized, and the detection head structure is used to reduce the parameter amount using the H-MBConv module.
Without significantly increasing the number of parameters, the detection performance of the vehicle detection network is improved, the detection accuracy of small-scale vehicles and shading vehicles is improved, and the real-time requirements are met.
Smart Images

Figure CN120279531A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and autonomous driving, and particularly relates to a method for detecting night vehicles based on an optimized YOLOv11 model, which is particularly applicable to scenarios of small-scale target detection, enhancement of occlusion scene features, and deployment of in-vehicle devices with high real-time requirements under low-light conditions. Background Art
[0002] Night vehicle detection is one of the core tasks of autonomous driving environmental perception. Traditional YOLO series models are commonly used in this task. They are representative algorithms for single-stage object detection, with good real-time performance and detection accuracy. And YOLOv11 is the latest model in the YOLO series and has the best effect. Therefore, YOLOv11s is selected as the benchmark model for night vehicle detection. However, due to problems such as low contrast and blurred target details in night images, there are prone to missed detections of small-scale vehicles, misdetections of occluded targets, and incorrect class recognition. The specific manifestations are as follows:
[0003] (1) Low detection accuracy for small-scale vehicles: The pixel area of vehicles at a long distance is small, and the blurred night images result in the loss of texture features, making it difficult for deep networks to effectively extract detailed information;
[0004] (2) Degraded performance in occlusion scenarios: The features of partially occluded targets are broken, and it is difficult for the model to recover the complete target information;
[0005] (3) Contradiction between computational efficiency and accuracy: Although YOLOv11s has a small number of parameters, the fixed structure of the original model in the feature extraction and fusion modules cannot adapt to complex night scenarios.
[0006] Existing improvement methods mostly improve accuracy by increasing model complexity, but it is difficult to meet the real-time requirements of in-vehicle devices. Therefore, there is an urgent need for an optimized night vehicle detection solution that takes into account both accuracy and efficiency. Summary of the Invention
[0007] In view of the above problems, the present invention proposes a method for detecting night vehicles based on an optimized YOLOv11s, which improves the detection performance through collaborative improvement of multiple modules. The specific technical solutions are as follows:
[0008] (1) Construct a night vehicle detection dataset, including night scene images screened from the BDD100K dataset and a self-built domestic night driving dataset, and perform resolution unification and annotation processing on all images;
[0009] Select 12,500 night scene images from the BDD100K dataset. The BDD100K dataset is a dataset dedicated to object detection, covering dozens of categories and can be obtained through public network channels. Eliminate blurred and vehicle-free frames. Build a domestic night driving dataset of 3,500 images and generate YOLO format labels through the LabelImg annotation tool. LabelImg is a commonly used software for object detection standards. Divide the training set, validation set, and test set in a ratio of 8:1:1 to ensure balanced class distribution.
[0010] (2) Optimize the feature extraction module of the YOLOv11s model. Replace the conventional stride convolution of YOLOv11s with the SPD Conv module and introduce the Monte Carlo attention module based on the dynamic mask mechanism (DM-MCAttn) to enhance the feature selection ability.
[0011] SPD Conv module: Replace the conventional stride convolution. Achieve lossless downsampling through space-to-depth transformation, retain local features of small targets. SPD Conv is a module that replaces direct convolution with channel recombination during downsampling. Monte Carlo attention module based on the dynamic mask mechanism: Introduce dynamic mask weights in the Monte Carlo attention module. The Monte Carlo attention module is a method that uses the Monte Carlo method to randomly select a scale of feature map to optimize feature selection.
[0012] (3) Optimize the feature fusion module of the YOLOv11s model, including introducing the large kernel selection module (LSK Block), parallel patch-aware attention module (PPA Block), hierarchical offset Dysample upsampling operator, and improved bidirectional feature pyramid network (D-BiFPN).
[0013] Large kernel selection module: A module that uses 5×5 convolution and 7×7 dilated convolution to extract multi-scale context and dynamically fuse through channel attention. In this invention, the large kernel selection module is inserted into the feature fusion module of YOLOv11s. Parallel patch-aware attention module: A method that divides local patches and calculates global template similarity to enhance the perception of occluded target details. In this invention, the parallel patch-aware attention module is inserted into the feature fusion module of YOLOv11s. Hierarchical offset Dysample upsampling operator: The Dysample upsampling operator is an operator that optimizes uniform sampling through learnable parameters. Hierarchical offset is proposed in this invention for high vehicle speeds in vehicle detection, and a two-stage offset grid is proposed to achieve high-fidelity upsampling. Introduce the bidirectional feature pyramid network, which is a method that optimizes the feature fusion path. Based on the bidirectional feature pyramid network, this invention proposes an improved bidirectional feature pyramid network that adds shallow feature information.
[0014] (4) Optimize the detection head of the YOLOv11s model. Replace the conventional convolution in the original regression branch with the H-MBConv module, and combine the ECA channel attention mechanism to reduce the number of parameters.
[0015] The H-MBConv module combines depthwise separable convolution and standard 3×3 convolution to balance the receptive field and computational efficiency of YOLOv11s. Among them, H-MBConv is an improvement on MBConv. MBConv is a module that uses lightweight convolution to balance real-time performance and accuracy. H-MBConv replaces the SE module with the ECA module in it. The ECA module is a one-dimensional convolution to compress the computational amount of channel attention, and the SE module is a module that calculates channel attention through convolution.
[0016] (5) Use the training set to train the optimized model, and evaluate the detection performance through the validation set and the test set to complete the vehicle detection task in the night scene.
[0017] Further, the dataset described in step (1) consists of the following: 12,500 night scene images screened from the BDD100K dataset, excluding blurred and vehicle-free frames; the self-built domestic night driving dataset contains 3,500 domestic night driving images, and the vehicle category and bounding box labels in YOLO format are generated through the LabelImg annotation tool; the training set, validation set, and test set are divided in a ratio of 8:1:1 to ensure balanced class distribution.
[0018] Further, the SPD Conv module in step (2) realizes lossless downsampling through space-to-depth transformation and retains the local features of small-scale targets; the Monte Carlo attention module based on the dynamic mask mechanism adjusts the Monte Carlo attention module through dynamic mask weights, as shown in formulas (1) and (2):
[0019] A mask =σ(W4·δ(W3·x)) (1)
[0020] A final =A×A mask (2)
[0021] In the formula, W3 and W4 are 1×1 weight matrices, δ(·) represents the ReLU activation function to increase the non-linear expression ability, σ(·) represents the Sigmoid activation function to normalize the weight result to [0,1], A is the attention weight obtained by the Monte Carlo attention module, and the two are combined to obtain the final weighted weight A final , which is the final Monte Carlo attention module based on the dynamic mask mechanism, and then input into the SE layer and subsequent weighted on the input x.
[0022] Further, in step (3), the large kernel selection module uses 5×5 convolution and 7×7 dilated convolution to extract multi-scale context information and dynamically fuse it through channel attention. The parallel patch-aware attention module divides the input features into local patches, calculates the similarity weights in combination with the global template, and enhances the local detail perception of occluded targets. The improved bidirectional feature pyramid network fuses shallow detail and high-level semantic features and optimizes the multi-scale feature competition problem by using weighted bidirectional connections. The hierarchical offset Dysample upsampling operator realizes high-fidelity upsampling through a two-stage offset grid, as shown in formulas (3) and (4):
[0023]
[0024] where O (1) represents the original offset obtained by inputting the feature map into a 1*1 convolution, and O' (1) is the final offset obtained in the above one-time offset process.
[0025] Further, in step (4): The H-MBConv module mixes depthwise separable convolution and standard 3×3 convolution in the regression branch to balance the receptive field and computational efficiency, and uses the ECA module to replace the SE module to compress the computational amount of channel attention through one-dimensional convolution.
[0026] Further, the verification of the autonomous driving target detection network model in step (4) specifically means: setting the number of batch-processed pictures in step (3) to 4, selecting the number of training iterations to be 200 rounds, setting the initial learning rate to 0.01, using the stochastic gradient descent SGD (Stochastic Gradient Descent) as the optimizer, setting the momentum to 0.937, setting the weight update decay to 0.0005, and setting the IOU intersection over union to 0.6 for verification.
[0027] The present invention relates to a method for night vehicle detection based on an optimized YOLOv11 model, which includes the following steps: (1) constructing a night vehicle detection dataset; (2) optimizing the feature extraction module of the YOLOv11s model; (3) optimizing the feature fusion module of the YOLOv11s model; (4) optimizing the detection head of the YOLOv11s model; (5) training the optimized model using the training set, and evaluating the detection performance through the validation set and the test set to complete the vehicle detection task in the night scene. In the feature extraction module, SPDConv is introduced to replace the stride convolution, and the dynamic mask Monte Carlo attention mechanism (DM-MCAttn) is adopted. In the feature fusion module, the large kernel selection module (LSKBlock) and the parallel patch attention module (PPABlock) are integrated. At the same time, hierarchical offset Dysample and BiFPN are introduced, and the detection head structure is optimized. The lightweight H-MBConv module is used to reduce the number of parameters. Finally, an optimized YOLOv11 night vehicle detection network model is constructed for training, and it is verified using the images in the validation dataset, and finally the purpose of vehicle detection is achieved. The present invention can make the vehicle detection network achieve better performance without significantly increasing the number of parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 FIG. is the detection flow chart of a method for night vehicle detection based on YOLOv11s involved in the present invention.
[0029] Figure 2 FIG. is the schematic diagram of the principle structure of a method for night vehicle detection based on YOLOv11s involved in the present invention.
[0030] Figure 3 FIG. is the snapshot of the dataset used.
[0031] Figure 4 FIG. is the schematic diagram of the optimized feature extraction module.
[0032] Figure 5 FIG. is the schematic diagram of the optimized feature fusion module.
[0033] Figure 6 FIG. is the schematic diagram of the result of the optimized detection head.
[0034] Figure 7 FIG. is the application effect diagram in the embodiment of a method for night vehicle detection based on YOLOv11s involved in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0035] Embodiment: A method for night vehicle detection based on YOLOv11s, as Figure 1 、 Figure 2 shown, is characterized by including the following steps:
[0036] (1) Select 12,500 night scene images from the publicly available dataset BDD100K, and remove invalid frames that are blurred, have no vehicles, or have abnormal lighting. Build a domestic night driving dataset, collect 3,500 images, covering scenarios such as urban roads, highways, and tunnels. Generate vehicle bounding box labels in YOLO format (categories include sedans, trucks, buses, etc.) through the LabelImg annotation tool. The dataset snapshot is as Figure 3 shown.
[0037] (2) Data preprocessing: Resolution unification: Resize all images to a resolution of 640×640, and use bilinear interpolation to maintain the image ratio. Data augmentation: Apply Mosaic augmentation (four-image stitching), random flipping (horizontal probability 0.5), and HSV color space perturbation (hue ±0.1, saturation ±0.7, value ±0.4) to improve the generalization ability of the model. Anchor box clustering: Perform K-means clustering on the annotated boxes to generate 9 groups of initial anchor box sizes, adapting to the distribution characteristics of small targets at night.
[0038] (3) Dataset division: Divide the dataset into a training set (12,800 images), a validation set (1,600 images), and a test set (1,600 images) in a ratio of 8:1:1 to ensure balanced class distribution in each subset.
[0039] (4) Improve the YOLOv11 algorithm using the SPD Conv module, the Monte Carlo attention module based on the dynamic mask mechanism, the large kernel selection module, the parallel patch-aware attention module, the hierarchical offset Dysample upsampling operator, the improved bidirectional feature pyramid network, and the H-MBConv module to build a night vehicle detection network;
[0040] This network consists of an input module, a feature extraction module, a feature fusion module, and a detection head. As Figure 2 shown, among them, the input end of the input module collects the picture signal of the autonomous driving scene, performs data augmentation operations such as random cropping, random scaling, and random flipping on it, and then outputs it to the feature extraction module. The feature extraction module extracts features from it and passes the extracted feature information into the feature fusion module for feature pyramid pooling and feature fusion processing. At this time, the feature map after feature fusion is passed into the detection head for vehicle target detection, and finally the detection result picture signal is output.
[0041] Among them, the specific construction method of the night vehicle detection network based on optimizing YOLOv11s includes the following stages:
[0042] The first stage: Use the input module to collect the picture signal of the autonomous driving scene at night. After performing random cropping, random scaling, and random flipping, output the data after the data augmentation operation to the feature extraction module;
[0043] The second stage: Execute the SPDConv module and the Monte Carlo attention module based on the dynamic mask mechanism in the feature extraction module of the night vehicle detection network based on the optimized YOLOv11s. This process optimizes the feature extraction ability of the feature extraction module;
[0044] As Figure 4 shown, the SPD Conv module is responsible for optimizing feature extraction, and the Monte Carlo attention module based on the dynamic mask mechanism is responsible for optimizing feature selection. The combination of the two further improves the detection accuracy of small targets and occluded targets at night. The calculation formulas of the Monte Carlo attention module based on the dynamic mask mechanism are as shown in formulas (1) and (2):
[0045] A mask = σ(W4·δ(W3·x)) (1)
[0046] A final = A × A mask (2)
[0047] where W3 and W4 are 1×1 weight matrices, δ(·) represents the ReLU activation function, which increases the non-linear expression ability, and σ(·) represents the Sigmoid activation function, which normalizes the weight result to [0,1]. A is the attention weight obtained by the Monte Carlo attention module, and the combination of the two gives the final weighted weight A final , and then it is input into the SE module and subsequent weighted on the input x.
[0048] The third stage: Process the feature map extracted in the second stage in the feature fusion module. Among them, the large kernel selection module uses 5×5 convolution and 7×7 dilated convolution to extract multi-scale features in parallel, and dynamically fuses through channel attention. The parallel patch-aware attention module: divides the input features into 2×2 and 4×4 local patches, calculates the similarity weight in combination with the global template, and enhances the local details of the occluded target. The hierarchical offset Dysample upsampling operator: uses a two-stage offset grid to achieve high-fidelity upsampling, as shown in formulas (3) and (4):
[0049]
[0050] where O (1) represents the original offset obtained by inputting the feature map into a 1*1 convolution, and O' (1) is the final offset obtained in the above one-time offset process. The second-layer offset is for refinement and local fine-tuning. After the first upsampling uses the offset of the first layer, the first-layer upsampled feature map Y (1), and then input it into the second layer for offset calculation. Calculate the offset and scale of the second layer based on the upsampled feature map of the first layer, which can better distinguish similar categories or perform fine alignment on the local features of the vehicle. The final feature fusion module is optimized as shown in Figure 5 shown below.
[0051] Phase 4: Process the feature map extracted in the second phase in the feature fusion module, and decouple the classification and regression tasks separately in the detection head module, as shown in Figure 6 shown below. Specifically, it means: Use the H-MBConv module to mix depthwise separable convolution and standard 3×3 convolution to balance the receptive field and computational efficiency, replace the SE module with the ECA module, and compress the channel attention parameters through one-dimensional convolution.
[0052] Phase 5: Calculate the bounding box regression loss for the classification information predicted by the classification branch in the third phase, the detection box information and confidence information predicted by the regression branch:
[0053] First, calculate the classification loss. The calculation formula is shown in Formulas (5) and (6):
[0054] L cls_element = -[y·logσ(x) + (1 - y)·log(1 - σ(x))] (5)
[0055]
[0056] where x represents the original class score of the model without passing through the Sigmoid function, y represents the class true value, σ represents the Sigmoid function, N represents the number of valid anchor points, and C is the number of classes.
[0057] Calculate the regression loss according to Formulas (7), (8), and (9):
[0058]
[0059] where (b, b gt ) represents the center point coordinates of the candidate box and the ground truth box, ρ 2 represents using the Euclidean distance to measure the distance between the center point coordinates of the candidate box and the ground truth box, (w gt , h gt ) represents the width and height of the candidate detection box, and (w, h) represents the width and height of the ground truth annotation box. Then calculate the DFL Loss according to Formula (10):
[0060] DFL(S i , S i+1 ) = -((y i+1 - y)·log(S i ) + (y - y i )·log(S i+1)) (10)
[0061] (5) Use the images in the training dataset obtained in step (3) to train the night vehicle detection network model based on the optimized YOLOv11s: Adjust the resolution of all images in the training dataset to a fixed resolution of 640×640. Set the initial learning rate to 0.01, and as the number of iterations increases, the learning rate decreases. To improve the training speed, set the training batch size to 4. To prevent overfitting, set the number of training epochs to 200 for training, and finally obtain the trained night vehicle detection network model based on the optimized YOLOv11s;
[0062] (6) Use the images in the validation dataset obtained in step (3) to validate the trained night vehicle detection network model based on the optimized YOLOv11s obtained in step (5): Load the trained model in step (5) into the YOLOv11s network for validation. To improve the validation speed, set the validation batch size to 32 and the IOU (Intersection over Union) to 0.6 for validation;
[0063] (7) Use the images in the autonomous driving scenarios collected during the night autonomous driving process as input, and perform object detection in the trained night vehicle detection network model based on the optimized YOLOv11s obtained in step (5), so as to accurately identify the types of vehicles in the picture and complete the vehicle object detection task.
[0064] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0065] As Figure 2 shown in the embodiments, the present invention provides a night vehicle detection method and application based on the optimized YOLOv11s, and its operation process is as follows:
[0066] Step 1: Data input and preprocessing
[0067] Adopt the 2D object detection dataset in the BDD100K autonomous driving dataset and the self - collected dataset to train the model. This dataset contains a total of 3 categories detected in autonomous driving scenarios, namely "car", "truck", and "lorry". It consists of 16,000 real - image data collected in scenarios such as urban areas, rural areas, and highways. Each image contains up to 15 vehicles and 30 pedestrians at most.
[0068] Divide the 2D vehicle detection dataset in the constructed autonomous driving dataset into training set, validation set, and test set according to the ratio of 8:1:1, and adjust the resolution of all dataset images to a fixed resolution of 640×640.
[0069] Step 2: Model construction
[0070] The network structure of this model is as follows Figure 2 shown, and it consists of four parts: an input end, a feature extraction module, a feature fusion module, and a detection head. The input module inputs images, and the feature extraction module is used to extract the features of the images. The network mainly extracts the features of the images by the SPD Conv module and the Monte Carlo attention module based on the dynamic mask mechanism. The optimized structure of the feature extraction module is as follows Figure 3 shown. The feature fusion module connects the feature extraction module and the detection head. The input is multiple feature maps output by the feature extraction module, which is used to perform upsampling processing on the feature maps and fuse the features, and outputs three enhanced feature maps of different sizes. The optimized feature fusion module is as follows Figure 4 shown. As Figure 5 shown, the optimized detection head is used for object detection, decouples the classification task and the regression task separately, introduces the accurate bounding box regression loss and the classification loss, and realizes the classification and regression of the target.
[0071] Step 3: Train the model
[0072] The transfer learning method is adopted to train the model. The original YOLOv11s pre-trained model is loaded into the night vehicle detection network based on the optimized YOLOv11s for training. To prevent overfitting, the number of training iterations is set to 150 times; to make the objective function converge to the local minimum within a suitable time, the initial learning rate is set to 0.01, and it becomes smaller as the number of iterations increases; the optimization strategy selects the SGD (Stochastic Gradient Descent) optimization algorithm. Save the training weights and load the validation samples for subsequent verification of the model.
[0073] Step 4: Verify and apply the model
[0074] The evaluation criteria of the experiment are evaluated by the average precision (AP: Average Precision) and the mean average precision (mAP: Mean Average Precision). The AP value is calculated from the area formed by the PR curve composed of the precision (Precision) and the recall (Recall) and the horizontal and vertical coordinates. The calculation methods of Precision and Recall are as follows:
[0075]
[0076] Among them, TP is the positive class judged as the positive class, FP is the negative class judged as the positive class, FN is the positive class judged as the negative class, and TN is the negative class judged as the negative class. The mAP value represents the average of all class APs, and its calculation method is as follows:
[0077]
[0078] Experimental environment: A Python compilation environment with PyTorch 1.8.0, torchvision = 0.9.0, and CUDA 11.1 as the deep learning framework is built. The programming language and software used are Python and PyCharm respectively.
[0079] Experimental equipment: Windows 10 system, and the graphics card is NVIDIA GeForce GTX 1660Ti. Ablation experiments are used to test the influence of the optimized feature extraction module, optimized feature fusion module, and optimized detection head on the detection results, and comparative experiments are carried out with multiple networks. The experimental results are shown in Table 1, Table 2, and Table 3.
[0080] "√" indicates that the optimization measure is used, and "×" indicates that the optimization measure is not used. The experimental results are shown in the table. Among them, using the SPD Conv module in the feature extraction module is denoted as A, and using the Monte Carlo attention module based on the dynamic mask mechanism is denoted as B; using the large kernel selection module in the feature fusion module is denoted as C, the parallel patch-aware attention module is denoted as D, the hierarchical offset Dysample upsampling operator is denoted as E, the improved bidirectional feature pyramid network is denoted as F, and H-MBConv is denoted as G.
[0081] Table 1 Ablation study 1 of each optimization module on the self-built dataset
[0082]
[0083] The feature extraction module is denoted as H, the feature fusion module is denoted as I, and the detection head is denoted as J. The experimental results are shown in Table 2.
[0084] Table 2 Ablation study 1 of each optimization module on the self-built dataset
[0085]
[0086]
[0087] Table 3 Comparison of the detection effects of the present invention and other networks
[0088]
[0089] As can be seen from Table 1, the mAP of the original YOLOv11s model on the night vehicle detection dataset constructed in 3.3.1 of this paper 0.5 is 81.1, and the number of parameters is 9.43M. After introducing the SPD Conv module in the feature extraction module, the number of parameters is reduced by 1.06M, and the mAP 0.5It has increased by 0.4%. The SPD Conv module belongs to the sparse computing method and only performs calculations in necessary areas, significantly improving the operation efficiency. Through information-preserving downsampling and efficient channel compression, while reducing the number of parameters, it enhances the feature expression ability. To further improve the model's feature extraction ability for key regions, on the basis of introducing the SPD Conv module, a Monte Carlo attention module based on the dynamic mask mechanism is inserted into the feature extraction module. The number of model parameters has increased by 0.03M, but the mAP 0.5 has increased by 1.2%. It brings a large gain at the expense of a small increase in parameters. The reason is the adaptive selection ability brought by this module, which can effectively improve the model's feature extraction ability. The combination of the SPD Conv module and the Monte Carlo attention module based on the dynamic mask mechanism can capture the fine feature information of small-scale vehicles and give higher attention weights to occluded vehicles, enabling rich feature information to be extracted for both small-scale vehicles and occluded vehicles in the night scene. It can effectively improve the feature extraction ability of the feature extraction module and the performance of the night vehicle detection model, and further improve the detection average precision.
[0090] At the same time, to further solve the problems of low detection accuracy of small target vehicles and the decline in vehicle detection performance under occlusion, this paper further optimizes the feature fusion module. After adding the large kernel selection module to the feature fusion module, the overall number of model parameters has decreased by 0.07M, and the mAP 0.5 has increased by 0.02%. To enable the night vehicle detection model to simultaneously focus on context information and local feature information, a parallel patch-aware attention module is introduced into the feature extraction module. The number of model parameters has increased by 6.77M, and the mAP 0.5 has increased by 0.5%. The reason is that this module can capture local information at different scales through multiple parallel convolutional branches and finely focus on feature information at different scales. To solve the problem of uniform upsampling caused by the original upsampling operator in the feature network, resulting in the distortion of small target details in the upsampling stage, the Dysample upsampling operator is introduced. And to further reduce the instability during the training process, the hierarchical offset Dysample upsampling operator is proposed. Since the number of parameters of this operator is small, the increase in the number of final model parameters is negligible, and the mAP 0.5It has increased by 0.5%. The reason is that the hierarchical offset Dysample upsampling operator can ensure the high-fidelity reconstruction of features, enabling the details of small targets to be reconstructed during the sampling process. To further address the difficulties in nighttime vehicle detection mentioned in 3.1.3, from the perspective of feature fusion, since certain optimization work has been done on feature extraction in this paper, obtaining fine shallow detail information and high-level semantic information, and these information may be weakened due to the competition and weight allocation between multi-scale features during fusion. Therefore, referring to the idea of the bidirectional feature pyramid network, the feature fusion path of the feature pyramid is optimized, and an improved bidirectional feature pyramid network is proposed. Finally, the increase in the number of parameters is negligible, and the mAP 0.5 It has increased by 1.6%. This is because small-scale vehicles often need to fuse shallow detail information in a larger proportion, while occluded vehicles require more high-level semantic information for assistance. These two types of features have been obtained through optimization measures in this paper. After introducing the improved bidirectional feature pyramid network in the present invention, the fusion of shallow features and the weighting of features with different resolutions are newly added in the feature fusion stage, improving the efficiency and effect of feature fusion. Without the need to add additional modules, by enriching the fusion path of feature information, the detection accuracy of the model can be effectively improved.
[0091] Finally, since the previous optimization measures have increased a certain number of parameters, therefore, on the premise of ensuring the average detection accuracy, by optimizing the parameters of the detection head and introducing H-MBConv into the regression branch of YOLOv11s, the overall number of parameters of the final model has decreased by 0.5M, and the mAP 0.5 It has increased by 1.6%. The reason is that while the improved detection head reduces the number of parameters through depthwise separable convolution, it can maintain a strong feature expression ability through the ECA module.
[0092] As can be seen from Table 2, the optimization measures in this paper for the feature extraction module, feature fusion module, and detection head all play a positive promoting role in the performance of the nighttime vehicle detection model. Among them, the optimized feature extraction module combines the feature extraction ability of the SPD Conv module, improves the problem of small-scale vehicle detection through information-lossless downsampling, implicitly retains multi-scale local features, reconstructs the global representation of occluded vehicles through channel complementarity, and the feature selection ability brought by the Monte Carlo attention module based on the dynamic mask mechanism, making the overall performance of the algorithm greatly improved, with the precision rate, recall rate, and mAP 0.5They have increased by 4.7%, 1.4%, and 1.6% respectively. Regarding the improvement of the feature fusion module, it combines the complementary relationship between the large kernel selection module and the parallel patch-aware attention module to achieve wide-area global information attention, local detail attention, refined feature screening, high-fidelity reconstruction details of the hierarchical offset Dysample upsampling operator, and multi-path fusion of the improved bidirectional feature pyramid network. The optimized feature fusion module can effectively enhance information flow, precision, recall, and mAP 0.5 They have increased by 6%, 2.5%, and 3.2% respectively. Regarding the optimization of the detection head, it is because the H-MBConv module further reduces redundant parameters, can maintain a certain detection accuracy while reducing the number of parameters, precision, recall, and mAP 0.5 They have increased by 2.1%, 0.1%, and 0.4% respectively. When any two of the feature extraction module, the feature fusion module, and the detection head are combined and optimized, they can all bring a positive promotion effect to the model. When all three are fully combined, the comprehensive detection performance reaches the optimal, precision, recall, and mAP 0.5 They have increased by 8%, 4.6%, and 4.6% respectively.
[0093] As can be seen from Table 3, the night vehicle detection algorithm designed in this paper has achieved the optimal performance in comprehensive detection, and the mAP 0.5 evaluation index has increased by 9.2%, 10.2%, and 10% respectively compared with single-stage detection algorithms such as SSD, YOLOv5s, and YOLOv7-tiny; compared with the relatively similar structured YOLOv8s and YOLOv11s, it has increased by 9.6% and 4.6% respectively; compared with the two-stage detection algorithm Faster RCNN, it has increased by 11.5%, and the number of model parameters is much less than that of Faster RCNN. The current experimental results show that the night vehicle detection algorithm designed in this paper has superiority in average detection accuracy.
[0094] Finally, the images in the autonomous driving scenario are input into the trained detection model to detect various vehicles. The example effect diagram is as Figure 7 shown. The results show that regardless of whether the target in the image is occluded or truncated, the detection method of the present invention can accurately identify the types of objects in the picture and precisely complete the target detection task, verifying the effectiveness of the target detection using the night vehicle detection method based on the optimized YOLOv11s.
Claims
1. A method for detecting night-time vehicles based on an optimized YOLOv11 model, characterized in that, It includes the following steps: (1) Construct a nighttime vehicle detection dataset, including nighttime scene images screened from the BDD100K dataset and a self-built domestic nighttime driving dataset, and perform resolution unification and annotation processing on all images; (2) Optimize the feature extraction module of the YOLOv11s model, replace the conventional stride convolution of YOLOv11s with the SPD Conv module, and introduce the Monte Carlo attention module based on the dynamic mask mechanism (DM-MCAttn) to enhance the feature selection ability; (3) Optimize the feature fusion module of the YOLOv11s model, including introducing the large kernel selection module (LSK Block), the parallel patch-aware attention module (PPA Block), the hierarchical offset Dysample upsampling operator, and the improved bidirectional feature pyramid network (D-BiFPN); (4) Optimize the detection head of the YOLOv11s model, replace the conventional convolution of the original regression branch with the H-MBConv module, and combine the ECA channel attention mechanism to reduce the number of parameters; (5) Use the training set to train the optimized model, and evaluate the detection performance through the validation set and the test set to complete the vehicle detection task in the nighttime scene.
2. The night vehicle detection method based on the optimized YOLOv11 model according to claim 1, wherein The dataset described in step (1) consists of the following: 12,500 nighttime scene images screened from the BDD100K dataset, excluding blurred and vehicle-free frames; the self-built domestic nighttime driving dataset contains 3,500 domestic nighttime driving images, and generates vehicle category and bounding box labels in YOLO format through the LabelImg annotation tool; the training set, validation set, and test set are divided in the ratio of 8:1:1 to ensure balanced class distribution.
3. A nighttime vehicle detection method based on an optimized YOLOv11 model according to claim 1, characterized in that, In step (2), the SPD Conv module realizes lossless downsampling through space-to-depth transformation, retaining the local features of small-scale targets; the Monte Carlo attention module based on the dynamic mask mechanism adjusts the Monte Carlo attention module through dynamic mask weights, as shown in formulas (1) and (2): A mask = σ(W4 · δ(W3 · x)) (1) A final = A × A mask (2) Where W3 and W4 are 1×1 weight matrices, δ(·) represents the ReLU activation function to increase the non-linear expression ability, σ(·) represents the Sigmoid activation function to normalize the weight result to [0,1], and A is the attention weight obtained by the Monte Carlo attention module. The two are combined to obtain the final weighted weight A final , which is the final Monte Carlo attention module based on the dynamic masking mechanism, and then input into the SE layer and subsequent weighted processing of the input x 4. A nighttime vehicle detection method based on an optimized YOLOv11 model according to claim 1, wherein In step (3), the large kernel selection module uses 5×5 convolution and 7×7 dilated convolution to extract multi-scale context information, and dynamically fuses it through channel attention. The parallel patch-aware attention module divides the input features into local patches, calculates the similarity weights in combination with the global template, and enhances the local detail perception of occluded targets. The improved bidirectional feature pyramid network fuses shallow details and high-level semantic features, and optimizes the multi-scale feature competition problem through weighted bidirectional connection. The hierarchical offset Dysample upsampling operator realizes high-fidelity upsampling through a two-stage offset grid, as shown in formulas (3) and (4): Among them, O (1) represents the original offset obtained by inputting the feature map into a 1*1 convolution, and O' (1) is the final offset obtained in the above-mentioned one-time offset process.
5. A nighttime vehicle detection method based on an optimized YOLOv11 model according to claim 1, characterized in that, In step (4): The H-MBConv module mixes depthwise separable convolution and standard 3×3 convolution in the regression branch to balance the receptive field and computational efficiency, replaces the SE module with the ECA module, and compresses the computational amount of channel attention through one-dimensional convolution.
6. The method for detecting nighttime vehicles based on an optimized YOLOv11 model according to claim 1, wherein, The verification of the autonomous driving target detection network model in the step (4) specifically refers to: setting the number of batch-processed pictures in the step (3) to 4, selecting the number of training iterations to be 200 rounds, setting the initial learning rate to 0.01, using the Stochastic Gradient Descent (SGD) optimizer, setting the momentum to 0.937, setting the weight update decay to 0.0005, and setting the Intersection over Union (IOU) to 0.6 for verification.
Citation Information
Cited By
Semantic fingerprint adaptive training method for teaching service robot
CN120653994A
Crown block hook identification method and system based on YOLOv8
CN120807959A