Night vehicle detection method and system and storage medium

Through the improved YOLOv5 network and low-light image enhancement algorithm, combined with the bidirectional feature pyramid structure and vehicle morphology random transformation module, the accuracy and robustness of vehicle detection in low-illumination environments at night are solved, and the detection accuracy and stability are improved.

CN120298980AInactive Publication Date: 2025-07-11HUBEI CENT CHINA TECH DEV OF ELECTRIC POWER +2
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510271391.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-08
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing vehicle detection methods have deteriorated performance in low-illumination environments at night, especially the detection effect of driving vehicles is poor, making it difficult to effectively distinguish between vehicles and backgrounds. The existing low-light enhancement technology has limited effects in this scenario.

Method used

The YOLOv5 network adopts a bidirectional feature pyramid structure, combining low-light image enhancement algorithm and vehicle morphology random transformation module, improves the accuracy and robustness of night vehicle detection through improved training strategies and feature fusion technology.

Benefits of technology

It improves the accuracy and stability of vehicle detection at night, reduces the missed detection rate, enhances the detection performance of the model in complex lighting environments, and optimizes the computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298980A_ABST
    Figure CN120298980A_ABST
Patent Text Reader

Abstract

The invention discloses a night vehicle detection method and system and a storage medium. The method comprises the following steps: acquiring a night vehicle detection data set; based on a low-light image enhancement algorithm, performing illumination enhancement operation preprocessing on image data in the night vehicle detection data set; the method comprises the following steps: modifying a training strategy of an original YOLOv5 network, and adding a vehicle form random transformation module on the basis of an original Mosaic enhanced training strategy; the preprocessed image data are transmitted to a YOLOv5-ours network for training, and optimal weight data are obtained; and loading the optimal weight data into a YOLOv5-ours network, and detecting the input image data to be detected. According to the night vehicle detection method, the bidirectional feature pyramid structure is adopted, multi-scale feature information can be fully extracted and fused, information interaction between feature maps is enhanced, and therefore efficient and accurate detection of night vehicles is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and particularly relates to a method and system for detecting vehicles at night and a storage medium. Background Art

[0002] Human-centered computer vision tasks, such as vehicle detection, vehicle re-identification, vehicle search, pose estimation, and face detection, have attracted much attention in the past decade. Among these tasks, vehicle detection is undoubtedly the most fundamental and crucial one. Its goal is to accurately locate and classify all vehicle instances in a given image. Vehicle detection not only provides a basis for advanced tasks such as autonomous driving, intelligent transportation, and traffic monitoring, but also is particularly important because it is directly related to people's lives. Therefore, the requirements for the detection rate and real-time performance of vehicle detection are extremely strict.

[0003] Currently, there are various methods for vehicle detection, including manual methods, deep learning methods, and hybrid applications of the two. These methods generally include three consecutive steps: proposal generation, classification (and regression), and post-processing. However, not all methods follow this fixed process. For example, the method proposed by Liu et al. (2019) omits the proposal generation step. With the continuous progress of neural network technology, methods based on deep features have far exceeded traditional methods based on manual features in terms of performance. In recent years, pure CNN-based vehicle detection methods have emerged in an endless stream, such as scale-aware methods, part-based methods, attention-based methods, data augmentation-based methods, loss-based methods, multi-task methods, and post-processing-based methods. The powerful representation learning ability of deep convolutional neural networks (CNNs) has brought significant progress to vehicle detection. Detectors such as Cascade R-CNN, Faster R-CNN, BGCNet, PedHunter, F2DNet, MAPD, and SMPD have all shown excellent performance in existing datasets under well-lit conditions.

[0004] However, for low-light night scenes, vehicle detection faces many challenges. Nighttime images are usually noisy and have poor lighting conditions. The instability of monitoring light sources further exacerbates the uneven illumination, resulting in drastic changes in image contrast. When detecting vehicles in such low-contrast areas, color or shape information is often easily lost, making it extremely difficult to distinguish between foreground and background areas. Therefore, even the most advanced vehicle detectors generally experience a decline in performance at night, and detection metrics such as precision mAP and miss rate MR are often not as ideal as during the day. However, considering the higher risk of nighttime driving and the loopholes in nighttime safety monitoring, the research on nighttime vehicle detection is particularly important, and more and more researchers are starting to focus on the vehicle detection problem in this low-light scene.

[0005] The research on low-light vehicle detection can be roughly divided into two directions: one is the input enhancement using fusion strategies, and the other is the optimization of vehicle detection model algorithms. The input enhancement can be further subdivided into a CNN-based fusion scheme and a graph attention network (GAN)-based fusion scheme. The vehicle detection model mainly focuses on how to use deep learning methods to detect vehicles under low-light conditions, and this part of the research can be divided into two major categories: region-based methods and non-region-based methods.

[0006] The structure of the Yolov5 network is mainly divided into three main parts: Backbone, Neck, and Head. This structure is a common model division method in visual deep learning.

[0007] Backbone: This is the backbone network of the model, responsible for extracting features from the input image. In YOLOv5, the Backbone uses a series of convolutional and deconvolutional layers to extract features, and uses residual connections and bottleneck structures to optimize the size of the network and improve performance. In addition, the C3 module is also used as the basic building block in the Backbone.

[0008] Neck: This part is an intermediate layer used to fuse the features from the Backbone to improve the performance of the model. The Neck part of YOLOv5 uses multi-scale feature fusion technology to fuse the feature maps from different stages of the Backbone, thereby enhancing the feature representation ability. Specifically, the Neck part of YOLOv5 includes the FPN module, PAN module, Conv module, Upsample module, and Concat module. However, it should be noted that in some descriptions of YOLOv5, the Neck part is not clearly divided, but its function is incorporated into the Head part.

[0009] Head: This is the last layer of the model, and its structure varies according to different tasks. In the object detection task, the Head part usually includes a bounding box regressor and a classifier. The Head part of YOLOv5 adopts a decoupled head design, that is, each scale has an independent detector, and each detector consists of a group of convolutional and fully connected layers for predicting the bounding boxes at that scale. This design enables the model to perform object detection in parallel at different scales.

[0010] Although the Yolov5 network achieves a good balance between accuracy and speed in vehicle detection in general scenarios, in low-light scenarios such as at night, like other deep learning algorithms, it still has certain detection defects and urgently needs further improvement and optimization.

[0011] The goal of low-light image enhancement (LLIE) technology is to improve the visibility and interpretability of images captured in poorly lit environments. In recent years, significant progress has been made in this field, mainly due to deep learning-based solutions. These solutions involve various learning strategies, network architectures, loss functions, and training data, etc.

[0012] Traditional low-light enhancement methods include histogram equalization-based methods (such as the research of H. Ibrahim et al., 2007) and Retinex model-based methods (such as the research of X. Guo et al., 2017; X. Ren et al., 2020; M. Li et al., 2018). Among them, Retinex model-based methods have received more attention. Such methods usually decompose low-light images into reflection components and illumination components through prior knowledge or regularization means, and use the estimated reflection components as the enhancement results.

[0013] However, in recent years, deep learning-based LLIE methods have achieved remarkable results. Compared with traditional methods, deep learning-based solutions have shown significant advantages in terms of accuracy, robustness, and speed. The learning strategies adopted in these solutions are diverse, including supervised learning (SL), reinforcement learning (RL), unsupervised learning (UL), zero-shot learning (ZSL), and semi-supervised learning (SSL), etc. For example, the GSAD research team proposed a diffusion-based framework to solve the low-light image enhancement problem, improving the performance of low-light enhancement from the perspective of regularized ODE trajectories. Compared with the state-of-the-art methods, this algorithm has made significant progress in image quality, noise suppression, and contrast amplification, etc.

[0014] However, when applying the above low-light enhancement technology to night-time vehicle detection, some challenges will be encountered. Specifically, for stationary vehicles on the roadside, this technology can indeed effectively capture more detailed information through low-light enhancement; but for vehicles moving on the road, due to the glare caused by the direct illumination of vehicle lights on the monitoring equipment, the body information is severely lost. In this case, further adopting low-light enhancement measures will not achieve better results. Summary of the Invention

[0015] Aiming at the vehicle images captured by night-time cameras, due to problems such as dim light and uneven illumination, it is difficult to distinguish vehicles from the background, and challenges such as insufficient information interaction between output feature maps, the present invention proposes a night-time vehicle detection method, aiming to accurately detect vehicle targets in night-time scenes, and while ensuring the detection speed, effectively solve the technical defects of existing algorithms in night-time scenes. By adopting a bidirectional feature pyramid structure, the present invention can fully extract and fuse multi-scale feature information, enhance the information interaction between feature maps, and thus achieve efficient and accurate detection of night-time vehicles.

[0016] A nighttime vehicle detection method disclosed by the present invention includes the following steps:

[0017] Step S1, obtaining a nighttime vehicle detection data set;

[0018] Step S2, based on a low-light image enhancement algorithm, performing preprocessing of illumination enhancement operation on the image data in the nighttime vehicle detection data set;

[0019] Step S3, modifying the training strategy of the original YOLOv5 network, adding a vehicle shape random transformation module on the basis of the original Mosaic enhancement training strategy. The vehicle shape random transformation module improves the training effect by randomly inserting other vehicles on other pictures in the same scene into random positions of the current picture. The vehicle shape random transformation module is used to train and build the YOLOv5-ours network suitable for vehicle detection;

[0020] Step S4, inputting the image data preprocessed in Step S2 into the YOLOv5-ours network for training to obtain the best weight data;

[0021] Step S5, loading the best weight data into the YOLOv5-ours network to detect the input image data to be detected.

[0022] Further, the YOLOv5-ours network includes:

[0023] A backbone network, constructed based on the backbone network of the original YOLOv5, used for extracting image data features at different scales, including a starting layer, a feature extraction layer 1, a feature extraction layer 2, a feature extraction layer 3, and a feature extraction layer 4 connected in sequence;

[0024] A neck network, including a bidirectional feature pyramid network composed of multiple conv modules and C3 modules, used for feature fusion of the outputs of the feature extraction layer 2, the feature extraction layer 3, and the feature extraction layer 4 to obtain three feature maps at different scales;

[0025] A head network, constructed based on the head network of the original YOLOv5, used for detecting the three feature maps at different scales.

[0026] Further, Step S4 includes:

[0027] Step S401, setting training parameters, where the training parameters include learning rate momentum, initial learning rate, number of iterations, and weight decay coefficient;

[0028] Step S402, inputting the image data preprocessed in Step S2 into the YOLOv5-ours network built in Step S3;

[0029] Step S403: Train the YOLOv5-ours network using the stochastic gradient optimization algorithm. Adjust the learning rate and the number of iterations based on the change trends of the average precision and the loss until the changes in precision and loss reach a stable state, and determine the final learning rate and the number of iterations.

[0030] Step S404: Complete the training of the YOLOv5-ours network based on the learning rate and the number of iterations determined in Step S403 to obtain the optimal weight data.

[0031] Further, Step S1 includes:

[0032] Step S101: Obtain the VD-NUS dataset.

[0033] Step S102: Prune the image data based on the VD-NUS dataset, and randomly select a part from the remaining image data as the original dataset.

[0034] Step S103: Perform data augmentation operations on the original dataset in Step S102 based on data augmentation methods to obtain a nighttime vehicle detection dataset; Step S104: Divide the nighttime vehicle detection dataset into a training set and a validation set.

[0035] Further, Step S102 includes:

[0036] Delete the image data without vehicle annotations, and randomly select one-third of the image data from the remaining image data as the original dataset. The targets of the original dataset are divided into 4 categories: vehicle, cyclist, motorcyclist, and ignore. The "ignore" includes posters, sculptures, and vehicles that are difficult to separate in a group.

[0037] Further, the data augmentation method in Step S103 includes one of the Mosaic data augmentation method, image flipping method, image rotation method, image cropping method, image adding salt and pepper noise method, and color transformation method.

[0038] Further, the low-light image enhancement algorithm in Step S2 includes one of the global structure-aware diffusion process algorithm, zero-reference depth curve estimation algorithm, and unsupervised generative adversarial network algorithm.

[0039] Further, it also includes Step S6: Evaluate the performance of the trained YOLOv5-ours network through the average miss detection rate, specifically including:

[0040] Step S601: Perform preprocessing of light enhancement operations on the image data in the nighttime vehicle detection dataset based on different low-light image enhancement algorithms.

[0041] Step S602: Input the preprocessed image data in Step S601 into the YOLOv5-ours network for training to obtain corresponding weight data;

[0042] Step S603: Load the corresponding weight data into the YOLOv5-ours network to detect the input image data to be detected;

[0043] Step S604: Based on the detection result in Step S603, select the low-light image enhancement algorithm with the lowest missed detection rate through the result of the average missed detection rate.

[0044] A night vehicle detection system includes: a computer-readable storage medium and a processor;

[0045] The computer-readable storage medium is used to store executable instructions;

[0046] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the night vehicle detection method described above.

[0047] A non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the night vehicle detection method described above.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] (1) Based on the low-light image enhancement algorithm, the image data in the night vehicle detection data set is preprocessed, and only the areas with relatively dark light are enhanced, while the parts with too strong light are appropriately weakened, thereby realizing more refined image processing, making the vehicle features more obvious and prominent, thus reducing the miss rate (MR), helping to improve the detection accuracy and stability, and reducing the sensitivity of the preprocessed image data to the environment such as light changes and shadow interference, thereby enhancing the robustness.

[0050] (2) Based on the deep learning object detection model YOLOv5 network, a vehicle morphology random transformation module is introduced, aiming to increase the number of samples and the types of changes in the training set. This module significantly broadens the coverage of the training samples, laying a solid foundation for improving the detection performance of the model in complex glare environments. The optimized YOLOv5-ours network not only improves the detection accuracy but also optimizes the model complexity and computational efficiency to a certain extent, making the model more suitable for the actual vehicle target detection task. Description of the Drawings

[0051] Figure 1 It is a technical implementation flowchart of an embodiment of a night vehicle detection method of the present invention;

[0052] Figure 2 The network structure of the YoloV5-ours algorithm for an embodiment of a nighttime vehicle detection method of the present invention;

[0053] Figure 3 The network structure of the Light-Effects Suppression algorithm for an embodiment of a nighttime vehicle detection method of the present invention; Figure 4 is Figure 2 The structural schematic diagram of the Head part in Detailed implementation manners

[0054] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0055] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the specification of this application herein are only for the purpose of describing specific implementation manners and are not intended to limit this application. The term "or / " used herein includes any and all combinations of one or more of the related listed items.

[0057] In addition, in the present invention, descriptions such as "first" and "second" are only for descriptive purposes and do not particularly refer to the order or sequence. They are not used to limit the present invention. They are only used to distinguish components or operations described with the same technical terms, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those skilled in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0058] A nighttime vehicle detection method of this embodiment, such asFigure 1 As shown in Figure 1 , it includes the following steps:

[0059] Step S1: Obtain a nighttime vehicle detection dataset;

[0060] Specifically, it is implemented by the following method:

[0061] Step S101: Obtain the VD-NUS dataset;

[0062] The VD-NUS dataset is a comprehensive public nighttime vehicle detection dataset. The dataset is captured by 50 high-definition cameras and contains a total of 70,000 high-quality images. Each image has been carefully screened and annotated, with a total of 510,000 vehicle bounding boxes annotated, ensuring the accuracy and integrity of the data.

[0063] Step S102: Based on the VD-NUS dataset, delete some of the image data, and randomly select some of the remaining image data as the original dataset;

[0064] In the VD-NUS dataset, since there are fewer vehicles moving at night, the number of frames without vehicles at night is much higher than that during the day. To reduce these invalid pictures, first, all pictures without vehicle annotations are deleted. Then, one-third of the remaining image data is randomly selected as the original dataset. The targets of the original dataset are divided into 4 categories: vehicles, cyclists, motorcyclists, and ignored (people who are difficult to separate in posters, sculptures, and groups). It should be noted that this invention only analyzes the effect of vehicle detection.

[0065] Step S103: Based on the data augmentation method, perform data augmentation operations on the original dataset in step S102 to obtain a nighttime vehicle detection dataset;

[0066] The data augmentation method includes the Mosaic data augmentation method, image flipping method, image rotation method, image cropping method, image adding salt-and-pepper noise method, and color transformation method. In this embodiment, the Mosaic data augmentation method is adopted, and 4 pictures within a batch are combined to obtain a picture with more abundant information for detection.

[0067] Step S104: Divide the nighttime vehicle detection dataset into a training set and a validation set.

[0068] In this embodiment, the training set includes 63,000 image data, and the validation set includes 7,000 image data.

[0069] Step S2: Based on a low-light image enhancement algorithm (such as the Light-Effects Suppression algorithm), perform preprocessing of light enhancement operations on the image data in the nighttime vehicle detection dataset in step S1;

[0070] Step S3: Modify the training strategy of the original YOLOv5 network. On the basis of the original Mosaic augmentation training strategy, add a vehicle form random transformation module for training, and build the YOLOv5-ours network suitable for vehicle detection.

[0071] Specifically, as Figure 2 shown, it is implemented by the following method:

[0072] Backbone network: It is constructed based on the backbone network (Backbone network) of the original YOLOv5 and is used for extracting different-scale features of image data, including a starting layer (Stem Layer), a feature extraction layer 1 (Stem Layer 1), a feature extraction layer 2 (Stem Layer 2), a feature extraction layer 3 (Stem Layer 3), and a feature extraction layer 4 (Stem Layer 4) connected in sequence.

[0073] Starting layer (Stem Layer): It includes layer 0 and is set as a conv module.

[0074] Feature extraction layer 1 (Stem Layer 1) includes layer 1 and layer 2 connected in sequence. Layer 1 is set as a conv module connected to the output end of the starting layer (Stem Layer), and layer 2 is set as a C3 module connected to the conv module.

[0075] Feature extraction layer 2 (Stem Layer 2) includes layer 3 and layer 4 connected in sequence. Layer 3 is set as a conv module connected to the output end of the feature extraction layer 1 (Stem Layer 1), and layer 4 is set as a C3 module connected to the conv module.

[0076] Feature extraction layer 3 (StemLayer3) includes layer 5 and layer 6 connected in sequence. Layer 5 is set as a conv module connected to the output end of the feature extraction layer 2 (StemLayer2), and layer 5 is set as a C3 module connected to the conv module.

[0077] Feature extraction layer 4 (StemLayer4) includes layer 7, layer 8, and layer 9 connected in sequence. Layer 7 is set as a conv module connected to the output end of the feature extraction layer 3 (StemLayer3), layer 8 is set as a C3 module connected to the conv module, and layer 9 is set as a SPPF module connected to the C3 module.

[0078] In this embodiment, the conv module, C3 module, and SPPF module all adopt the existing technologies of the YOLOv5 network, which will not be elaborated here. The input image data is 640*640*3. After feature extraction by the backbone network, the output of the feature extraction layer 2 (StemLayer2) is an 80*80*256 feature map, the output of the feature extraction layer 3 (StemLayer3) is a 40*40*512 feature map, and the output of the feature extraction layer 4 (StemLayer4) is a 20*20*1024 feature map.

[0079] The neck network includes a bidirectional feature pyramid network composed of multiple conv modules and C3 modules, which is used to perform feature fusion on the outputs of the feature extraction layer 2, feature extraction layer 3, and feature extraction layer 4 to obtain three feature maps of different scales;

[0080] In this embodiment, the output ends of the feature extraction layer 2 (StemLayer2), feature extraction layer 3 (StemLayer3), and feature extraction layer 4 (StemLayer4) are respectively connected to the 10th conv module, 11th conv module, and 12th conv module.

[0081] Preferably, the parameter settings of the 10th conv module, 11th conv module, and 12th conv module are ch_out = 256, kernel = 1, stride = 1, and padding = 0. The sizes of the obtained feature maps are 80*80*256, 40*40*256, and 20*20*256 respectively.

[0082] The output end of the 12th layer is connected to the 13th layer, and the 13th layer is set as an upsampling module (Upsample module). The Upsample module is an existing technology and will not be elaborated here. In this embodiment, the sampling multiple of the upsampling module is set to 2, and the nearest neighbor interpolation sampling algorithm is used, and the output is a 40*40*256 feature map.

[0083] The output ends of the 13th layer and the 11th layer are connected to the 14th layer. The 14th layer is set as an FPN module. The FPN module performs weighted summation according to the weights learned by the feature map to obtain an integrated new feature map output with a size of 40*40*256. The output end of the 14th layer is connected to the 15th layer, and the 15th layer is set as a C3 module. In the C3 module of the 15th layer, there are 3 Bottleneck modules connected in sequence, as Figure 4 shown. In the Bottleneck module, the shortcut is connected in the false manner. The output channel of the C3 module of the 15th layer is set to 256, and the output is a 40*40*256 feature map.

[0084] The output of the 15th layer is connected to the 16th layer. The 16th layer is set as an upsampling module (Upsample module). The sampling multiple of the upsampling module is set to 2, and the nearest neighbor interpolation sampling algorithm is used. The output is an 80*80*256 feature map.

[0085] The outputs of the 16th layer and the 10th layer are connected to the 17th layer. The 17th layer is set as an FPN module. The FPN module performs weighted summation according to the weights learned by the feature map to obtain an integrated new feature map output with a size of 80*80*256.

[0086] The output of the 17th layer is connected to the 18th layer. The 18th layer is set as a C3 module, and the output is 80*80*256. The output of the 18th layer is connected to the 19th layer. The 19th layer is set as a conv module. The parameters of the conv module are set as ch_out = 256, kernel = 3, stride = 2, padding = 1, and the size of the obtained feature map is 40*40*256.

[0087] The outputs of the 10th layer, the 18th layer, and the 19th layer are connected to the 20th layer. The 20th layer is set as an FPN module. The FPN module performs weighted summation according to the weights learned by the feature map to obtain an integrated new feature map output with a size of 80*80*256.

[0088] The output of the 20th layer is sequentially connected to the 21st layer and the 22nd layer. The 21st layer is set as a C3 module, and the output channel of the C3 module is set to 256, and the size of the output feature map is 80*80*256. The 22nd layer is set as a conv module. The parameters of the conv module are set as ch_out = 256, kernel = 3, stride = 2, padding = 1. The size of the obtained feature map is 40*40*256.

[0089] The outputs of the 22nd layer, the 11th layer, and the 15th layer are connected to the 23rd layer. The 23rd layer is set as an FPN module, and a feature map with a size of 40*40*256 is obtained correspondingly.

[0090] The output of the 23rd layer is sequentially connected to the 24th layer and the 25th layer. The 24th layer is set as a C3 module, and the output channel of the C3 module is set to 512, and the size of the output feature map is 40*40*512. The 25th layer is set as a conv module. The parameters of the conv module are set as ch_out = 256, kernel = 3, stride = 2, padding = 1. The size of the feature map output by the 25th layer is 20*20*256.

[0091] The output ends of the 25th layer and the 12th layer are connected to the 26th layer. The 26th layer is set as an FPN module. The FPN module performs weighted summation according to the weights learned from the feature maps to obtain an output of a new integrated feature map with a size of 20*20*256. The output end of the 26th layer is connected to the 27th layer. The 27th layer is set as a C3 module. The output channel of the C3 module is set to 1024, and the size of the output feature map is 20*20*1024.

[0092] The head network is constructed based on the head network of the original YOLOv5 and is used to detect feature maps of three different scales.

[0093] As Figure 4 shown, the Head network adopts a decoupled head structure, separating the classification and detection heads, and also changing from Anchor-Based to Anchor-Free. The outputs of the 21st, 24th, and 27th layers, 80*80*256, 40*40*512, and 20*20*1024 respectively, are fed into three detection heads for detection. During detection, the positive and negative sample assignment strategy adopts a dynamic distribution strategy, that is, the weights are dynamically adjusted according to the progress of training and the characteristics of the samples. The matching strategy of TaskAlignedAssigner can be simply summarized as: selecting positive samples according to the scores weighted by the classification and regression scores. The core formula of TAL is as follows:

[0094] t = s α + u β

[0095] where s and u are the classification score and the IoU value respectively, and α and β are weight hyperparameters (set to α = 0.5, β = 6.0 in this example). Through training, t can achieve task alignment by optimizing both the classification score and the IoU, and use the weighted product of the classification quality and the regression quality to describe the prediction quality of the anchor, which can guide the network to dynamically focus on high-quality anchor prediction boxes.

[0096] In this embodiment, the loss function includes classification and regression branches. The classification branch adopts a binary cross-entropy loss function (BCEWithLogitsLoss) combined with the Sigmoid function.

[0097]

[0098] is the Sigmoid function, and log is the natural logarithm. p i represents the probability that sample x i is predicted as a positive example, and y i represents the true label of sample x i . In a binary classification problem, y iUsually 0 or 1, p(yi) is the probability that the output belongs to the label, and N represents the number of groups of objects predicted by the model.

[0099] The regression branch includes the Distribution Focal Loss function and the Complete Intersection over Union Loss function (CIoU Loss).

[0100] Distribution Focal Loss(S i ,S i+1 ) = -((y i + 1 - y)log(S i ) + (y - y i )log(S i+1 ))

[0101] y = the distance from the center to a certain side / the current downsampling factor, and after rounding, the predicted value y is obtained i , the adjacent predicted value y i + 1.

[0102] Where d is the distance between the centers of the predicted box and the ground truth box, and c is the diagonal distance of the minimum bounding rectangle. γ is a correction factor used to further adjust the loss function, considering the shape and orientation of the target box. The specific calculation method is (w G ,h G ),(w p ,h p ) are the widths and heights of the target box and the predicted box respectively.

[0103] CIoU Loss = 1 - CIoU

[0104] The ratio of the classification loss, the distribution focal loss, and the complete intersection over union loss is 0.5:1.5:7.5.

[0105] Step S4, input the image data preprocessed in step S2 into the YOLOv8-ours network for training to obtain the optimal weight data;

[0106] Specifically, it is implemented by the following method:

[0107] Step S401: Set the training parameters. The learning rate momentum is set to 0.937, the initial learning rate and the final learning rate are both set to 0.01, the number of iterations is set to 300, the weight decay coefficient is set to 0.0005, close_mosaic is set to 10, and the mosaic data augmentation method is disabled in each epoch of the last 10 training epochs. mask_ratio: 4, indicating that the mask is downsampled by 4 times; iou: 0.7, meaning that only when the intersection over union of two boxes is greater than 0.7, are they considered overlapping.

[0108] Step S402: Input the image data preprocessed in Step S2 into the YOLOv5-ours network built in Step S3;

[0109] Step S403: Train the YOLOv5-ours network using the Stochastic Gradient Descent (SGD) algorithm. Adjust the learning rate and the number of iterations according to the average precision change and the loss change trend until the precision change and the loss change are in a stable state, and determine the final learning rate and the number of iterations;

[0110] Step S404: Complete the training of the YOLOv5-ours network based on the learning rate and the number of iterations determined in Step S403 to obtain the optimal weight data.

[0111] Step S5: Load the optimal weight data into YOLOv5-ours to detect the input image data to be detected.

[0112] This embodiment further includes Step S6: Evaluate the performance of the trained YOLOv5-ours network through the average miss detection rate, specifically including:

[0113] Step S601: Perform preprocessing of light enhancement operations on the image data in the night vehicle detection dataset based on different low-light image enhancement algorithms;

[0114] Step S602: Input the image data preprocessed in Step S601 into the YOLOv5-ours network for training respectively to obtain the corresponding weight data;

[0115] Step S603: Load the corresponding weight data into the YOLOv5-ours network to detect the input image data to be detected;

[0116] Step S604: Based on the detection results in Step S603, select the low-light image enhancement algorithm with the lowest miss detection rate through the results of the average miss detection rate.

[0117] Mean Average Precision (mAP) is used as an evaluation metric. mAP is a key metric for evaluating the performance of algorithms in fields such as object detection and image recognition. It is obtained by calculating the average precision (AP) for each class and taking the average, which reflects the prediction accuracy and recognition ability of the model for positive samples. The larger the value of mAP, the better the performance of the vehicle detector.

[0118] In the present invention, through the vehicle form random transformation module in the yolov5-ours network, compared with the baseline, the mAP has increased by nearly four points.

[0119] The vehicle form random transformation module is implemented as follows:

[0120] First, the present invention randomly pairs the input image with another image in the same scene to form an image pair. Then, random scale jittering and level flipping are applied to these two images respectively to simulate the perspective changes and light fluctuations in the real world. On this basis, the present invention further randomly selects a subset of vehicles from one of the paired images and, using advanced image processing techniques, seamlessly "pastes" these vehicle objects to the corresponding positions in the other image. This process not only requires a high degree of accuracy to maintain the coherence of the scene but also ensures the integration of the pasted vehicles with the surrounding environment, thus simulating a more realistic traffic scene.

[0121] In the yolov5-ours network of the present invention, for the images enhanced by the Light-EffectsSuppression algorithm for low light, compared with the baseline, the mAP has increased by nearly one point.

[0122] In this embodiment, the Light-EffectsSuppression algorithm is used to preprocess the nighttime vehicle detection dataset in step S2.

[0123] As Figure 3 shown, the Light-EffectsSuppression algorithm includes the following parts:

[0124] The decomposition of the present invention is based on the following image layer model

[0125] I = R⊙L + G

[0126] where I represents the input nighttime image, G represents the light effect layer, and R and L represent the reflection layer and the shadow layer respectively. The symbol ⊙ represents element-wise multiplication. In this equation, the present invention assumes a linear gamma function. However, this equation is not directly used in the method of the present invention. Instead, the present invention only uses it to guide Figure 3Design of the middle network (i.e., the layer decomposition network). When using non-linear images with non-linear gamma functions during training, the background scene is an approximation of the physically correct value. The decomposition objective of the present invention is to obtain a background scene that is not affected by the lighting effect, that is, the present invention wants to estimate the background scene Jinit = R ⊙ L.

[0127] The decomposition network is based on the image layer model proposed by the present invention in the equation. Given the input image (I), the present invention first performs image decomposition. The present invention uses three independent networks and a novel unsupervised loss function proposed by the present invention to separately obtain the lighting effect (G), shadow (L), and reflectance (R) layers.

[0128] Learning the lighting effect, shadow, and reflectance layers: To obtain the lighting effect (G), shadow (L), and reflectance (R) layers, the present invention uses three networks respectively: the lighting effect network the shadow network and the reflectance network where these three networks are trained using unsupervised loss, which will be discussed in the subsequent paragraphs.

[0129] Initialization of the lighting effect and shadow: To solve the decomposition ambiguity problem, it is very important to provide appropriate initial estimates for the layers. For the shadow layer, the present invention adopts the shadow map Li obtained by taking the maximum value of the three color channels of each pixel. For the lighting effect layer, the present invention uses the lighting effect map Gi, which is calculated using a relatively smooth technique. This is extracted from the input image using a second-order Laplacian filter because the lighting effect changes smoothly. The present invention defines the loss function of the initialization step as:

[0130] L init = |G - G i | l + |L - L i | l

[0131] Gradient exclusion loss: The gradient of the lighting effect layer has a short-tail distribution, similar to the gradient distribution of the "glowing" phenomenon. In contrast, the gradient of the background image has a long-tail distribution. Therefore, the present invention adopts gradient exclusion loss to recover the uncorrelated layers {G, Jinit}, and its objective is to separate these two layers as much as possible in the gradient space. The definition of this loss follows:

[0132]

[0133] Color constancy loss: To minimize any color shift in the decomposition output of the present invention, inspired by the gray world assumption, the present invention uses a color constancy prior, which encourages the intensity value ranges of the three color channels in the background image Jinit to remain balanced:

[0134]

[0135] Reconstruction Loss: For the decomposition task of the present invention, recombining the estimated layers should be able to recover the original input image. Therefore, the present invention defines the reconstruction loss as:

[0136] L recon = I - (R⊙L + G) l

[0137] This method combines a layer decomposition network and a light effect suppression network. Given a night image as input, the decomposition network of the present invention learns to decompose the shadow, reflectance, and light effect layers under the guidance of an unsupervised layer-specific prior loss. The light effect suppression network of the present invention further suppresses the light effect while enhancing the illumination of dark areas. This light effect suppression network uses the estimated light effect layer as a guide and focuses on the light effect area. To restore background details and reduce hallucinations / artifacts, the present invention proposes a structure and high-frequency consistency loss.

[0138] On the other hand, the present invention provides a night vehicle detection system, including: a computer-readable storage medium and a processor;

[0139] The computer-readable storage medium is used to store executable instructions;

[0140] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the night vehicle detection method described in the first aspect.

[0141] On the other hand, the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the night vehicle detection method described in the first aspect.

[0142] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0143] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more flows and / or one or more blocks. Figure 1 in one or more flows and / or one or more blocks Figure 1 in the block or blocks.

[0144] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more flows and / or one or more blocks. Figure 1 in one or more flows and / or one or more blocks Figure 1 in the block or blocks.

[0145] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or one or more blocks. Figure 1 in one or more flows and / or one or more blocks Figure 1 in the block or blocks.

[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific embodiments of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A method for detecting vehicles at night, characterized in that, It includes the following steps: Step S1, obtain the night vehicle detection dataset; Step S2, based on the low-light image enhancement algorithm, perform preprocessing of light enhancement operation on the image data in the night vehicle detection dataset; Step S3, modify the training strategy of the original YOLOv5 network, and add a vehicle form random transformation module on the basis of the original Mosaic enhancement training strategy. The vehicle form random transformation module improves the training effect by randomly inserting other vehicles on other pictures in the same scene into random positions of the current picture. The vehicle form random transformation module is used to train and build the YOLOv5-ours network suitable for vehicle detection; Step S4, input the image data preprocessed in Step S2 into the YOLOv5-ours network for training to obtain the optimal weight data; Step S5, load the optimal weight data into the YOLOv5-ours network to detect the input image data to be detected.

2. The method for detecting vehicles at night according to claim 1, characterized in that, The YOLOv5-ours network includes: The backbone network, constructed based on the backbone network of the original YOLOv5, is used for extracting image data features at different scales, including a starting layer, a feature extraction layer 1, a feature extraction layer 2, a feature extraction layer 3, and a feature extraction layer 4 connected in sequence; The neck network, including a bidirectional feature pyramid network composed of multiple conv modules and C3 modules, is used for feature fusion of the outputs of the feature extraction layer 2, the feature extraction layer 3, and the feature extraction layer 4 to obtain three feature maps at different scales; The head network, constructed based on the head network of the original YOLOv5, is used for detecting the three feature maps at different scales.

3. A nighttime vehicle detection method according to any one of claims 1, characterized in that, Step S4 includes: Step S401, set the training parameters, and the training parameters include learning rate momentum, initial learning rate, number of iterations, and weight decay coefficient; Step S402, input the image data preprocessed in Step S2 into the YOLOv5-ours network built in Step S3; Step S403, use the stochastic gradient optimization algorithm to train the YOLOv5-ours network, and adjust the learning rate and the number of iterations according to the average precision change and the loss change trend until the precision change and the loss change are in a stable state, and determine the final learning rate and the number of iterations; Step S404, based on the learning rate and the number of iterations determined in Step S403, complete the training of the YOLOv5-ours network to obtain the optimal weight data.

4. The method for detecting vehicles at night according to claim 1, characterized in that, Step S1 includes: Step S101, obtain the VD-NUS dataset; Step S102, based on the VD-NUS dataset, delete the image data, and randomly select a part from the remaining picture data as the original dataset; Step S103, based on the data augmentation method, perform data augmentation operation on the original dataset in Step S102 to obtain the night vehicle detection dataset; Step S104, divide the night vehicle detection dataset into a training set and a validation set.

5. A method for detecting vehicles at night according to claim 4, characterized in that: Step S102 includes: Delete the image data without vehicle annotations, and randomly select one-third of the remaining image data as the original dataset. The targets of the original dataset are divided into four categories: vehicles, cyclists, motorcyclists, and ignored. The ignored category includes posters, sculptures, and vehicles that are difficult to separate in a group.

6. A method for detecting vehicles at night according to claim 4, characterized in that: The data augmentation method in step S103 includes one of the Mosaic data augmentation method, image flipping method, image rotation method, image cropping method, image adding salt and pepper noise method, and color transformation method.

7. A method for detecting vehicles at night according to claim 1, characterized in that: The low-light image enhancement algorithm in step S2 includes one of the global structure-aware diffusion process algorithm, zero-reference depth curve estimation algorithm, and unsupervised generative adversarial network algorithm.

8. A method for detecting vehicles at night according to claim 7, characterized in that, It also includes step S6, which evaluates the performance of the trained YOLOv5-ours network through the average miss detection rate, specifically including: Step S601, based on different low-light image enhancement algorithms, perform preprocessing of light enhancement operations on the image data in the night vehicle detection dataset; Step S602, input the preprocessed image data in step S601 into the YOLOv5-ours network for training respectively to obtain corresponding weight data; Step S603, load the corresponding weight data into the YOLOv5-ours network to detect the input image data to be detected; Step S604, based on the detection results in step S603, select the low-light image enhancement algorithm with the lowest miss detection rate through the result of the average miss detection rate.

9. A night vehicle detection system, comprising: A computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the night vehicle detection method described in any one of claims 1-8.

10. A non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the night vehicle detection method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Automatic driving scene vehicle detection method based on improved YOLOv5

    CN116434170A

  • Night pedestrian detection method, computer equipment, device and storage medium

    CN118570760A

  • Target detection method, vehicle and storage medium

    CN118692050A

  • Parking space state identification method based on improved YOLOv5s model

    CN119478889A

  • Target object detection method and apparatus, electronic device, and readable storage medium

    WO2024139763A1