Night vehicle detection system and method

By using a dual-branch frame and feature similarity perception module in night vehicle detection, the problem of low accuracy of night vehicle detection in the prior art is solved, and higher detection accuracy and reliability are achieved.

CN120014564APending Publication Date: 2025-05-16SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510032302.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the detection of night vehicle, the prior art has low detection accuracy due to excessive reliance on offline light recognition results and insufficient utilization of reflective area characteristics on the vehicle.

Method used

A dual-branch framework is adopted, combining vehicle detection branches and highlighted area segmentation branches, and collaborative learning and information exchange between each branch is achieved through feature similarity perception module. This module optimizes the feature comparison process through window division strategy and feature similarity calculation, and deeply analyzes the channel and spatial distribution of features through highlighted area segmentation branches to identify the significant highlighted area of ​​the vehicle.

Benefits of technology

It significantly reduces the impact of complex lighting conditions at night on vehicle detection, improves the accuracy and reliability of vehicle detection, and enhances the accuracy of vehicle identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014564A_ABST
    Figure CN120014564A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and discloses a night vehicle detection system and method, and the system comprises a vehicle detection branch which is used for receiving an input image and outputting a detection result, and a region segmentation branch which is used for receiving the input image and outputting a segmentation result. The vehicle detection branch comprises i-stage cascaded vehicle detection feature collectors and a detection head in communication connection with the last-stage vehicle detection feature collector, the region segmentation branch comprises i-stage cascaded region segmentation feature collectors and a segmentation head in communication connection with the last-stage region segmentation feature collector, and the detection head is in communication connection with the segmentation head. The problem that in the prior art, the accuracy of vehicle detection under the night condition is low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and in particular to a nighttime vehicle detection system and method. Background Art

[0002] Vehicle detection at night is a key research topic in the field of intelligent transportation systems (ITS) and has received widespread attention in the industry in recent years. Systems designed to improve vehicle detection capabilities under nighttime conditions are of great significance for enhancing the perception of road networks, reducing traffic accidents, and reducing labor costs. However, compared with daytime, vehicle detection at night faces greater challenges due to the complexity of the lighting environment. As two core research topics in the field of nighttime vehicle analysis, nighttime vehicle detection and headlight recognition have made significant progress in recent years, mainly due to the widespread application of deep learning in object detection and segmentation models. In order to cope with the problem of nighttime detection, a variety of techniques have been proposed, including low-light image enhancement and style transfer. Among them, the main purpose of image enhancement is to improve the contrast between the vehicle and the background, while style transfer uses generative adversarial networks (GANs) to achieve the conversion between daytime and nighttime images.

[0003] However, these methods may introduce noise or change the original visual information during the style transfer process, thus affecting the detection effect. The latest night analysis methods based on deep learning use headlight recognition as an important basis for vehicle positioning, but over-reliance on offline headlight recognition results may lead to the introduction of noise, thus affecting the accuracy of vehicle detection. In addition, the significant reflective area on the vehicle, as an important feature of vehicle recognition, has not been fully taken into account in existing methods. Summary of the invention

[0004] In order to overcome the deficiencies of the prior art, the present invention provides a nighttime vehicle detection system and method, which solve the problems of low accuracy in vehicle detection under nighttime conditions in the prior art.

[0005] The technical solution adopted by the present invention to solve the above problems is:

[0006] A nighttime vehicle detection system includes a vehicle detection branch for receiving an input image and outputting a detection result, and a region segmentation branch for receiving an input image and outputting a segmentation result. The vehicle detection branch includes an i-stage cascaded vehicle detection feature collector and a detection head communicatively connected to the last-stage vehicle detection feature collector. The region segmentation branch includes an i-stage cascaded region segmentation feature collector and a segmentation head communicatively connected to the last-stage region segmentation feature collector. The detection head is communicatively connected to the segmentation head.

[0007] As a preferred technical solution, it includes a feature similarity perception module which is respectively communicated with the vehicle detection branch and the region segmentation branch. The feature similarity perception module is used to compare the similarity between the features of the vehicle detection branch and the features of the region segmentation branch.

[0008] As a preferred technical solution, when the vehicle detection branch is working, the input image is divided into multiple local windows of k×k dimensions, and the feature comparison between the vehicle detection branch and the region segmentation branch is limited to be performed only within their respective corresponding windows; during the comparison process, the features of the current branch are regarded as queries, and the features of the other branch are regarded as keys; wherein k is the length or width of the local window.

[0009] As a preferred technical solution, the calculation formula of the feature similarity perception module is:

[0010]

[0011] or,

[0012]

[0013] In the formula, A represents the vehicle detection feature vector, B represents the region segmentation feature vector, and M att Indicates the similarity between vehicle detection features and regional segmentation features. represents the operation of calculating the mean across the channel dimension, Indicates a window reversal operation, represents the window partitioning operation, d represents the dimension of the matrix, T represents the matrix transpose calculation, and Ones_like(·) represents initializing a matrix vector whose element values ​​are all 1 and whose number of channels, width, and height are the same as A or B.

[0014] As a preferred technical solution, when the region segmentation branch works, first, the output of the feature similarity perception module is combined with the features of the region segmentation branch through a pixel-by-pixel multiplication operation; then, the number of channels of the combined features is halved through a convolution layer with a kernel size of 1; then, the weight distribution of the features with the number of channels halved in different channels is adjusted through a group of convolution layers with a kernel size of 3 and a convolution layer with a kernel size of 1; then, the weight distribution of the features at different positions after the weight distribution of the features with the number of channels halved in different channels is adjusted using convolution layers with different kernel sizes; finally, the spatial distribution weights are generated by pixel-by-pixel classification through a convolution layer with a kernel size of 1.

[0015] As a preferred technical solution, the calculation formula of the region segmentation branch is:

[0016]

[0017] In the formula,

[0018]

[0019] Among them, represents the output of the region segmentation branch, Indicates feature enhancement from the channel direction. Indicates feature enhancement from the spatial direction. represents the spatially aware convolution operation, represents the channel-aware convolution operation, and σ(·) represents the sigmoid function.

[0020] As a preferred technical solution, the nighttime vehicle detection system can train the detection head and the segmentation head; during the training, the vehicle detection feature vector and the region segmentation feature vector are updated respectively by the following formulas:

[0021]

[0022] Among them, A' represents the updated vehicle detection feature vector, and B' represents the updated region segmentation feature vector.

[0023] As a preferred technical solution, the training loss function Including detection loss and segmentation loss The calculation formula is:

[0024]

[0025] Here, λ represents a hyperparameter.

[0026] As a preferred technical solution, The calculation formula is:

[0027]

[0028] in, Represents the loss calculation item of classification recognition in the detection task, Represents the loss calculation term for location recognition in the detection task.

[0029] A nighttime vehicle detection method adopts the nighttime vehicle detection system to perform nighttime vehicle detection.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] (1) This paper proposes a dual-branch framework that significantly reduces the challenges of vehicle detection under complex lighting conditions at night by effectively integrating vehicle highlight information;

[0032] (2) The present invention proposes a feature similarity perception attention module to promote collaborative learning and information exchange between branches within the framework. This module makes full use of the multi-level features extracted by each branch and captures the similarity between features, thereby achieving effective information exchange and fusion between the vehicle detection branch and the highlight area segmentation branch;

[0033] (3) The present invention designs a highlight region segmentation branch, which optimizes the spatial distribution and channel distribution of features, and can accurately identify and predict the significant highlight regions of the vehicle, providing more reliable information support for vehicle detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic diagram of the overall framework of the present invention;

[0035] Figure 2 for Figure 1 One of the partial enlarged pictures;

[0036] Figure 3 for Figure 1 The second partial enlarged image. DETAILED DESCRIPTION

[0037] The present invention will be further described in detail below in conjunction with embodiments and drawings, but the embodiments of the present invention are not limited thereto.

[0038] Example 1

[0039] like Figures 1 to 3 As shown, the present invention proposes an innovative nighttime vehicle detection scheme, which uses the information of the vehicle highlight area (mainly including the light of the headlights and the external reflective surface of the vehicle) as the key guide to significantly improve the accuracy of vehicle recognition. Specifically, the present invention regards the vehicle detection task and the vehicle highlight area recognition task as two independent and interdependent tasks, and conducts joint training. The adoption of this joint training strategy enables the two tasks to exchange information and enhance each other during the feature learning representation process, thereby effectively improving the execution efficiency of each task. To achieve the above goals, the present invention designs a dual-branch architecture, which fully considers the differences in the focus features between vehicle detection and vehicle highlight area segmentation. In addition, the present invention incorporates a feature enhancement module into the backbone network, aiming to effectively identify the positional features of vehicle highlight information at different levels. This module guides the model to focus more on significant and recognizable features by adjusting the distribution of feature weights, while effectively weakening the interference of background information.

[0040] The goal of this invention is to identify a vehicle input image in a nighttime traffic scene, and the recognition process includes the position recognition of the corresponding vehicle and the segmentation of the highlight area on the vehicle. The highlight area is defined as the pixels on the vehicle with a grayscale value greater than 246. The overall architecture is as follows Figure 1 shown.

[0041] The overall process of the framework consists of two branches: vehicle detection branch and vehicle highlight region segmentation branch (i.e., region segmentation branch). The model input includes an RGB image, and its output consists of two independent parts: one is the vehicle detection result from the vehicle detection branch, and the other is the vehicle highlight region generated by the highlight region perception branch. The feature similarity perception module uses the information from the backbone network (A 1 -A i , B 1 -B i ) The features of different stages are used as input to analyze the similarity between features; where i represents the level of the feature, A i represents the i-th level vehicle detection feature, B i Represents the i-th level region perception feature. The highlight region segmentation branch can be used to perceive the highlight region. Figure 1 middle, The transmission direction of vehicle detection features is shown. The transmission direction of the area-aware features is shown (the arrow shapes are different for distinction).

[0042] 1. Implementation of feature similarity perception module

[0043] The visual representation of a vehicle is significantly affected by lighting conditions, and headlights become a prominent feature of the vehicle at night. Previous nighttime vehicle detection methods have mitigated the interference caused by background lighting by identifying headlights. Traditionally, a two-stage approach has been adopted, which first detects the headlights and then identifies the vehicle itself. In this paper, a parallel approach is proposed to solve these two tasks simultaneously by integrating feature similarity perception modules to improve their interdependent learning process.

[0044] Let A∈R C×W×H and B∈R C×W×H The sizes are the same (i.e., C, W, and H are the same), respectively representing the relevant features of the detection branch and the segmentation branch at a certain stage in the backbone network. Among them, C, W, and H represent the number of channels, width, and height of the feature map, respectively. A represents the vehicle detection feature vector, B represents the region perception feature vector, and A i is an element in A, B iis an element in B. Directly calculating the similarity between A and B not only consumes a lot of resources, but is also highly sensitive to changes in image resolution. In order to effectively address these challenges, the present invention adopts a window partitioning strategy similar to Swin-Transformer. Specifically, the original feature map is divided into multiple local windows of k×k dimensions, where k is the size of the local window (that is, the length and width of the local window are both k), and the feature comparison between the two branches is limited to be performed only within their respective corresponding windows, thereby generating two similarity graphs. During the comparison process, the features of the current branch are regarded as queries (Query), and the features of the other branch are used as keys (Key). The specific calculation method of these two features is as follows:

[0045]

[0046] In the formula, A represents the vehicle detection feature vector, B represents the region segmentation feature vector, and M att Indicates the similarity between vehicle detection features and regional segmentation features. represents the operation of calculating the mean across the channel dimension, Indicates a window reversal operation, represents the window partitioning operation, d represents the dimension of the matrix, T represents the matrix transpose calculation, and Ones_like(·) represents initializing a matrix vector whose element values ​​are all 1 and whose number of channels, width, and height are the same as A or B.

[0047] 2. Implementation of highlight area segmentation branch

[0048] The feature similarity perception module has the ability to efficiently identify similar areas between two branches without introducing additional learning parameters, showing excellent performance. However, in practical applications, in addition to the lighting provided by the vehicle's headlights and taillights, other light sources such as street lights may interfere with the accurate identification of the vehicle's corresponding highlight areas. To address this problem, the present invention proposes a highlight area segmentation branch that can deeply analyze the channel and spatial distribution of features, and mine the existence of highlight features by acquiring the attributes of different channels and combining cross-scale analysis.

[0049] In the specific implementation, first, the output of the feature similarity perception module is combined with the features of the region segmentation branch through a pixel-by-pixel multiplication operation; then, the number of channels of the combined features is halved through a convolution layer with a kernel size of 1; then, the weight distribution of the features with the number of channels halved in different channels is adjusted through a set of convolution layers with a kernel size of 3 and a convolution layer with a kernel size of 1; then, the weight distribution of the features at different positions after the weight distribution of the features with the number of channels halved in different channels is adjusted using convolution layers with different kernel sizes; it can adapt to the perception of the distribution of highlight information at different scales; finally, a convolution layer with a kernel size of 1 is used to generate spatial distribution weights by pixel-by-pixel classification. The specific implementation method can be described as follows:

[0050]

[0051] in, represents the output of the region segmentation branch, Indicates feature enhancement from the channel direction. Indicates feature enhancement from the spatial direction. represents the spatially aware convolutional layer operation, represents the channel-aware convolutional layer operation, and σ(·) represents the sigmoid function.

[0052] The backbone features of the two branches are This realizes the adaptive adjustment of feature weights, and the features belonging to the highlighted area will be emphasized.

[0053] 3. Model training and supervision loss design

[0054] The model training loss function designed by the present invention consists of two independent parts: detection loss and and segmentation loss The parameters of the model are trained and learned separately (including the backbone network, detection and segmentation heads, feature similarity perception module, and highlight area segmentation branch, etc.), which are defined as:

[0055]

[0056] Where λ is a hyperparameter that balances the loss updates of different tasks, and They are the loss calculation items for classification and location recognition in the detection task respectively. Through the simultaneous supervised training of the detection task and the segmentation task, the mutual influence between the two tasks is achieved, thus achieving common improvement.

[0057] Example 2

[0058] like Figures 1 to 3As shown, as a further optimization of Example 1, based on Example 1, this embodiment also includes the following technical features:

[0059] Due to the adoption of the technical solution of the present invention, the following technical effects are achieved: experimental verification is carried out on the publicly available night vehicle detection dataset BBD100K-Night. The proposed model is integrated into the existing target detection model in a plug-and-play manner for training and evaluation. The initial learning rate is set to 0.01, four GPUs are used, the batch size is set to 32, the number of training rounds is uniformly set to 12 rounds, and the dimension k of the feature block window is set to 7. The loss threshold λ for both detection and segmentation tasks is fixed to 0.5. The present invention adopts COCO target detection evaluation indicators, which provide additional statistics for objects classified by size (small, medium, and large), while incorporating a wider range of threshold settings (AP, AP 0.5 ,AP 0.75 ,APS,APM,APL). From the comparison results in Table 1, we can see that our proposed scheme surpasses all previous methods and obtains the best performance, which proves the effectiveness of the proposed scheme in solving the nighttime vehicle detection task.

[0060] Table 1 Comparison of experimental results of BBD100K-Night dataset

[0061]

[0062] As described above, the present invention can be preferably implemented.

[0063] All features disclosed in all embodiments in this specification, or steps in all methods or processes implicitly disclosed, except for mutually exclusive features and / or steps, can be combined and / or expanded or replaced in any manner.

[0064] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. According to the technical essence of the present invention, within the spirit and principles of the present invention, any simple modification, equivalent replacement and improvement made to the above embodiment still falls within the protection scope of the technical solution of the present invention.

Claims

1. A nighttime vehicle detection system, characterized in that: It includes a vehicle detection branch for receiving an input image and outputting a detection result, and a region segmentation branch for receiving an input image and outputting a segmentation result. The vehicle detection branch includes an i-level cascaded vehicle detection feature collector and a detection head communicatively connected to the last-level vehicle detection feature collector. The region segmentation branch includes an i-level cascaded region segmentation feature collector and a segmentation head communicatively connected to the last-level region segmentation feature collector. The detection head is communicatively connected to the segmentation head.

2. A nighttime vehicle detection system according to claim 1, characterized in that: It includes a feature similarity perception module which is respectively communicated with the vehicle detection branch and the region segmentation branch. The feature similarity perception module is used to compare the similarity between the feature of the vehicle detection branch and the feature of the region segmentation branch.

3. A nighttime vehicle detection system according to claim 2, characterized in that: When the vehicle detection branch is working, the input image is divided into multiple local windows of k×k dimensions, and the feature comparison between the vehicle detection branch and the region segmentation branch is limited to be performed only within their respective corresponding windows; during the comparison process, the features of the current branch are regarded as queries, and the features of the other branch are regarded as keys; where k is the length or width of the local window.

4. A nighttime vehicle detection system according to claim 3, characterized in that: The calculation formula of the feature similarity perception module is: or, In the formula, A represents the vehicle detection feature vector, B represents the region segmentation feature vector, and M att Indicates the similarity between vehicle detection features and regional segmentation features. represents the operation of calculating the mean across channel dimensions, Indicates a window reversal operation, represents the window partitioning operation, d represents the dimension of the matrix, T represents the matrix transpose calculation, and Ones_like(·) represents initializing a matrix vector whose element values ​​are all 1 and whose number of channels, width, and height are the same as A or B.

5. A nighttime vehicle detection system according to claim 4, characterized in that: When the region segmentation branch is working, first, the output of the feature similarity perception module is combined with the features of the region segmentation branch through a pixel-by-pixel multiplication operation; then, the number of channels of the combined features is halved through a convolutional layer with a kernel size of 1; then, the weight distribution of the features with the number of channels halved in different channels is adjusted through a set of convolutional layers with a kernel size of 3 and a convolutional layer with a kernel size of 1; then, convolutional layers with different kernel sizes are used to adjust the weight distribution of features at different positions after the weight distribution of the features with the number of channels halved in different channels is adjusted; finally, a convolutional layer with a kernel size of 1 is used to generate spatial distribution weights by pixel-by-pixel classification.

6. A nighttime vehicle detection system according to claim 4, characterized in that: The calculation formula of the region segmentation branch is: In the formula, in, represents the output of the region segmentation branch, Indicates feature enhancement from the channel direction. Indicates feature enhancement from the spatial direction. represents the spatially aware convolution operation, represents the channel-aware convolution operation, and σ(·) represents the sigmoid function.

7. A nighttime vehicle detection system according to any one of claims 1 to 6, characterized in that: The nighttime vehicle detection system can train the detection head and the segmentation head. During the training, the vehicle detection feature vector and the region segmentation feature vector are updated by the following formulas: Among them, A' represents the updated vehicle detection feature vector, and B' represents the updated region segmentation feature vector.

8. A nighttime vehicle detection system according to claim 7, characterized in that: The loss function for training Including detection loss and segmentation loss The calculation formula is: Here, λ represents a hyperparameter.

9. A nighttime vehicle detection system according to claim 8, characterized in that: The calculation formula is: in, Represents the loss calculation item of classification recognition in the detection task, Represents the loss calculation term for location recognition in the detection task.

10. A nighttime vehicle detection method, characterized in that: A nighttime vehicle detection system as claimed in any one of claims 1 to 9 is used to perform nighttime vehicle detection.