Day and night scene monocular 3D target detection method and system
By introducing brightness monitoring, low-brightness image enhancement and field adaptation modules into the monocular 3D object detection method, the poor performance of monocular 3D object detection under low light conditions is solved, and high-quality image enhancement and high-precision object detection are achieved.
Patent Information
- Application Number
- CN202510290070.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The existing monocular 3D object detection method performs poorly in low-light conditions, mainly due to changes in light, data set limitations, noise increase and depth information loss.
A method of monocular 3D object detection in day and night scenes is proposed, including backbone network, brightness monitoring module, low-brightness image enhancement module and domain adaptation module. Through adversarial learning and domain adaptation training, effective transfer learning across data sets is achieved, and detection accuracy is improved through image enhancement under low brightness conditions.
Significantly improve image quality and detail retention capabilities under low light conditions, reduce noise and distortion, achieve high-precision object detection, and the overall frame structure is highly portable and easy to transplant and deploy.
Smart Images

Figure CN120220131A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image processing, and particularly to a monocular 3D object detection method and system for day and night scenes. Background Art
[0002] In autonomous driving and intelligent transportation systems, accurate 3D object detection is crucial for ensuring driving safety. Existing 3D object detection methods mainly rely on multi-sensor fusion (such as lidar and cameras), but in practical applications, single-camera solutions have attracted much attention due to advantages such as low cost and easy deployment.
[0003] However, existing monocular 3D object detection methods perform poorly under low-light conditions, mainly for the following reasons:
[0004] 1. Lighting variation: The lighting conditions during the day and at night are significantly different, making it difficult for the model to adapt to different lighting environments.
[0005] 2. Dataset limitation: Most publicly available datasets (such as KITTI) mainly contain daytime scenes and lack sufficient nighttime scene samples, resulting in insufficient model training.
[0006] 3. Increased noise: Under low-light conditions, image noise increases significantly, affecting the accuracy of object detection.
[0007] 4. Loss of depth information: A monocular camera cannot directly obtain depth information and needs to estimate it through complex algorithms, increasing the error.
[0008] Although some solutions have been proposed in existing research, such as using generative adversarial networks (GANs) to generate nighttime scene images or adopting augmented reality technology to improve image quality, these methods still have the following problems:
[0009] Instability: Generative adversarial models (such as cyclegan) are prone to performance fluctuations during long-term training, resulting in overfitting.
[0010] High complexity: Although conditional diffusion models generate excellent image quality, they require a large amount of computing resources and time for training, and have a complex structure, making it difficult to transplant and deploy into monocular 3D object detection models. Summary of the Invention
[0011] In view of this, the purpose of the present invention is to propose a monocular 3D object detection method and system for day and night scenes to solve the problems of instability and high complexity of existing methods.
[0012] Based on the above purpose, the present invention provides a monocular 3D object detection method for day and night scenes, including the following steps:
[0013] Build a monocular 3D object detection model, where the monocular 3D object detection model includes a backbone network, a brightness monitoring module, a low-brightness image enhancement module, and a domain adaptation module. The domain adaptation module includes a gradient reversal layer, a convolutional layer, and a domain classification layer.
[0014] Input the paired daytime and nighttime image samples into the backbone network of the monocular 3D object detection model in sequence for feature extraction.
[0015] Learn the feature maps from different levels through adversarial learning methods, including passing through the gradient reversal layer, automatically reversing the gradient direction during backpropagation, and performing domain adaptation training on the feature maps through the convolutional layer and the domain classification layer to achieve effective transfer learning of cross-dataset distributions, and obtain the trained monocular 3D object detection model.
[0016] Input the image to be detected into the trained monocular 3D object detection model. After being monitored by the brightness monitoring module, if the brightness is lower than the set threshold, the low-brightness image enhancement module is used to enhance the brightness.
[0017] Input the original image data higher than the set threshold or the image data after brightness enhancement processing into the trained monocular 3D object detection model to obtain the object detection result.
[0018] Preferably, the low-brightness image enhancement module includes a generator module and a discriminator module. The generator includes an encoding layer, a transformation layer, and a decoding layer. The discriminator is used to determine whether the image is a real image or a generated image.
[0019] Preferably, the training process of the low-brightness image enhancement module includes:
[0020] Data preprocessing, select several images that meet the preset conditions as samples.
[0021] Random cropping: Randomly crop the image samples to a size of 512x512.
[0022] Independent pre-training: Adopt the method of independent pre-training to train the low-brightness image enhancement module to avoid overfitting.
[0023] Real-time supervision of training quality: Evaluate the generated images output in each epoch during the training process using the no-reference image quality assessment method TRES and the full-reference image quality assessment method FID.
[0024] Preferably, the brightness monitoring module determines whether brightness enhancement is required by calculating the average brightness of the RGB image data.
[0025] Preferably, among the paired daytime and nighttime image samples, the daytime images are from the KITTI dataset, and the nighttime images are generated from the images in the KITTI dataset through a low-brightness image generation network.
[0026] Preferably, the process of adversarial learning includes:
[0027] Define the data from the KITTI dataset as the source domain, denoted as S, and define the enhanced nighttime data as the target domain, denoted as T. Define u s ∈S, u t ∈T;
[0028] Using the H divergence, let h: X → {0, 1} be a binary classifier, where the sample u in the source domain s is labeled 0, and the sample u in the target domain t is labeled 1. The H divergence is used to represent the distance between the two domains, and the formula is as follows:
[0029]
[0030] where m is the number of samples for which the classifier h outputs the corresponding class, and I is the indicator function. When the condition inside the square brackets holds, the value of I is 1, and when the condition does not hold, the value of I is 0;
[0031] To align the source domain and the target domain and minimize the domain distance of the network, that is:
[0032]
[0033] where,
[0034] E S and E T represent the prediction errors on the source domain samples and the target domain samples respectively.
[0035] Preferably, the low-brightness image generation network includes an SD decoder, an SD encoder, a U-Net, and a CLIP text encoder;
[0036] The CLIP text encoder is used to convert the input text information into a vector representation to guide the image generation process;
[0037] The U-Net connects the SD decoder and the SD encoder and is used to generate images.
[0038] The present invention also provides a monocular 3D object detection system for day and night scenes, which is characterized in that it includes a monocular 3D object detection model. The monocular 3D object detection model includes a backbone network, a brightness monitoring module, a low-brightness image enhancement module, and a domain adaptation module. The domain adaptation module includes a gradient reversal layer, a convolutional layer, and a domain classification layer.
[0039] The backbone network is used to extract features from paired daytime and nighttime image samples;
[0040] The model learns through adversarial learning methods from feature maps at different levels, including passing through a gradient reversal layer that automatically reverses the gradient direction during backpropagation, and the feature maps are subjected to domain adaptation training through convolutional layers and domain classification layers to achieve effective transfer learning of cross-dataset distributions;
[0041] The brightness monitoring module is used to monitor the brightness of the image. If the brightness is lower than the set threshold, brightness enhancement is performed through the low-brightness image enhancement module;
[0042] The model is also used to perform monocular 3D object detection on the original image data above the set threshold or the image data after brightness enhancement processing to obtain the object detection result.
[0043] Advantages of the present invention:
[0044] High-quality generation: The low-brightness image generation model can perform high-quality style transfer between daytime and nighttime scenes, generating images close to real nighttime driving scenes. Multiple image quality evaluation metrics prove the core technical breakthroughs of the generation algorithm in multi-scale feature preservation and high-frequency detail enhancement.
[0045] Efficient enhancement: This method of the low-brightness image enhancement network performs excellently in multiple no-reference and full-reference evaluations. The highest MANIQA metric is improved by up to 36.26% compared to the second method, and other metrics are also improved to varying degrees. Therefore, it can significantly improve the image quality and detail retention ability under low-light conditions, reducing noise and distortion.
[0046] High-precision detection: The monocular 3D object detection model combining domain adaptation and key point detection can achieve high-precision object detection under different lighting conditions.
[0047] Easy portability: The overall framework structure is highly portable, easy to transplant and deploy, and applicable to a variety of application scenarios. Description of the Drawings
[0048] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0049] Figure 1 For the single-view Figure 3 D object detection method framework diagram of the embodiment of the present invention;
[0050] Figure 2 Schematic diagram of the low-brightness image generation framework according to an embodiment of the present invention;
[0051] Figure 3 Schematic diagram of the domain adaptation framework according to an embodiment of the present invention;
[0052] Figure 4 Schematic diagram of the low-brightness image enhancement network framework according to an embodiment of the present invention;
[0053] Figure 5 Display diagram of the image generated by the low-brightness image generation network according to an embodiment of the present invention;
[0054] Figure 6 Comparison diagram of the BDD100K dataset and the synthetic dataset according to an embodiment of the present invention;
[0055] Figure 7 Visualization diagram of the image enhancement result according to an embodiment of the present invention;
[0056] Figure 8 Visualization diagram of the daytime detection result according to an embodiment of the present invention;
[0057] Figure 9 Visualization diagram of the nighttime detection result according to an embodiment of the present invention. Detailed implementation manners
[0058] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific embodiments.
[0059] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those with ordinary skills in the field to which the present invention belongs. The "first", "second", and similar terms used in the present invention do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0060] As Figure 1 shown, the embodiments of this specification provide a monocular 3D object detection method for day and night scenes, including the following steps:
[0061] Build a monocular 3D object detection model, which includes a backbone network, a brightness monitoring module, a low-light image enhancement module, and a domain adaptation module. The domain adaptation module includes a gradient reversal layer, a convolutional layer, and a domain classification layer.
[0062] Input the paired day and night image samples into the backbone network of the monocular 3D object detection model in sequence for feature extraction.
[0063] Learn the feature maps from different levels through adversarial learning methods, including passing through the gradient reversal layer, automatically reversing the gradient direction during backpropagation, and performing domain adaptation training on the feature maps through the convolutional layer and the domain classification layer to achieve effective transfer learning of cross-dataset distributions, and obtain the trained monocular 3D object detection model.
[0064] Input the image to be detected into the trained monocular 3D object detection model. After being monitored by the brightness monitoring module, if the brightness is lower than the set threshold, the low-light image enhancement module is used to enhance the brightness.
[0065] Input the original image data higher than the set threshold or the image data after brightness enhancement processing into the trained monocular 3D object detection model to obtain the object detection result.
[0066] The pre-training process of the low-light image enhancement network consists of two major modules: a generator and a discriminator. The generator includes three main parts: an encoding layer, a transformation layer, and a decoding layer. The specific framework is as Figure 4 shown:
[0067] The encoding layer consists of three convolutional layers and an instance normalization layer; the transformation layer contains nine residual blocks; the decoding layer combines transposed convolution and ordinary convolutional layers. The discriminator adopts a 70×70 PatchGAN architecture to determine whether an image is a real image or a generated image by classifying overlapping image patches (70×70).
[0068] Data preprocessing: Select 10,000 samples with normal exposure, proper focus, and no blur from the BDD100K dataset for training.
[0069] Random cropping: The input image is randomly cropped to a size of 512x512 during training to better handle fine details.
[0070] Independent pre-training: During multiple training processes, it is observed that the training results of the enhanced model are relatively unstable. Especially compared with the object detection model, the training loss and the image generation quality fluctuate greatly. Therefore, the method of independent pre-training is adopted to improve the stability of the model and avoid overfitting.
[0071] Real-time supervision of training quality: To further accurately understand the quality of the generated images, the generated images output in each epoch during the training process are evaluated using the no-reference image quality assessment method TRES and the full-reference image quality assessment method FID.
[0072] The monocular 3D object detection model of the embodiments of the present invention combines the characteristics of the MonoFlex network design. Objects are recognized through their representative points and predicted through the peaks in the heat map. The specific framework is as Figure 1 shown.
[0073] (1) Inference stage (the network within the dashed box is not effective)
[0074] The input single RGB image data (day or night image) first passes through the "brightness monitoring" module to calculate the average brightness of the image. If the brightness is lower than the set threshold, it is processed through the pre-trained brightness enhancement network. The brightness enhancement process includes three steps: encoding, transformation, and decoding. The processed enhanced data or the original data will be input into the DLA34 backbone network for feature extraction.
[0075] The feature maps of different sizes extracted by the backbone network are upsampled and fused, and then passed as input to multiple regression heads to obtain the attributes of the instance, including the 2D bounding box, object size, pose, key points, and depth. The fused depth estimation regression is based on the object size and multiple key points and is completed using the uncertainty-guided weighted average method.
[0076] (2) Training stage (the network within the dashed box is effective)
[0077] Compared with the inference stage, in the training stage, paired day and night images are used. The night image still undergoes low-brightness image enhancement, while the day image does not undergo image enhancement. Then the two images are sequentially input into the backbone network for feature extraction.
[0078] Next, the feature maps from different levels pass through the gradient reversal layer (GRL), which automatically reverses the gradient direction during the backpropagation process. These feature maps are then subjected to domain adaptation training through convolutional layers and domain classification layers, aiming to make the feature representations between the source domain and the target domain as consistent as possible, so as to achieve effective transfer learning of cross-dataset distributions.
[0079] The subsequent instance attribute regression still only regresses the attributes of the day image, which is exactly the same as the inference stage. The advantage of doing this is that it can effectively avoid potential problems such as data leakage.
[0080] Backbone Network: It adopts an enhanced DLA-34 architecture and integrates deformable convolutions. The core idea of this network is to enhance the network's ability to extract semantic and spatial information through Hierarchical Deep Aggregation and Iterative Deep Aggregation. In DLA-34, the layer-by-layer aggregation module establishes iterative connections between adjacent network stages, enabling the output of each stage to fuse with the information of the previous and subsequent layers, thus effectively avoiding information loss in the deep network. Different from traditional network structures, DLA-34 can effectively combine low-level features with high-level features through these cross-layer connections, enhancing the network's learning ability. In addition, Iterative Deep Aggregation further improves the network's information transmission efficiency. Through connections spanning multiple network stages, Iterative Deep Aggregation can deepen the network depth while refining the expression of spatial information. This iterative process enables the network to propagate features and gradients between different levels, promoting more accurate feature extraction and optimization. Under this structure, DLA-34 can not only improve semantic understanding ability but also effectively enhance the representation of spatial information, and is widely used in object detection algorithms.
[0081] Brightness Monitoring Module: It automatically determines whether brightness enhancement is required by calculating the average brightness of RGB image data. The design of the brightness monitoring module follows traditional processing methods. First, the input RGB image is converted into a grayscale image, and the conversion formula is as follows:
[0082] Y = 0.299R + 0.587G + 0.114B
[0083] where R, G, and B represent the pixel values of the red, green, and blue channels of each pixel respectively, and Y is the grayscale value corresponding to that pixel. Through this formula, the RGB values of each pixel in the color image can be converted into a single grayscale value.
[0084] Next, the average brightness value of the image is calculated. Let the height of the image be H and the width be W, then the average brightness of the image can be expressed by the following formula:
[0085]
[0086] where Y ij represents the grayscale value of the pixel located in the i-th row and j-th column. By summing up the grayscale values of all pixels in the entire image and taking the average, the overall brightness of the image can be obtained. At the same time, by statistically calculating the average brightness values of all images in the dataset, the brightness threshold for distinguishing between day and night images can be obtained.
[0087] The brightness monitoring module can evaluate the lighting conditions of an image by calculating the average brightness of the image, enabling the system to make adaptive adjustments according to different brightness conditions and ensuring the accuracy and stability of target detection.
[0088] Domain adaptation training: The purpose of the domain adaptation module is to calibrate the distribution differences in the convolutional feature maps between the source domain and the target domain. The basic assumption is that if the overall distributions of images in different domains are similar, then their distributions at the object level will also be approximately similar. In other words, the main factor causing domain differences is the change in the overall image distribution. In a deep network, the intermediate layers of the convolutional feature maps capture features such as the shape, contour, and edges of the image. To minimize the differences between domains, this embodiment proposes a hierarchical domain feature alignment module, which consists of multiple adversarial domain classifiers embedded in different convolutional blocks. Figure 3 Shows a single domain feature alignment module, defining the source domain and target domain image features obtained from the backbone network as: F S ,F T , after passing through the Gradient Reversal Layer (GRL), channel concatenation is performed to obtain F con , and then it is fed into the domain discriminant network D, and the output obtains a two-channel prediction result.
[0089] First, define the data from KITTI as the source domain, denoted as S, and the enhanced night data as the target domain, denoted as T. Define, u s ∈S, u t ∈T. To measure the distribution difference or distance between the two domains, the H-divergence is used. Let h: X → {0, 1} be a binary classifier, where the sample u in the source domain s is labeled as 0, and the sample u in the target domain t is labeled as 1. When the classifier is more difficult to distinguish between the two domains, it indicates that the distribution difference between the domains is smaller. According to this definition, the H-divergence can be used to represent the distance between the two domains, and the formula is as follows:
[0090]
[0091] where m is the number of samples for which the classifier h outputs the corresponding class, and I is I
[0092] is the Indicator Function. When the condition inside the square brackets holds, the value of I is 1; when the condition does not hold, the value of I is 0.
[0093] To align these two domains, it is necessary to minimize the domain distance of the network, ensuring that the features of the source domain and the target domain are difficult to distinguish; in other words, the classification error rate of the network should be maximized as much as possible. The two terms in the formula:
[0094]
[0095] Need to reach their maximum values, so define:
[0096]
[0097] Among them, E S and E T represent the prediction errors on the source domain samples and the target domain samples respectively. Therefore, the minimum domain distance is expressed as:
[0098]
[0099] This adversarial learning method is optimized by combining adversarial training with a Gradient Reversal Layer (GRL). Specifically, the Gradient Reversal Layer reverses the sign of the gradient, making the optimization direction of the domain classifier opposite to that of the standard classification task. As a result, the network tries to maximize the error rate of the domain classifier, thus promoting the alignment of the feature distributions between the source domain and the target domain.
[0100] In actual operation, the domain classifier is integrated into the backbone network, especially in the image classification task in the unsupervised domain adaptation scenario, after the feature output. Through this adversarial training, the network can better learn domain-invariant features, namely geometric structure features, thus minimizing the distribution difference between the source domain and the target domain and improving the performance of the model in the target domain. It should be noted that the domain classification component is only effective during the training phase. During the testing or inference phase, the domain classifier does not participate in the calculation, and its main role is to help the model achieve better domain adaptation during training.
[0101] Monocular Figure 3 The task of 3D object detection is to identify objects of interest from a two-dimensional image and predict their corresponding three-dimensional attributes, including 3D position (x, y, z), size (h, w, l), and orientation θ. To simplify the prediction process, the 3D position is usually converted into 2.5D information (u c , v c , z) for prediction, where u c and v c represent the center coordinates of the object on the image plane, and z represents the depth of the object in the camera coordinate system. The recovery process of the 3D positions x and y can be expressed as:
[0102]
[0103] Among them, c x and c y are the optical center coordinates of the camera respectively, and f x and f y are the focal lengths of the camera. By predicting u c , v cGiven x, y, and z, the position (x, y) of the object in 3D space can be restored, thus realizing the position estimation in monocular Figure 3 D object detection.
[0104] It should be noted that the loss of depth information during the imaging process is an important factor limiting the performance of monocular Figure 3 D detectors. To improve the accuracy of depth estimation, this object detection method combines the method in MonoFlex. Specifically, the final depth estimate is calculated by weighted averaging all candidate depth values, where the weight of each candidate depth value corresponds to its associated uncertainty.
[0105] This method means that when calculating the final depth estimate, the study not only considers the likelihood of different candidate depth values but also the uncertainty of each estimate. In this way, the result can effectively reduce the error caused by the loss of depth information of individual objects, thereby improving the overall performance of the 3D object detector. The final depth estimate Z soft is calculated by the formula:
[0106]
[0107] where, z i is the i-th candidate depth value, and w i is the uncertainty weight associated with the i-th depth value z i satisfying
[0108] Therefore, combined with the above-mentioned adaptive domain training, the total loss in the training stage can be expressed as:
[0109] Loss = Loss 2D + Loss 3D + Loss soft_combined_depth + Loss domain_classifier
[0110] where, the Loss 2D term includes instance classification, 2D size, and heatmap, and these attributes mainly involve the 2D visual features of the image. The Loss 3D term involves the size (h, w, l), orientation, and key points of the object.
[0111] To generate better low-brightness image samples for training the model, the embodiments of this specification also provide a low-brightness image generation network.
[0112] Architectural design: The framework introduced in this chapter aims to generate the output of the target domain (night image) based on the text and the input of the source domain (daytime image), realizing the style transfer task, and keeping the structure of the input image unchanged during this process. Specifically, the low-brightness image generation framework is asFigure 2 As shown, the framework design is based on the classic method of StableDiffusion Turbo (SD-Turbo). The overall model is an end-to-end model, mainly composed of four core components: the SD encoder, the SD decoder, the U-Net, and the CLIP text encoder.
[0113] Text Guidance: The CLIP text encoder is a model based on the Transformer architecture, aiming to convert the input text information (such as "night driving") into a vector representation. During training, CLIP randomly selects an image and its corresponding label text from the training set. The task is to extract the embedding vectors of the label text and the image through the text encoder and the image encoder respectively. Then, CLIP uses cosine similarity to compare the similarity of the two embedding vectors, thereby determining whether the text and the image match. Through gradient backpropagation, CLIP continuously optimizes the model to improve the matching ability. In generative models such as Stable Diffusion, the CLIP text encoder converts text into latent vectors to guide the image generation process.
[0114] Skip Connection Mechanism: Aiming at the visual detail degradation problem in the Stable Diffusion encoding and decoding architecture and to retain the input structure and high-frequency details, the present invention proposes to add a skip connection mechanism strengthened by zero convolution between the SD encoder and the SD decoder. Its optimization effects are reflected in the following three aspects:
[0115] (1) Dynamic Feature Calibration
[0116] Through the learnable characteristics of the zero convolution layer, the system can adaptively adjust the fusion weights of features at different levels, and establish an optimal matching relationship between the encoder features and the decoder features. This dynamic adjustment mechanism effectively avoids the problems of feature redundancy or insufficient information caused by traditional fixed-weight skip connections.
[0117] (2) High-Frequency Feature Preservation
[0118] The zero convolution module optimizes the sensitivity to high-frequency components through gradient update, and significantly reduces the loss rate of high-frequency visual elements (such as structural edges, texture details) during the image compression and reconstruction process.
[0119] (3) Multi-Scale Fusion Enhancement
[0120] Construct a hierarchical feature pyramid fusion network, and achieve lossless integration of cross-resolution features through upsampling operations regulated by zero convolution. This structure not only retains the spatial topological characteristics of the original image, but also enhances the decoder's ability to recover local details, especially showing stronger reconstruction robustness in regions with sudden illumination changes (such as building outlines, vegetation textures).
[0121] Figure 5 It shows the simulated night driving training data generated by using the KITTI dataset as input after fine-tuning training. It can be seen that, without destroying the original geometric structure of the image input, this method has successfully generated highly realistic night scenes. The generated training data not only closely resembles the visual effect of real nights but also maintains high resolution and low noise, ensuring the quality and accuracy of the training data and providing a reliable basis for the subsequent training of object detection models.
[0122] To verify the effectiveness of the above method, the following experiments were conducted.
[0123] 1. Low-brightness Image Generation Network
[0124] Following the general method of large model fine-tuning, the designed framework introduced LoRA adapter fine-tuning into each module of the generation network, including the encoder, U-Net, and decoder, and specifically retrained the first layer of the U-Net. The model was trained on the BDD100K dataset and has 330 million trainable parameters. During training and inference, the model used the Adam optimizer with a learning rate of 1×10 -6 , and a batch size of 8.
[0125] 2. Image Enhancement Network
[0126] Different from the image generation model, the quality of the training data has a significant impact on the effectiveness of the enhancement model. Therefore, 10,000 samples with normal exposure, correct focus, and no blur were selected from the BDD100K dataset for training the model. 7000 samples were randomly selected as training data in each epoch, and the model was trained for a total of 100 epochs. It should be noted that the resolution has an important impact on the model performance; there are significant differences in resolution among the training samples in the BDD100K and KITTI datasets. Therefore, the KITTI dataset was not used for training the image generation model, which can also maintain the independence of the enhancement module and demonstrate its generality. During the model training process, the input images were randomly cropped to a size of 512x512. Compared with the generation network, the image enhancement network is more simplified. Although training while maintaining the aspect ratio allows the model to capture more global cues, it performs poorly in dealing with details. For example, distant cars or pedestrians may appear with blurred outlines and increased noise, which has a negative impact on subsequent object detection. Therefore, the training method of random cropping better improves the model's ability to handle details. It is worth mentioning that the model performed best at the 20th epoch, indicating that the image enhancement network can achieve high efficiency within a relatively short training period. During the training process of the object detection model, the pre-trained image enhancement model no longer participates in the loss calculation, and its weights are no longer updated.
[0127] 3. Monocular 3D Object Detection Network
[0128] KITTI dataset and synthetic night dataset detection: The backbone network adopts an enhanced DLA-34 architecture and incorporates deformable convolutions. The input image is padded to a resolution of 384x1280 before processing. The AdamW optimizer is used during training, with an initial learning rate of 3×10 -4 , and a weight decay of 1×10 -5 . A total of 150 epochs are trained, with a batch size of 8. Each batch contains 4 pairs of matching data; for every daytime training image loaded, the corresponding nighttime image is loaded to train the domain adaptation network. It should be noted that paired training data is only required during the training phase, and the system can perform adaptive inference on any daytime or nighttime image during the validation and testing phases. At the 80th and 120th epochs, the learning rate is reduced to one-tenth of the original for more refined adjustment. Data augmentation is limited to horizontal flipping to enhance the model's robustness under different perspectives.
[0129] 4. Evaluation and Optimization:
[0130] Table 1 lists 4 no-reference image quality assessment methods (MUSIQ, MANIQA, TRES, NIQE), and the quality scores are calculated respectively in the nighttime images of the BDD100K dataset and the simulated nighttime driving images generated with low brightness. The upward arrow in the list represents that the higher the score, the better the image quality, and the downward arrow represents that the lower the score, the better the image quality.
[0131] Table 1
[0132]
[0133] From the data in the table, it can be seen that the nighttime images in the BDD100K dataset generally have low scores in multiple no-reference image quality assessment methods. Especially in the two indicators of MUSIQ (Multi-Scale Image Quality Assessment) and TRES (Texture Edge Strength), they only obtained 29.219 and 28.048 points respectively, while the low-brightness generated images achieved a leapfrog improvement with 40.165 and 37.239 points. The NIQE (Natural Image Quality Evaluation) score dropped from 6.874 to 4.035 (a decrease of 41.3%), indicating that the statistical characteristics of the generated images are closer to the natural imaging law, effectively avoiding the gamma distortion and color cast accumulation problems caused by dynamic range compression in traditional nighttime images. The above data prove the core technical breakthrough of the generation algorithm in multi-scale feature preservation and high-frequency detail enhancement.
[0134] On the other hand, the main reason for the low score of the dataset is that some of the image data collected at night is overexposed, blurred, or even both, which affects the quality of the images to a certain extent. Figure 6 The top two images show the data in the dataset with normal exposure and no blur; the middle two images show the blurred and overexposed data respectively. Overexposure causes the loss of details in some areas of the image, while blur causes a decrease in the clarity of the image. These factors combined result in a low score for the image quality assessment of the dataset. The low-light generated images ( Figure 6 the bottom two images) can better overcome these problems due to their generation method, so they score higher in most quality assessment methods and show better visual quality.
[0135] The evaluation of low-light image enhancement is as follows:
[0136] Table 2 shows the comparison between the daytime images in the KITTI dataset and several low-light image enhancement methods using the no-reference image quality assessment method (the best results are marked in red, and the second best are marked in blue). It can be seen that the low-light image enhancement method of the present invention performs the best in 3 no-reference image quality assessment criteria (MUSIQ, MANIQA, TRES), with improvements of 19.39%, 36.26%, and 21.12% compared to the second method respectively.
[0137] Table 2
[0138]
[0139] Table 3 shows the comparison between the output of the low-light image enhancement method and the KITTI dataset using 3 full-reference image quality assessment methods for the full-reference image quality assessment of low-light generated images (the best results are marked in red, and the second best are marked in blue). This method performs the best in two tests (PSNR, FID), with improvements of 5.79% and 26.24% respectively.
[0140] Table 3
[0141]
[0142] The reason why this method can achieve a high score in the image quality assessment is mainly due to the adoption of the strategy of real-time monitoring of image quality assessment performance during the training process, and carefully selecting night images with normal exposure and high quality as training data. In this way, the model can learn the changes in light, brightness, etc. between day and night more accurately, and thus perform better in the assessment. The advantages of this method can be more clearly observed through the visualization in Figure 7.
[0143] Comparison of daytime scene detection performance:
[0144] Figure 8 The visualization shows the detection results of this method on the KITTI test set. It can be seen that the model can effectively detect targets of the car category and distinguish targets of the van category. This is also the advantage of the visual algorithm over the point cloud algorithm. For distant targets, the model's depth prediction is also very accurate.
[0145] Table 4 shows a comprehensive evaluation of the 3D detection performance of this method on the KITTI validation dataset and compares it with the current mainstream methods. 3D ) as the core evaluation indicator, and introduce inference time and additional data requirements as the measure of efficiency and practicality. 3D , this manual calculates the average precision at a recall rate of 40, and sets the intersection over union (IoU) threshold to 0.7. Experimental results show that this method has significant advantages in balancing performance and efficiency, especially without relying on additional training data, and can still achieve detection accuracy comparable to that of multimodal methods.
[0146] Table 4
[0147]
[0148] In terms of detection performance, this method achieved AP of 24.34% and 17.98% in the Easy and Mod difficulty of the KITTI3D detection task, respectively. 3D The average accuracy surpasses all the comparison methods. In particular, compared with the LiDAR-based MonoRUN method (20.02% Easy, 14.65% Mod), this method improves by 4.32% at Easy difficulty and 3.33% at Mod difficulty. In addition, the reasoning time of this method is 53ms, including 37ms of detection and 16ms of enhancement. Compared with the reasoning time of MonoRUN assisted by LiDAR of 70ms, this method reduces 17ms, showing higher computational efficiency.
[0149] In comparison with other monocular methods, this method also performs excellently. Compared with the SMOKE method (14.76% Easy, 12.85% Mod), this method has improved by 9.58% and 5.13% respectively under Easy and Mod difficulties. Especially under the more challenging Mod difficulty, the improvement of this method is significant, verifying the robustness of the method to complex scenarios. This method is improved based on the MonoFlex framework. Benefiting from the augmentation of input data and the domain adaptation network, which increases the model's perception of structural invariance, this method also has an improvement compared with MonoFlex (0.70% Easy, 0.47% Mod, 0.18% Hard).
[0150] Regarding the inference efficiency, this method fully optimizes the use of computing resources through lightweight design. Although the inference time increases after introducing the enhancement module, compared with other methods, it still maintains a low latency. For example, compared with CaDDN (630ms), the inference time of this method is significantly shorter and can meet the requirements of high-real-time application scenarios such as autonomous driving.
[0151] Table 5
[0152]
[0153] In the APBEV@IoU = 0.7 evaluation of the KITTI dataset in Table 5, this method is significantly better than the comparison methods such as M3D-RPN, MonoPair, and MonoDLE under Easy, Mod, and Hard difficulties. Especially under Easy and Mod difficulties, the improvement amplitude is more obvious. Specifically, the APs of this method under Easy, Mod, and Hard difficulties BEV are 32.83%, 24.50%, and 20.92% respectively, showing a large improvement compared with other methods. The second-place MonoDLE has improvements of 7.86%, 5.17%, and 3.91% respectively under Easy, Mod, and Hard difficulties, demonstrating its robustness and high-precision performance in complex scenarios. Generally speaking, this method shows strong detection capabilities in various difficulty scenarios and is better than existing methods.
[0154] Comparison of night scene detection performance:
[0155] This specification evaluates the detection performance of this method on the synthesized low-light KITTI validation set and compares it with the performance of existing advanced methods under low-light conditions. As shown in Table 6, the average precision AP of 3D detection for the vehicle category is used this time 3DAs an evaluation metric, the Intersection over Union (IoU) threshold remains at 0.7, and the detection results are classified into three categories: Easy, Moderate, and Hard according to the difficulty level. The experimental results show that the proposed method exhibits significant advantages in low-light scenarios, especially with more prominent performance improvements in the more challenging Moderate and Hard difficulties.
[0156] Table 6
[0157]
[0158] Specifically, the proposed method achieves an Average Precision (AP) of 22.70 in the Easy difficulty, 3D surpassing the existing best method AMAE3D by 0.67%. In the Moderate and Hard difficulties, it reaches 16.60% and 13.79% respectively, with improvements of 5.73% and 7.73% compared to AMAE3D. Compared with the DEVIANT method, the proposed method achieves performance gains of 1.34%, 9.50%, and 11.84% in the three difficulty levels respectively.
[0159] To further illustrate the advantages of the method, Figure 9 comparative examples between the proposed method and the current best method CubeR-CNN are shown. The experiments find that CubeR-CNN is prone to missing distant targets under low-light conditions, while the proposed method significantly improves the detection accuracy of distant targets by fusing adaptive light enhancement and depth perception features. These results fully demonstrate that the technical solution proposed in this specification can effectively alleviate the problems of texture loss and depth estimation bias caused by low light, thus providing a more reliable solution for night-time environmental perception in autonomous driving.
[0160] Ablation experiments
[0161] Table 7
[0162]
[0163]
[0164] In the ablation experiment section, to further verify the effectiveness of the hierarchical domain adaptation module, Table 7 compares the performance of this method with two other models on the low-luminance generation dataset: one is the model without the domain adaptation module, and the other only uses a single domain adaptation module, where the input data is the feature map after feature fusion. Consistent with the previous evaluation, the AP value is calculated by setting the IoU threshold to 0.7 at 40 recall positions. The results in Table 3 show that this method performs best on the validation set. Domain adaptation training significantly improves the performance of the model. Compared with the model without the domain adaptation module, the performance at the easy, medium, and difficult levels on the validation set is improved by 12.60%, 11.63%, and 9.10% respectively.
[0165] In summary, while ensuring high precision, this method also improves the computational efficiency by optimizing the network structure, especially outperforming other similar methods in complex scenarios, showing its good practicality and application potential.
[0166] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present invention is limited to these examples; under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present invention as described above, which are not provided in detail for the sake of brevity. Any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A monocular 3D target detection method for day and night scenes, characterized in that: The method comprises the following steps: A monocular 3D target detection model is established, wherein the monocular 3D target detection model includes a backbone network, a brightness monitoring module, a low brightness image enhancement module and a domain adaptation module, wherein the domain adaptation module includes a gradient reversal layer, a convolution layer and a domain classification layer. The paired daytime and nighttime image samples are sequentially input into the backbone network of the monocular 3D object detection model for feature extraction; The feature maps from different levels are learned through adversarial learning methods, including the gradient reversal layer, which automatically reverses the gradient direction during the back propagation process. The feature maps are trained for domain adaptation through convolutional layers and domain classification layers to achieve effective transfer learning across data sets, and obtain a trained monocular 3D object detection model. The image to be detected is input into the trained monocular 3D object detection model and monitored by the brightness monitoring module. If the brightness is lower than the set threshold, the low brightness image enhancement module is used to enhance the brightness. The original image data or the image data after brightness enhancement that is higher than the set threshold is input into the trained monocular 3D target detection model to obtain the target detection result.
2. The monocular 3D target detection method for day and night scenes according to claim 1, characterized in that: The low-brightness image enhancement module includes a generator module and a discriminator module, the generator includes a coding layer, a transformation layer and a decoding layer, and the discriminator is used to determine whether the image is a real image or a generated image.
3. The monocular 3D target detection method for day and night scenes according to claim 2, characterized in that: The training process of the low brightness image enhancement module includes: Data preprocessing, selecting several images that meet the preset conditions as samples; Random cropping: randomly crop the image samples to 512x512 size; Independent pre-training: Use independent pre-training to train the low-brightness image enhancement module to avoid overfitting; Real-time supervision of training quality: The generated images output at each epoch during training are evaluated using the no-reference image quality assessment method TRES and the full-reference image quality assessment method FID.
4. The monocular 3D target detection method for day and night scenes according to claim 1, characterized in that: The brightness monitoring module determines whether brightness enhancement is required by calculating the average brightness of the RGB image data.
5. The monocular 3D target detection method for day and night scenes according to claim 1, characterized in that: In the paired daytime and nighttime image samples, the daytime images come from the KITTI dataset, and the nighttime images are generated by using the images in the KITTI dataset through a low-brightness image generation network.
6. The monocular 3D target detection method for day and night scenes according to claim 5, characterized in that: The adversarial learning process includes: The data from the KITTI dataset is defined as the source domain, denoted as S, and the enhanced nighttime data is defined as the target domain, denoted as T. Define, u s ∈S,u t ∈T; Using H divergence, let h:X→{0,1} be a binary classifier, where the sample u in the source domain s Labeled as 0, sample u in the target domain t Marked as 1, H divergence is used to represent the distance between two domains, and the formula is as follows: Where m is the number of samples of the corresponding category output by classifier h, and I is the indicator function. When the condition in the square brackets is met, the value of I is 1, and when the condition is not met, the value of I is 0; In order to align the source domain and the target domain, the domain distance of the network is minimized, that is: in, E S and E T They represent the prediction errors on source domain samples and target domain samples respectively.
7. The monocular 3D target detection method for day and night scenes according to claim 5, characterized in that: The low brightness image generation network includes an SD decoder, an SD encoder, a U-Net and a CLIP text encoder; The CLIP text encoder is used to convert the input text information into a vector representation to guide the image generation process; The U-Net connects the SD decoder and the SD encoder to generate an image.
8. A monocular 3D target detection system for day and night scenes, characterized in that: It includes a monocular 3D target detection model, which includes a backbone network, a brightness monitoring module, a low brightness image enhancement module and a domain adaptation module, and the domain adaptation module includes a gradient reversal layer, a convolution layer and a domain classification layer. The backbone network is used to extract features from paired daytime and nighttime image samples; The model learns from feature maps at different levels through an adversarial learning method, including a gradient reversal layer that automatically reverses the gradient direction during back propagation, and the feature map undergoes domain adaptation training through a convolutional layer and a domain classification layer to achieve effective transfer learning across data set distributions; The brightness monitoring module is used to monitor the brightness of the image, and if the brightness is lower than a set threshold, the brightness is enhanced by the low-brightness image enhancement module; The model is also used to perform a monocular 3D target detection model on original image data above a set threshold or image data after brightness enhancement processing to obtain target detection results.