Zero sample dark domain adaptive target detection method
By combining the CRF camera response function and the traditional single-graph CRF method in the dark domain environment, zero-sample dark domain adaptive object detection is achieved, solving the problem of object detection accuracy and robustness in the dark domain environment, and significantly improving the accuracy and flexibility of detection.
Patent Information
- Application Number
- CN202510218648.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
AI Technical Summary
In dark-domain environments, traditional object detection methods are difficult to accurately identify targets under insufficient lighting or complex lighting conditions, mainly due to insufficient lighting, resulting in image quality degradation and blurred target features.
Using the zero-sample dark domain adaptive object detection method, the normal brightness data is converted into low-brightness images through the principle of the CRF camera response function, and combined with the traditional single-picture CRF method, the object detection in the dark domain environment is realized. The specific steps include: inversely obtaining CRF parameters, reducing the irradiation map based on CRF parameters and enhancing it, dynamically optimizing the pre-training detection model, and designing a multi-task joint loss function.
It significantly improves the accuracy and robustness of object detection in dark domain environments, reduces the impact of light changes on object detection, and improves the flexibility and accuracy of detection.
Smart Images

Figure CN120125901A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to a zero-shot dark domain adaptive object detection method for dark domain environments, aiming to improve the accuracy and robustness of object detection under insufficient lighting or complex lighting conditions. Background Art
[0002] In object detection tasks, lighting conditions have a significant impact on detection performance. Especially in dark domain environments, due to insufficient lighting or uneven lighting, the image quality deteriorates, and object features become blurred, making it difficult for traditional object detection methods to accurately identify objects. Currently, the object detection technology in dark domain environments has the following defects:
[0003] 1) The traditional method for low brightness detection is to train the enhancement and detection networks separately and then connect them, but it cannot guarantee good compatibility between the enhanced image and the detection network. In particular, the ghost problem caused by enhancement is prone to misjudgment.
[0004] 2) The end-to-end detection model based on deep learning connects the enhancement and detection networks first and then trains the network using a special dataset. Although the design of the direct connection layer can enable the front end of the network to obtain a relatively good training effect, it is extremely difficult to obtain a dark domain object detection dataset.
[0005] 3) Most low brightness object detection datasets contain both real low brightness images and synthetic low brightness images, and synthetic low brightness images will inevitably have a negative impact on training.
[0006] Therefore, it is necessary to develop an object detection method that can actively adapt to dark domain environments. Summary of the Invention
[0007] To overcome the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a zero-shot dark domain adaptive object detection method. Based on the principle of the CRF camera response function, by synthesizing normal brightness data into low brightness images and combining the traditional method of obtaining CRF from a single image, accurate detection of objects in dark domain environments is achieved.
[0008] The technical solution of the present invention is as follows:
[0009] A zero-shot dark domain adaptive object detection method, characterized by including the following steps:
[0010] Step 1. Obtain the camera response function (CRF) parameters through the CRF inverse solution network:
[0011] Convert the object detection dataset with normal brightness into a synthetic dark image, use the CRF inverse solution network based on the RepVGG multi-branch residual structure to predict the CRF curve parameters. At the same time, conduct distillation training with the reference CRF curve generated by the traditional CRF calculation method, and use the L2 norm loss of the discrete point set to constrain the network convergence;
[0012] Step 2. Restore the irradiated image based on the CRF parameters and enhance the irradiated image:
[0013] Perform inverse CRF operation on the RGB channels of the synthetic dark image to obtain the irradiated image, generate a residual matrix through the enhancement network with a cross-layer Bottleneck structure, combine the irradiated image and the residual matrix to output the enhanced image, and introduce the SSIM loss to constrain the perceptual consistency between the enhanced image and the normal brightness image;
[0014] Step 3. Dynamic optimization based on the pre-trained detection model:
[0015] Input the enhanced image and the normal image into the object detection network, align the feature distribution differences through the difference between the two camera response functions in Step 1 and the KL divergence in Step 2, and combine the regression loss and classification loss of object detection to jointly optimize the CRF inverse solution module, the enhancement module and the detection module;
[0016] Step 4. Design of the multi-task joint loss function:
[0017] The total loss function includes the CRF reconstruction loss, the enhanced SSIM loss, the feature KL divergence loss and the detection loss. Adjust the weights of each loss through hyperparameters to achieve end-to-end collaborative training.
[0018] Furthermore, the specific steps of Step 1 are as follows:
[0019] Step 1.1 Process the object detection dataset with normal brightness to generate a synthetic dark image;
[0020] Step 1.2 For the synthetic dark image, obtain the CRF curve parameters through two paths respectively: one path predicts the CRF curve parameters through the CRF inverse solution network, and the other path obtains the reference CRF curve using the traditional CRF calculation method;
[0021] Step 1.3 Represent the CRF curve parameters output by both paths in the form of a discrete point set, and use the L2 norm to evaluate the distance between the two CRF curves as the loss function;
[0022] Step 1.4 Design the CRF network architecture, which is based on the residual network and uses the RepVGG structure internally to extract image features and predict the CRF curve parameters;
[0023] Step 1.5 Select the dimension of the output parameter according to the accuracy requirement of the CRF curve, and train the CRF network using the loss function to minimize the difference between the two CRF curves.
[0024] Further, in step 1.3, the L2 norm is used to evaluate the distance between the CRF curve obtained by the CRF inverse network and the reference CRF curve obtained by the traditional CRF method as the loss function, and the formula is as follows:
[0025] L CRF = ||CRF 参考 - CRF x || 2
[0026] Further, in step 1.2, the CRF inverse network adopts a RepVGG multi-branch residual structure, including a convolutional layer, a max pooling layer, and a ResNet18 improvement module. The image first undergoes scale reduction through the initial convolutional layer and pooling layer, and then passes through 4 ResBlocks in sequence. Each ResBlock will first reduce the length and width of the image to half of the input and then increase the data dimension. The ResBlock passes through a RepVGG unit, a non-linear function, and a RepVGG unit in sequence, and then is superimposed with the input data and undergoes non-linear processing to obtain the output. The RepVGG unit utilizes the mathematical principle of the convolutional kernel. During training, the unit has a multi-branch topological structure, and the final CRF curve output is determined by the method of discrete point sets.
[0027] Further, step 2 specifically includes:
[0028] Step 2.1 Restore the irradiance map: Based on the camera response function, perform an inverse operation on the RGB three-channel pixel matrix of the synthesized dark map using the CRF curve obtained in step 1 to obtain an irradiance map representing the three-channel light intensity distribution;
[0029] Step 2.2 Enhance the irradiance map: Design an enhancement network using the residual learning method. The enhancement network outputs a residual matrix with the same size as the input irradiance map, and the final enhanced map is obtained by adding the residual matrix to the input irradiance map.
[0030] Furthermore, the enhanced network adopts a path aggregation architecture, including a low-level feature enhancement path (the layer where P3 is located) and a high-level detection-oriented path (the layer where P1 and P2 are located). It fuses multi-level features through a cross-layer Bottleneck module and outputs a residual matrix with the same size as the irradiated image. The inputs of the path aggregation network are irradiated images X1, X2, and X3 with different resolutions obtained by successive downsampling from bottom to top. Each input has corresponding outputs with the same resolution, namely P1, P2, and P3. At the same time, there is a guiding relationship between different layer paths. X1 is aggregated with X2 after being processed by the Bottleneck and ResBlock, and there is also such a relationship between X2 and X3. In the final image output stage, P1 - P3 are aggregated with the upper-layer output with a smaller resolution through downsampling, and finally the output P3 of the entire image is obtained. The Bottleneck consists of 4 ResBlocks, and a direct connection layer is established between the outputs with the same resolution.
[0031] Furthermore, in step 3, the object detection network is based on the pre-trained YOLOv5s model, adapts to the feature extraction of the enhanced image through fine-tuning, and uses the KL divergence loss to constrain the difference in feature distributions between the enhanced image and the normal image. The specific formula is:
[0032] L fea = D KL (F 原 ||F 增 )
[0033] L det = λ cls L cls + λ reg L reg
[0034] Among them, the KL divergence is an index to measure the difference between two probability distributions. D KL uses the feature distribution extracted from the normal image as a reference to measure the difference between the feature distribution extracted from the enhanced image and the reference distribution. Among them, F 原 and F 增 are the feature distributions extracted from the normal image and the enhanced image by the detection network respectively. The L drt loss is obtained by the linear combination of the classification loss L cls and the regression loss L reg . λ cls and λ reg are the corresponding coefficients respectively.
[0035] Furthermore, in step 4, the total loss function is:
[0036]
[0037] Where L det is the detection loss, L CRF is the CRF loss, L enh is the loss between the enhanced image and the normal image, L fea is the feature map loss, λ enh and λ fea are the corresponding weight values respectively.
[0038] Compared with the prior art, the present invention can significantly improve the accuracy and robustness of target detection in a dark domain environment. Specifically, it is manifested as follows:
[0039] 1) By performing reverse CRF processing, the light intensity received by the camera sensor is obtained, reducing the influence of light changes on target detection.
[0040] 2) The application of the adaptive detection strategy enables the detection method to dynamically adjust according to the actual lighting conditions and target features of the image, improving the flexibility and accuracy of detection.
[0041] 3) The refinement, classification, and recognition of the bounding boxes in the post-processing stage improve the accuracy of target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a schematic flowchart of the zero-shot dark domain adaptive target detection method according to an embodiment of the present invention;
[0043] Figure 2 is a schematic diagram of the CRF inverse solution network in an embodiment of the present invention;
[0044] Figure 3 is a schematic diagram of the ResBlock in an embodiment of the present invention;
[0045] Figure 4 is a schematic diagram of the RepVGG in an embodiment of the present invention
[0046] Figure 5 is a schematic diagram of the enhanced matrix acquisition network in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The present invention will be further described below in conjunction with the embodiments and the drawings, but the protection scope of the present invention should not be limited thereby.
[0048] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the zero-shot dark domain adaptive target detection method according to an embodiment of the present invention. As shown in the figure, a zero-shot dark domain adaptive target detection method includes the following steps;
[0049] S1. Assist in adjusting the parameters of the CRF network through distilling the traditional CRF method. The main goal is to optimize the parameters of the inverse CRF network module so that the CRF network can more accurately simulate and adapt to the conditional probability distribution of image data when performing object detection under low-light conditions. To achieve this goal, this method adopts the strategy of knowledge distillation, combining the advantages of the traditional CRF method and the deep learning network. The specific steps are as follows:
[0050] S1.1 Process the object detection dataset with normal brightness to generate synthetic dark images, including converting from the sRGB format to the RAW format, simulating low-light conditions, then converting back from the RAW format to the sRGB format, and enhancing the images to improve the visibility of the images under low-light conditions and the accuracy of object detection;
[0051] S1.2 For the synthetic dark images, obtain the CRF curve parameters through two paths: one path predicts the CRF curve parameters through the CRF inverse network; the other path obtains the reference CRF curve using the traditional CRF calculation method. Among them, the CRF inverse network is a deep learning-based network, whose input is the synthetic dark image and the output is the predicted CRF curve parameters.
[0052] The CRF parameters output by these two paths are represented in the form of discrete points, that is, the CRF curve is represented by a set of normalized coordinates on a discrete curve.
[0053] To train the CRF network, a loss function is designed, and the L2 norm is used to evaluate the distance between the CRF curve obtained through the CRF inverse network and the reference CRF curve obtained through the traditional method:
[0054] L CRF =||CRF 参考 -CRF x || 2
[0055] S1.3 Specially design a CRF network architecture based on the Residual Network (ResNet) and RepVGG structure. This network architecture is improved on the basis of ResNet18, and the network units of the original ResNet are replaced with the RepVGG model. The RepVGG model has a multi-branch structure, which can better extract features without having an adverse impact on forward propagation. The specific network structure is as Figure 2 shown. First, compress the image using a 7*7 convolution with a stride of 2 and a 3*3 max pooling; then design ResBlocks on the structure of the main part ResNet18, as Figure 2 shown on the right, and replace the network units of the original ResNet with the RepVGG model.
[0056] S1.4 Select the dimension of the output parameter according to the accuracy requirement of the CRF curve. This embodiment uses a discrete point set method to represent the CRF curve. When selecting the reference CRF curve generation method, the fitting model should be consistent and the sampling density should be similar.
[0057] S2. Restore the irradiance map and enhance the irradiance map:
[0058] The restoration of the irradiance map is based on the processing principle of the camera response function (CRF). Using the CRF curve obtained in step S1, the RGB three-channel pixel matrix of the synthetic dark image is inversely operated to obtain the irradiance map, that is, the three-channel light intensity matrix.
[0059] Irradiance map, using residual learning method, a new nonlinear transformation is learned through the network, making the enhanced image more suitable for detection network perception.
[0060] The augmentation matrix is a residual matrix with the same size as the input irradiance map.
[0061] By adding the enhancement matrix to the irradiance map, a final enhancement map is obtained.
[0062] Image enhancement not only considers the difference between human eye perception and network perception, but also considers the perception requirements of the detection network. It plays a key role in the joint optimization system. Figure 2 , Figure 3 As shown. Image enhancement for target detection is a process that takes into account both low-level and high-level feature extraction. The enhancement network uses a path aggregation network architecture. On the basis of highlighting low-level features (such as edges, textures, etc.), it also takes into account high-level enhancement for detection purposes (such as semantic information, target detection, etc.). Based on the ResBlock (residual block) in step S1, the BottleNeck structure and the cross-layer feature aggregation path are designed, as shown in Figure 3 , in order to improve the feature extraction ability of the network.
[0063] The loss function introduces the difference between the normal image and the enhanced image as a constraint to accelerate the convergence speed in the early stage of network training.
[0064] The SSIM (Structural Similarity Index) model is chosen as the method to evaluate the loss. SSIM focuses on the differences in brightness, contrast, and structure between two images. The calculation formula is as follows:
[0065] L enh =SSIM(I 增强 , I 原图)
[0066]
[0067] In the formula, the structural similarity index SSIM is a commonly used image quality evaluation index, which is used to measure the structural similarity between the enhanced image and the reference image. It mainly compares the brightness, contrast, and structure of the two images. Among them, μ x , μ y , σ x , σ y correspond to the data distribution expectations and standard deviations of the two images respectively, and C1 and C2 are constants that need to be externally specified.
[0068] S3. Dynamic optimization of the object detection model based on the pre-trained YOLOv5s: Using the pre-trained YOLOv5s (You Only Look Once v5s) as the basis, and making adjustments and optimizations for specific requirements. The specific steps are as follows:
[0069] Input the enhanced image and the normal image (original image) obtained in step S2 into the detection network at the same time. The difference in feature extraction of the normal image is used as a constraint condition to improve the detection ability under complex or low-light conditions. In this way, the detection module can learn how to accurately identify targets under different lighting conditions.
[0070] The detection network part is adjusted based on the pre-trained YOLOv5s. Considering that there are significant differences in the two parts of image feature extraction and final result output in the adaptive detection task compared with object detection under normal brightness, therefore, fine-tuning is performed through training on the basis of the pre-trained YOLOv5s.
[0071] In terms of sample processing method, considering that the detection of extremely small objects under low brightness does not meet the normal requirements, therefore, a certain resolution is sacrificed, that is, the samples in the sample set with the target bounding box smaller than a certain value are regarded as negative samples.
[0072] In this embodiment, the loss function design includes two parts: feature extraction loss and detection loss:
[0073] Feature extraction loss L fea : Use the Kullback-Leibler Divergence to measure the difference in the output feature distribution. By ensuring the maximum difference, the upper limit of the accuracy of the detection network can be improved, making it better adapt to the feature extraction requirements under different lighting conditions. The formula is as follows:
[0074] L fea = D KL (F 原 ||F 增 )
[0075] In the formula, the KL divergence is an index for measuring the difference between two probability distributions. This method uses the feature distribution extracted from normal images as a reference to measure the difference between the feature distribution extracted from the enhanced image and the reference distribution, where F 原 and F 增 are the feature distributions extracted from the normal image and the enhanced image by the detection network, respectively.
[0076] The detection loss L det is calculated as the linear sum of the regression loss and the classification loss commonly used in object detection:
[0077] L det = λ cls L cls + λ reg L reg
[0078] In the formula, the loss L det is obtained by linearly combining the classification loss L cls and the regression loss L reg . λ cls and λ reg are the corresponding coefficients, respectively.
[0079] The regression loss is used to measure the difference between the predicted bounding box and the ground truth bounding box, while the classification loss is used to measure the difference between the predicted class and the ground truth class. This design of the loss function helps the detection network to accurately identify the location and class of the target during the training and testing processes.
[0080] S4: Design of the loss function.
[0081] Four loss functions are designed, namely the CRF curve loss, the enhancement effect loss, the feature extraction loss, and the detection result loss. The total loss function designed by this method introduces 4 hyperparameters to handle the constraints of different loss functions in the early and late stages of training. Among them, more attention is paid to the CRF loss and the detection loss:
[0082]
[0083] It should be noted that the above specific embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A zero-sample dark domain adaptive target detection method, characterized in that: The following steps are involved: Step 1. Obtain the camera response function (CRF) parameters through the CRF inversion network: The normal brightness object detection dataset is converted into a synthetic dark image, and the CRF curve parameters are predicted using the CRF inversion network based on the RepVGG multi-branch residual structure. At the same time, the reference CRF curve generated by the traditional CRF method is used for distillation training, and the network convergence is constrained by the L2 norm loss of the discrete point set. Step 2. Restore the irradiance map based on the CRF parameters and enhance the irradiance map: An inverse CRF operation is performed on the RGB channels of the synthetic dark image to obtain an irradiance map, a residual matrix is generated through an enhancement network with a cross-layer BottleNeck structure, the irradiance map and the residual matrix are combined to output an enhanced image, and a SSIM loss is introduced to constrain the perceptual consistency of the enhanced image and the normal brightness image; Step 3. Dynamic optimization based on pre-trained detection model: The enhanced image and the normal image are input into the target detection network. Through the difference of the two camera response functions in step 1 and the difference of the KL divergence alignment feature distribution in step 2, the regression loss and classification loss of target detection are combined to jointly optimize the CRF inversion module, the enhancement module and the detection module. Step 4. Multi-task joint loss function design: The total loss function includes CRF reconstruction loss, enhanced SSIM loss, feature KL divergence loss and detection loss. The weights of each loss are adjusted through hyperparameters to achieve end-to-end collaborative training.
2. The zero-sample dark domain adaptive target detection method according to claim 1, characterized in that: The specific steps of step 1 are as follows: Step 1.1 processes the target detection dataset of normal brightness to generate a synthetic dark image; Step 1.2: For the synthetic dark image, two paths are used to obtain CRF curve parameters respectively: one path predicts the CRF curve parameters through the CRF inverse network, and the other path uses the traditional CRF method to obtain the reference CRF curve; Step 1.3: The CRF curve parameters output by the two paths are expressed in the form of discrete point sets, and the distance between the two CRF curves is evaluated using the L2 paradigm as the loss function; Step 1.4 Design a CRF network architecture, which is based on a residual network and uses the RepVGG structure internally to extract image features and predict CRF curve parameters; Step 1.5 selects the dimension of the output parameters according to the accuracy requirement of the CRF curve, and trains the CRF network using the loss function to minimize the difference between the two CRF curves.
3. The zero-sample dark domain adaptive target detection method according to claim 2, characterized in that: In step 1.3, the L2 paradigm is used to evaluate the distance between the CRF curve obtained by the CRF inversion network and the reference CRF curve obtained by the traditional CRF method as the loss function, and the formula is as follows: L CRF =||CRF 参考 -CRF x ||2。 4. The zero-sample dark domain adaptive target detection method according to claim 1 or 2, characterized in that: In step 1.2, the CRF inverse network adopts a RepVGG multi-branch residual structure, including a convolution layer, a maximum pooling layer, and a ResNet18 improved module. The image first passes through the initial convolution layer and the pooling layer for scale reduction, and then passes through 4 ResBlocks in sequence. Each ResBlock will first reduce the length and width of the image to half of the input and then increase the data dimension. ResBlock passes through a RepVGG unit, a nonlinear function, and a RepVGG unit in sequence, and then superimposes the input data and obtains the output through nonlinear processing. The RepVGG unit utilizes the mathematical principle of the convolution kernel. During training, the unit has a multi-branch topological structure, and the final CRF curve output selection is determined by the discrete point set method.
5. The zero-sample dark domain adaptive target detection method according to claim 1, characterized in that: The step 2 specifically includes: Step 2.1: restoring the irradiance map: based on the camera response function, using the CRF curve obtained in step 1, an inverse operation is performed on the RGB three-channel pixel matrix of the synthetic dark image to obtain an irradiance map representing the three-channel light intensity distribution; Step 2.2 Enhance the irradiance map: Use the residual learning method to design an enhancement network, which outputs a residual matrix with the same size as the input irradiance map. The final enhanced map is obtained by adding the residual matrix to the input irradiance map.
6. The zero-sample dark domain adaptive target detection method according to claim 5, characterized in that: The enhancement network adopts a path aggregation architecture, including a low-level feature enhancement path (the layer where P3 is located) and a high-level detection guidance path (the layer where P1P2 is located). Multi-level features are fused through a cross-layer BottleNeck module to output a residual matrix consistent with the size of the irradiance map. The input of the path aggregation network is irradiance maps X1, X2, and X3 of different resolutions obtained by continuous downsampling from bottom to top. Each input has outputs of corresponding resolutions, namely P1, P2, and P3. At the same time, there is a guiding relationship between routes at different layers. X1 is aggregated with X2 after being processed by BottleNeck and ResBlock, and this relationship also exists between X2 and X3. In the final image output stage, P1-P3 are downsampled and then aggregated with the upper layer output with a smaller resolution to finally obtain the output P3 of the entire image. BottleNeck consists of 4 ResBlocks, and a direct connection layer is established between outputs with the same resolution.
7. The zero-sample dark domain adaptive target detection method according to claim 1, characterized in that: In step 3, the target detection network is based on the pre-trained YOLOv5s model, and the feature extraction of the enhanced image is enhanced by fine-tuning and adapting, and the feature distribution difference between the enhanced image and the normal image is constrained by the KL divergence loss. The specific formula is: L fea =D KL (F 原 ||F 增 ) L det =λ cls L cls +λ reg L reg The KL divergence is an indicator that measures the difference between two probability distributions. KL Taking the feature distribution extracted from the normal image as a reference, the difference between the feature distribution extracted from the enhanced image and the reference distribution is measured, where F 原 and F 增 are the feature distributions extracted by the detection network for normal images and enhanced images, L det loss, which is composed of the classification loss L cls With the regression loss L reg The linear combination, λ cls With λ reg are the corresponding coefficients respectively.
8. The zero-sample dark domain adaptive target detection method according to claim 1, characterized in that: In step 4, the total loss function is: L=λ detLdet +λ CRF L CRF +λ enh L enh +λ fea L fea Where L det To detect the loss, L CRF is the CRF loss, L enh To enhance the loss of the image and the normal image, L fea is the feature map loss, λ detCRF , enh and λ fea are the corresponding weight values respectively.