Contrast learning framework FoCo for target detection in foggy scene

Through the unsupervised comparative learning framework FoCo and the foggy-day image synthesis method, the problem of low target detection accuracy in foggy-day environments is solved, and the detection accuracy is improved without increasing the calculation amount and parameter amount, which is suitable for scenarios with high real-time requirements.

CN120339826APending Publication Date: 2025-07-18BEIJING UNION UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510376186.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has low target detection accuracy in foggy environments, and the existing methods have increased the calculation amount and parameter amount, making it difficult to apply to scenarios with high real-time requirements.

Method used

An unsupervised contrast learning framework FoCo is designed to improve the robustness of the target detection model by learning similar features between clear images and synthetic foggy-day images by using the foggy-day image synthesis method based on depth information, and calculate the contrast loss and orthogonal constraint loss optimization model through FoCo.

Benefits of technology

Without increasing the calculation amount and parameter amount, the target detection accuracy in foggy environments is significantly improved, and is suitable for scenarios with high real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339826A_ABST
    Figure CN120339826A_ABST
Patent Text Reader

Abstract

The invention discloses a comparative learning framework for foggy day target detection and a method thereof, and relates to the field of computer vision and deep learning. According to the method, a shared backbone network is used for extracting features of a clear image and synthesizing a foggy day image, and a contrast learning module FoCo is used for calculating contrast loss so as to learn similar features between the foggy day image and the clear image. Meanwhile, a depth information synthetic fog generation method (DSFGM) is adopted to generate a more real synthetic fog day image based on monocular depth estimation, and the generalization ability of the model is improved. Target identification is carried out through a target detection head, and loss functions including detection loss, comparison loss and feature decoupling loss are optimized. The method improves the target detection precision in the foggy environment without increasing the calculation complexity, and is suitable for unmanned driving, intelligent traffic monitoring and other application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and intelligent perception, and particularly relates to a method for improving the target detection performance in foggy weather by using contrastive learning, which can be widely applied to intelligent systems such as autonomous driving, intelligent transportation, video surveillance, and drone navigation to improve the target recognition ability under adverse weather conditions. Background Art

[0002] Object Detection is a computer vision task aimed at identifying specific classes of objects in images or videos and drawing bounding boxes around them. Object detection technology has been widely applied in many fields, such as autonomous driving, video surveillance, drone navigation, maritime shipping, industrial defect detection, and medical image analysis. As an important task in the field of computer vision, Object Detection has experienced a development process from traditional handcrafted feature extraction methods to deep learning-driven end-to-end detection methods. Before the rise of deep learning, object detection mainly relied on manually designed feature extraction methods and combined with machine learning classifiers to complete object recognition.

[0003] In 1994, the Viola-Jones face detection algorithm (VJ algorithm) was proposed, which used Haar features and Adaboost classifiers to achieve a breakthrough in real-time object detection.

[0004] In 2008, the Deformable Part-based Model (DPM) was proposed, which was based on HOG features and used part-level modeling and was widely used for pedestrian and vehicle detection.

[0005] In recent years, with the rapid development of deep learning technology, significant progress has been made in object detection methods based on convolutional neural networks (CNNs).

[0006] In 2015, Joseph Redmon et al. first proposed You Only Look Once (YOLO), which can complete object detection through a single forward propagation, achieving end-to-end real-time detection.

[0007] Current object detection technologies have achieved high detection accuracy in ordinary environments. However, in foggy environments, due to the scattering effect of fog on light, the contrast and clarity of images are greatly reduced, posing many challenges to object detection tasks. The impact of foggy environments on object detection is mainly reflected in the following aspects: 1. Contrast reduction: Fog causes the color contrast between the target and the background to decrease, making it difficult for traditional edge- and texture-based detection methods to effectively segment the target. 2. Color shift and light attenuation: Light is scattered and absorbed by fog during propagation, resulting in an overall white or color-distorted image, affecting detection models based on color features. 3. Target blurring and detail loss: The scattering effect caused by fog weakens the high-frequency information in the image, making the edges of objects blurred and affecting the feature extraction ability of the detection model. 4. Depth dependence: The impact of fog is exponentially related to the distance from the target to the camera. Distant targets are more likely to be obscured by fog or even completely disappear, resulting in a decline in detection performance.

[0008] To overcome the impact of fog on object detection, researchers have proposed many methods.

[0009] The first type of method preprocesses the image through an image dehazing algorithm and then performs object detection. This type of method will seriously affect the overall detection time and is not suitable for autonomous driving scenarios with high real-time requirements.

[0010] The second type of method cascades the dehazing network and the detection network, and optimizes the dehazing loss and the detection loss simultaneously. For example, IA-YOLO cascades a differentiable image processing (DIP) module before the detection network and uses CNN-PP to predict the DIP module. During training, the recovery loss and the detection loss are optimized simultaneously. IA-YOLO was initially proposed by liu et al. in the 2022 paper "Image-Adaptive YOLO for Object Detection in Adverse Weather Conditions".

[0011] The third type of method connects the dehazing network and the detection network in parallel, uses the dehazing network as a branch, and uses the dehazing loss to guide the detection network to extract "clean" features, thereby improving the performance of the detector in foggy environments. For example, DSNet contains two subnets: a detection subnet and a restoration subnet. The restoration subnet shares the feature extraction layer with the detection subnet and uses a feature recovery module to enhance visibility. DSNet was proposed by Huang et al. in their 2021 paper "DSNet: Joint Semantic Learning for Object Detection in Inclement Weather Conditions". TogetherNet helps TogetherNet improve detection performance in adverse weather conditions by sharing the clean features produced by the restoration network. TogetherNet was proposed by Wang et al. in their 2022 paper "TogetherNet: Bridging Image Restoration and Object Detection Together via Dynamic Enhancement Learning". RDMNet introduces a degradation branch based on TogetherNet to model the degraded representation of the degraded image to guide the feature transformation in the restoration and detection branches. RDMNet was proposed by Wang et al. in their 2023 paper "Degradation Modeling for Restoration-enhanced Object Detection in Adverse Weather Scenes". These works have a common feature that they require foggy images as the training set of the detection network, and clear images as a guide, usually introducing a recovery branch, but this will cause a significant increase in the number of parameters and computational complexity of the model, making it difficult to apply in scenarios with high real-time requirements, such as autonomous driving. Unlike these works, we use clear images as training sets and consider unsupervised contrastive learning to improve the performance of object detection in foggy environments without increasing any computational complexity or parameters.

[0012] Contrastive Learning is an unsupervised learning method that aims to learn discriminative feature representations by shortening the representation distance between similar samples (positive samples) and increasing the distance between different samples (negative samples). The core idea is to optimize the neural network by constructing positive and negative sample pairs based on the similarities and differences between samples, so that it can learn more robust representations.

[0013] The Siamese network was proposed by Sumit Chopra et al. in their 2005 CVPR conference paper "Learning a similarity metric discriminatively, with application to face verification". The Siamese network is one of the earliest forms of contrastive learning. It consists of two identical neural networks that process two input samples respectively. After passing through the network with shared parameters, the similarity between sample pairs is finally calculated through metric learning. In this way, the Siamese network can learn whether two samples belong to the same category.

[0014] The Triplet Loss method was proposed by Elad Hoffer et al. in their 2015 paper "Deep metric learning using triplet network. In Similarity-based pattern recognition: third international workshop". It is trained with a triplet of samples, where each triplet consists of an anchor, a positive sample, and a negative sample. The goal is to minimize the distance between the anchor and the positive sample while maximizing the distance between the anchor and the negative sample. The advantage of Triplet Loss is that it can explicitly perform similarity learning and is applicable to many tasks such as face recognition and image retrieval.

[0015] SimCLR was proposed by Ting Chen et al. in their 2020 paper "A simple framework for contrastive learning of visual representations". SimCLR uses data augmentation techniques to generate positive samples, learns representations by maximizing the similarity of positive sample pairs, and uses contrastive loss to optimize the network.

[0016] Momentum Contrast (MoCo) was proposed by Kaiming He et al. in their 2020 paper "Momentum contrast for unsupervised visual representation learning". MoCo is a method for contrastive learning by constructing a dynamic dictionary. The core idea of MoCo is to maintain a dynamically updated dictionary and use the samples in the dictionary as negative samples during training. By introducing a momentum update mechanism, MoCo can alleviate the problem of lagging update of negative samples, thereby improving the stability and performance of the model.

[0017] There is very little research on object detection under foggy conditions using contrastive learning. The present invention learns the similarity features between clear images and corresponding foggy images through contrastive learning, thereby enhancing the robustness of the detector. In addition, in order to make the contrast samples more realistic, we studied the fog synthesis method. We found that the fog synthesis methods used in previous work were all based on the fog center, that is, the fog was thicker near the center point of the picture and thinner in places far from the center point. This is different from the characteristics of foggy images in the real world, where the fog is thicker in places farther from the camera and thinner in places closer to the camera. Therefore, in order to simulate more realistic foggy images, the present invention proposes a foggy image synthesis method based on depth information. Summary of the Invention

[0018] The present invention explores a new technical route for object detection in a foggy environment: designs an unsupervised contrastive learning framework Foggy Contrast (FoCo) to learn the similarity features between clear images and synthesized foggy images and distinguish the features of different images, and a fog addition algorithm based on depth information to simulate a more realistic foggy environment. The present invention effectively improves the robustness of the object detection model without increasing the burden of additional computational amount. The overall framework of the present invention is as Figure 1As shown, we first apply a fogging algorithm to clear images to create a foggy effect, simulating a real foggy environment. Then, we use a shared object detection Backbone to extract the features of both the clear images and the synthesized foggy images. Subsequently, we use FoCo to calculate the contrast loss (ctrloss) and the orthogonality constraint loss (disloss) for the features of the clear images and the foggy images. These two losses will be backpropagated to optimize the parameters of the Backbone. Meanwhile, the features of the clear images are used by the detection head to calculate the detection loss. The object detection Backbone is mainly used to extract the features of clear images and foggy images. The object detection backbone network is mainly designed with reference to TogetherNet. For the input image, we first use the Focus operation to extract the shallow features of the image and halve the spatial scale. Subsequently, the shallow features will go through a feature extraction process at four scales to extract multi-scale features. Dynamic Transformer Feature Enhancement (DTFE) is proposed by TogetherNet to enhance the feature extraction ability of the network. DTFE mainly consists of two Deformable Convolutional Networks (DCN) and a transformer block. Deformable convolution can adaptively deform the convolution kernel, enhancing the flexibility of convolution and thus more flexibly capturing target features. The entire network consists of standard convolutions, which can fully extract local feature information but lacks global information. Therefore, a Transformer module is introduced to model the long-range dependencies between features, which has a positive impact on the extraction of larger target features.

[0019] When the scene changes from normal weather to a foggy scene, the positions and categories of the objects in the scene remain unchanged, only the weather has changed. The features extracted by the model for these two scenes should be similar. Inspired by this idea, we developed an unsupervised contrast learning framework, FoCo, to learn the similar features between clear images and synthesized foggy images, while distinguishing the features of images from different scenes. Specifically, we input the clear images and the corresponding synthesized foggy images into the BackBone in the detection network simultaneously to obtain the clear image features Fc and the synthesized foggy image features Ff.

[0020] Then we use MLP to map Fc and Ff to the feature space while reducing the dimension to avoid a large computational load when calculating similarities. MLP consists of an average pooling layer, two linear layers, and a ReLu activation function. To ensure the independence and non-redundancy of the features of the clear images and the synthesized foggy images captured by the backbone network in the feature spaces Fc and Ff, we impose an orthogonality constraint on Fc and Ff:

[0021]

[0022] Fc and Ff are a pair of positive samples, and the similarity between Fc and Ff is calculated by dot product.

[0023] Since contrastive learning usually requires a large number of negative samples to enhance the discriminative ability of the model, inspired by MoCo, as Figure 2 shown, we construct a circular queue to store negative samples, and then calculate the similarity between Fc and a large number of negative samples in the circular queue.

[0024] Finally, we calculate the contrastive loss of FoCo:

[0025]

[0026] where τ is the temperature hyperparameter used to control the distribution shape. If τ is too large, the contrastive loss treats all negative samples equally, which will lead to the model learning without focus. If τ is too small, the model will only focus on particularly difficult negative samples and it is difficult to generalize. The present invention uses τ = 0.07, k i is the negative sample in the circular queue, and K is the total capacity of the circular queue.

[0027] The foggy day synthesis method used in the past uses the center point of the image as the center of fogging. The closer to the center of the image, the thicker the fog. The feature distribution of this synthesized foggy day image is quite different from that of the real foggy day image. In the real foggy day image, the fog is generally thicker in the place farther from the camera, and the fog is thinner in the place closer to the camera. Therefore, the present invention designs a fog adding algorithm based on depth information, and the overall process is as Figure 3 shown. There are many methods to obtain the depth information of the image, but most of them require relatively complex conditions. For example, using two cameras to shoot the same scene from different angles, and calculating the disparity between the images to infer the depth information. The larger the disparity, the closer the object is to the camera, and the smaller the disparity represents that the object is farther away. In actual situations, for the same scene, we may only have a single clear image. Therefore, we choose to use the pre-trained monocular depth estimation model Midas to estimate the depth information of the image using a single clear image. The saturation and brightness of the clear image are too large to provide sufficient color gradients or texture information, resulting in Midas misestimating the depth. Therefore, we adaptively adjust the brightness and saturation of the input clear image. First, we convert the color space of the clear image from the RGB space to the HSV space, and then adaptively adjust the brightness and saturation using the following formula.

[0028]

[0029] Then we use Midas to perform depth estimation on the preprocessed image to obtain the depth information, and use the depth information to synthesize fog on the original image through the atmospheric scattering model.

[0030] The loss function of FoCo is divided into three parts: (1) Detection loss (2) feature distangle loss (3) contrastive loss. Among them, the parts optimized by Detection loss include the detection Backbone and the detection head, while the distangle loss and contrastive loss only act on the Backbone part. The feature distangle loss is used to decouple the feature vectors extracted by the Backboone, and the contrastive loss helps the Backbone learn the similarity features between clear images and foggy images. Finally, the total loss expression of the method we proposed is as follows:

[0031] L total = αL Detection + β(L ctr + γL dis )

[0032] Among them, α, β, and γ are hyperparameters used to balance the detection loss, contrast loss, and decoupling loss.

[0033] Compared with previous foggy environment object detection methods, the present invention uses unsupervised contrast learning to greatly improve the accuracy of object detection without increasing the computational complexity and the number of parameters of the model, enabling the detector to be deployed in scenarios with high real-time requirements. In addition, our foggy image synthesis method is also more in line with real-world foggy images. We tested on multiple real-world foggy image test sets and compared with other methods, and visualized the test results as Figure 4 shown. It can be seen that the method of the present invention can detect more objects and has higher confidence. Description of the Drawings

[0034] Figure 1 Shows a schematic diagram of the network structure of an embodiment of the FoCo-YOLOx network for the object detection method in a foggy scenario of the unsupervised contrast learning framework FoCo.

[0035] Figure 2 Shows the principle of action of the unsupervised contrast learning framework FoCo

[0036] Figure 3 Shows the fog addition algorithm process based on depth information

[0037] Figure 4 Shows the comparison results of FoCo-YOLOX and other models on the RTTS test set

[0038] Figure 5 The flowchart of a preferred embodiment of the object detection method in a foggy scene showing the unsupervised contrast learning framework FoCo

[0039] Figure 6 Showing the fogging effects of different concentrations based on depth information Detailed implementation manners

[0040] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments

[0041] Embodiment 1:

[0042] As shown in Figure 5 , step 100 is executed to construct a clear image dataset for model training. First, download the VOC2012 and VOC2017 datasets, and screen out the target categories related to the driving scene (such as cars, pedestrians, buses, motorcycles, etc.) from them to construct a clear image training set VOC_Clean_Train

[0043] Step 110 is executed to construct a foggy image dataset for model training. To generate training data that is more in line with the real foggy environment, in this embodiment, the monocular depth estimation network MiDas is used to calculate the depth information of the image, and the atmospheric scattering model is combined for foggy image synthesis: 1. Calculate the depth map of the image and perform normalization processing; 2. Calculate and synthesize the foggy image according to the atmospheric scattering model, making the fog thicker in the distant area and thinner in the near area to improve the authenticity of the synthesized data; 3. Generate a foggy image training set (VOC_Foggy_Depth) based on the depth information

[0044] Step 120 is executed through the constructed VOC_Clean_Train and VOC_Foggy_Depth. As shown in Figure 2 , the foggy synthetic image and the clear image are input into the shared object detection Backbone to extract the features of the clear image and the synthetic foggy image, and then the FoCo algorithm is applied to the two features. To improve the adaptability of the detector to the foggy environment, in this embodiment, the orthogonality constraint loss and the contrast loss are adopted in the optimization process

[0045] Step 130 is executed to save the optimal weights of the model during the training process, and test the optimized object detector on the real foggy datasets RTTS and Foggy_Driving, and compare it with the existing methods to verify the effectiveness of this method. The comparison results are shown in the table

[0046] Table 1 Comparative experiments with other methods on RTTS

[0047]

[0048] Example 2: Foggy Image Synthesis Method Based on Depth Information

[0049] This example proposes a foggy image synthesis method based on depth information to generate more realistic foggy training data and improve the generalization ability of object detection in foggy environments. The method steps are as follows:

[0050] Step 1: Input clear image: Select a clear image dataset containing traffic scenes and convert it to the HSV color space to reduce the influence of highlight areas.

[0051] Step 2: Use a pre-trained monocular depth estimation network (such as MiDas) to calculate the depth map of the input image: Normalize the depth information to adapt it to the subsequent foggy synthesis model.

[0052] Step 3: Generate a synthetic foggy image. Calculate the transmittance using the atmospheric scattering model:

[0053] I(x) = J(x)t(x) + A(1 - t(x))

[0054] t(x) = e -βd(x)

[0055] where I(x) is the synthetic foggy image, J(x) is the clear image, A is the atmospheric light component, t(x) is the transmittance controlled by the depth information, d(x) is the depth information predicted by the monocular depth estimation model, and β is used to control the concentration of the synthetic fog. In this example, the value range of β is set to [0.005, 0.015]. The synthesis result is as Figure 6 shown.

[0056] Although the foregoing disclosure discusses exemplary solutions and / or embodiments, it should be noted that many changes and modifications can be made herein without departing from the scope of the described solutions and / or embodiments as defined by the claims. Moreover, although the elements of the described solutions and / or embodiments are described or claimed in the singular, the plural cases can also be contemplated unless explicitly stated to be limited to the singular. Additionally, all or part of any solution and / or embodiment can be used in combination with all or part of any other solution and / or embodiment, unless otherwise indicated.

Claims

1. A method for object detection using unsupervised contrastive learning, characterized in that, The framework includes: a shared backbone network for extracting features of clear images and synthetic foggy images; a contrastive learning module FoCo that constructs positive and negative sample pairs, learns the similar features of clear images and foggy images through contrastive loss, and ensures the independence of different features in the representation space through feature decoupling constraints; a contrastive learning module FoCo that constructs positive and negative sample pairs, learns the similar features of clear images and foggy images through contrastive loss, and ensures the independence of different features in the representation space through feature decoupling constraints; an object detection head for performing object detection using the features of clear images. The following steps are used for model training: Steps: Step 1: Construct a dataset, including a clear image dataset and a synthetic foggy image dataset, where the foggy images are generated by a fog-adding algorithm based on depth information. Step 2: Use the shared backbone network to extract the features of clear images and foggy images, and perform feature contrastive learning through FoCo to optimize the feature representation ability; Step 3: Use the object detection head to train the detector, calculate the detection loss, contrastive loss, and feature decoupling loss during the training process, and perform gradient optimization; Step 4: After the training is completed, use the trained model for object detection, which is applicable to object recognition tasks in foggy environments.

2. The contrastive learning framework according to claim 1, wherein The shared backbone network includes: a multi-scale feature extraction module that uses a convolutional neural network (CNN) or a Transformer to extract image features at different scales; a dynamic feature enhancement module (DTFE) that combines deformable convolution (DCN) and a Transformer to enhance the feature extraction ability.

3. The contrastive learning framework according to claim 1, wherein The contrastive learning module FoCo includes: a shared backbone network for extracting the features of clear images and foggy images; the positive sample pairs include clear images and their corresponding synthetic foggy images, and the negative sample pairs include other images and their synthetic foggy versions; a cyclic queue is used to store negative samples, and the similarity between clear images and negative samples is calculated to optimize the detector's ability to learn more stable representations; the contrastive loss calculation includes calculating the similarity of positive sample pairs while maximizing the distinguishability between clear images and negative samples.

4. The contrastive learning framework according to claim 1, wherein The depth information-based synthetic fog generation method (DSFGM) includes: using a monocular depth estimation model (such as Midas) to obtain the scene depth information; generating a transmission map based on the depth information to ensure that the synthetic foggy images conform to the characteristics of thick fog in the distance and thin fog near in a real foggy environment; using an atmospheric scattering model to adjust the image brightness and saturation to further optimize the authenticity of the foggy synthetic images.

5. The contrastive learning framework according to claim 1, wherein The object detection head uses a lightweight YOLO series detector, including: predicting the object category and bounding box; optimizing the backbone network through contrastive learning to enable the detector to maintain high detection accuracy in foggy scenarios.

6. The contrastive learning framework according to claim 1, wherein The FoCo framework improves the detection robustness in foggy environments without increasing the computational complexity of the detector.