Automatic driving image defogging method based on CTFormer

Through the CTFormer model, combined with convolutional feature extraction and improved Transformer architecture, the problems of insufficient global information modeling capabilities, poor detail recovery effects and lack of real-time performance of image defog technology in autonomous driving scenarios are solved, and efficient and accurate defog image generation and visual perception are improved.

CN119941576AActive Publication Date: 2025-05-06FUDAN UNIVERSITY

Patent Information

Application Number
CN202411909732.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-06
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

The existing image defog removal technology has problems such as insufficient global information modeling capabilities, poor detail recovery effect and lack of real-time performance in autonomous driving scenarios.

Method used

The CTFormer model combined with convolution feature extraction and improved Transformer architecture is adopted to achieve efficient and accurate fogging image generation through multi-scale feature modeling, Taylor expansion attention mechanism, contrast constraint module and detail enhancement module.

Benefits of technology

It significantly improves the quality and processing efficiency of the defog removal image, improves the robustness and reliability of visual perception in autonomous driving scenarios, and meets the real-time processing needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941576A_ABST
    Figure CN119941576A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of automatic driving, and particularly relates to an automatic driving image defogging method based on a CTFormer. The method comprises the following steps: by introducing a learnable position code, pre-processing an input fogged image, and enhancing the perception ability of a model to spatial position information; an improved encoder is utilized, local convolution and a multi-head self-attention mechanism based on Taylor expansion are combined, multi-scale features of the image are modeled, and global context information and local details are captured; gradually recovering a high-resolution image through jump connection and multi-scale up-sampling by using a decoder, and enhancing image edge and high-frequency information in combination with detail enhancement; and a clear defogged image is generated through the output module. According to the method, the contrast learning strategy and optimization design of various loss functions are combined, scene details under the foggy weather condition are accurately restored, the calculation efficiency and robustness of a defogging model are remarkably improved, and high-quality visual perception input is provided for an automatic driving system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and in particular relates to an image defogging method in an autonomous driving scenario. Background Art

[0002] With the rapid development of autonomous driving technology, the visual perception system, as one of the key components, is crucial for tasks such as scene understanding, object detection, and path planning. However, in severe weather conditions (such as fog), the quality of images collected by the camera is severely degraded, resulting in a significant reduction in the performance of the perception system. This is because fog causes light scattering and attenuation, making the image appear low-contrast, blurry, and with lost details.

[0003] Traditional image dehazing methods can be divided into two categories: physical model driven and data driven. For example, physical model driven methods: for example, atmospheric scattering models, restore clear images by calculating atmospheric transmittance and ambient light components. However, such methods are highly dependent on model parameters and are difficult to cope with complex real scenes. Another data driven method is image dehazing technology based on deep learning, which has made significant progress in recent years. These methods use neural networks to model the global and local features of fogged images, but most methods still have the following limitations:

[0004] Insufficient local features: Traditional convolutional networks have difficulty capturing long-range contextual relationships and cannot effectively handle the balance between global information and local details in images.

[0005] Weak model generalization ability: It relies on a specific dehazing dataset during the training phase and is difficult to cope with complex dehazing situations in different scenarios.

[0006] The Transformer architecture has achieved success in natural language processing and computer vision due to its excellent global feature modeling capabilities. In recent years, hybrid models combining convolutional neural networks (CNNs) and Transformers have shown great potential in image processing tasks. However, the direct application of existing Transformer models in image dehazing tasks still faces the following challenges:

[0007] High computational complexity: The traditional multi-head self-attention mechanism has high computational cost on high-resolution images and is difficult to process in real time.

[0008] Insufficient detail recovery capability: Modeling based solely on global features can easily lead to loss of image detail information. Summary of the invention

[0009] The purpose of the present invention is to provide an autonomous driving image defogging method based on CTFormer to solve the problems of insufficient global information modeling capability, poor detail restoration effect and lack of real-time performance of existing image defogging technologies in autonomous driving scenarios.

[0010] The CTFormer-based autonomous driving image defogging method proposed in this invention adopts a model combining convolutional feature extraction and improved Transformer architecture, denoted as CTFormer; through multi-scale feature modeling, Taylor unfolded attention mechanism (T-MHSA), contrast constraint module (CCM) and detail enhancement module, efficient and accurate defogging image generation is achieved, the quality and processing efficiency of defogging images are significantly improved, and higher robustness and reliability are provided for visual perception in autonomous driving scenarios;

[0011] The CTFormer model includes an input preprocessing module, an encoder module, a decoder module, a contrastive constraint module (CCM), a loss function and optimization module, and an output module; among them:

[0012] Input preprocessing module: Figure 5 As shown, the input fogged image is normalized and data augmented, and the spatial information expression capability of the input features is enhanced using learnable position encoding, so that the model can adapt to different weather and environmental conditions.

[0013] Encoder module: The improved Transformer architecture is adopted to optimize the traditional multi-head self-attention mechanism into an efficient attention mechanism based on Taylor expansion (T-MHSA), which is combined with local convolution operations to capture the multi-scale global features of the image and enhance the model's ability to represent the fogged area.

[0014] Decoder module: The decoder gradually restores the high resolution and detail information of the image through multi-scale upsampling and cross-layer skip connections; combined with the detail enhancement module, it further improves the clarity and edge performance of the image.

[0015] Contrastive Constraint Module (CCM): It uses contrastive learning strategy to construct the feature distribution of anchor points, positive samples and negative samples, optimizes the relative relationship between features through contrast loss, and improves the discriminative ability of dehazing features.

[0016] Loss function and optimization module: Design a variety of loss functions, including pixel-level loss, perceptual loss, and contrast loss, and optimize the model training process through dynamic weight combination to ensure the quality and consistency of the final output image.

[0017] The output module maps the feature map generated by the decoder into a high-definition dehazed image, and combines post-processing techniques such as edge smoothing and detail enhancement to output high-quality dehazing results that meet the needs of autonomous driving.

[0018] The CTFormer-based automatic driving image defogging method proposed in this invention has the following specific steps:

[0019] (1) The foggy image is preprocessed by the preprocessing module:

[0020] Preprocess foggy images collected in autonomous driving scenarios and convert the original images into standardized input tensors that can be processed by the model; specifically, the following steps are involved:

[0021] Normalization operation: Convert the pixel value I(x,y) of the input image to the interval [0,1] to eliminate the brightness difference between different input images. The formula is:

[0022]

[0023] Among them, min(I) and max(I) are the minimum and maximum values ​​of image pixels, respectively. ′ (x,y) is the normalized pixel value.

[0024] Data augmentation: Expanding the training dataset by:

[0025] Random cropping: crop random areas of the image to retain the main information of the image;

[0026] Flip operation: randomly flip the image horizontally or vertically;

[0027] Add noise: introduce Gaussian noise or salt and pepper noise to simulate sensor noise in real environment;

[0028] Lighting adjustment: Randomly adjust image brightness, contrast, and saturation to improve the lighting robustness of the model.

[0029] Learnable position encoding: By adding a learnable position encoding matrix PE, the model can perceive the spatial position information of the image. The calculation formula of the position encoding is:

[0030]

[0031] Where i represents the position, j represents the dimension index, and d is the encoding dimension. The position encoding matrix is ​​added element-wise to the input image tensor to enhance the model's understanding of spatial relationships.

[0032] (2) The encoder module encodes the preprocessed image, that is, the preprocessed foggy image is sent as the input tensor to the encoder module, and multi-scale features are extracted through a hybrid attention mechanism (such as T-MHSA) and local convolution; the feature map output by the encoder is used as the input of the subsequent module to provide a multi-scale, hierarchical feature representation. Specifically:

[0033] The encoder module extracts multi-scale features of the input image and models global dependencies; the encoder module combines the Taylor-expanded Multi-Head Self-Attention mechanism (T-MHSA) and local convolution feature extraction technology. The specific calculation process includes the following steps:

[0034] (2.1) Input feature normalization and block operation;

[0035] Layer Normalization: Normalize the input tensor Perform layer normalization to improve the stability of features. The calculation formula is as follows:

[0036]

[0037] Among them, μ is the characteristic mean, σ 2 is the characteristic variance, ∈ is the stability constant, T ′ is the normalized feature tensor.

[0038] Block operation: normalize the input image features T ′ Divide the window size into N non-overlapping sub-blocks according to P×P The size of each sub-block is P×P×C.

[0039]

[0040] (2.2) Global feature modeling based on T-MHSA;

[0041] Query, key and value generation: For each sub-block B i Generate the query matrix Q, key matrix K, and value matrix V:

[0042] Q=B i W Q ,K=B i W K ,V=B i W V , (5)

[0043] Among them, W Q ,W K , is a learnable weight matrix, and d represents the dimensions of query, key, and value.

[0044] Attention weight calculation (Taylor expansion optimization): Taylor expansion is used to approximate the Softmax function to reduce the computational complexity. The attention score calculation formula is as follows:

[0045]

[0046] Through Taylor expansion, the Softmax function can be approximated as:

[0047]

[0048] The calculation formula for attention output OOO is:

[0049] O=A(Q,K)V, (8)

[0050] Multi-head attention mechanism: The attention outputs of multiple heads are concatenated and linearly transformed. The calculation formula is:

[0051] O multi-head =Concat(O1,O2,…,O h )W o , (9)

[0052] in, is the linear transformation matrix, and h is the number of attention heads.

[0053] (2.3) Local feature extraction enhanced by hybrid convolution;

[0054] Local convolution operation: for each sub-block B i A convolution operation is performed to extract fine-grained features. The convolution kernel size is 3×3 and the calculation formula is:

[0055] F conv =Conv2D(B i ,W conv ), (10)

[0056] Among them, W conv is the convolution kernel parameter.

[0057] Fusion of local and global features: The local features F extracted by convolution conv With the global feature O multi-head After addition, the layers are normalized:

[0058] F fusion =LayerNorm(F conv +O multi-head ), (11)

[0059] (2.4) Output of encoder;

[0060] Multi-scale feature representation: The encoder aggregates the processed features through multi-scale convolution to generate feature maps of different resolutions:

[0061] F multi-scale =Concat(F scale 1,F scale 2,…,F scaleN ), (12)

[0062] Layer stacking: Multiple encoding layers are stacked to further enhance the feature representation capability, and the output feature map F encoder It will be used as input to the decoder.

[0063] (3) The decoder module restores the multi-scale features output by the encoder into a high-quality clear image, and gradually restores the image resolution and details through multi-layer upsampling and detail enhancement modules; the specific process is as follows:

[0064] (3.1) Cross-layer skip connection;

[0065] Feature fusion of skip connections: extracting multi-scale features from each layer of the encoder And pass it to the corresponding layer of the decoder through the skip connection:

[0066]

[0067] in, is the output of the previous layer of the decoder, The fused features.

[0068] Feature transformation: fusion of features Transformed by 3×3 convolution to enhance the expressiveness of features while reducing redundant information:

[0069]

[0070] Among them, W conv is the convolution kernel parameter.

[0071] (3.2) Multi-scale upsampling

[0072] Upsampling operation: The decoder uses a layer-by-layer upsampling module to restore the low-resolution feature map to high resolution, and uses the nearest neighbor interpolation method to achieve upsampling. The formula is:

[0073]

[0074] Multi-scale feature fusion: During the upsampling process, a multi-scale fusion module is introduced to combine feature maps of different resolutions to improve detail recovery capabilities:

[0075]

[0076] Among them, w i is the weight of each scale, and N represents the number of scales.

[0077] (3.3) Detail enhancement

[0078] High-frequency information extraction: Use the refined convolution operation to extract high-frequency information to enhance the clarity of details and edges:

[0079] F detail =Conv2D(F multi-scale ,W detail ), (17)

[0080] Among them, W detail is the weight of the refined convolution.

[0081] Residual enhancement: Combine the original features F output by the encoder original And the output of the current decoding layer to form a residual connection:

[0082]

[0083] (3.4) Decoder output

[0084] Restore the original resolution: The last layer of the decoder uses a fully connected layer to map the features to the original image size H×W and output a clear image:

[0085] I clear =Conv2D(F enhanced ,W output ), (19)

[0086] Dehazing result output: The image pixel values ​​are limited to the range of [0,1] through a nonlinear activation function to generate a dehazed high-definition image.

[0087] (4) The contrast constraint module (CCM) optimizes the distinguishing ability of features and enhances the comprehensive performance of the encoder and decoder through contrastive learning strategy. The specific process is as follows:

[0088] (4.1) Construction of anchor samples and positive and negative samples;

[0089] Anchor sample definition: Select the global feature f of the current input image from the output of the encoder a as anchor feature.

[0090] Positive sample definition: Generate an enhanced version of the same input image through random perturbations (such as cropping, rotation, etc.) and extract its features f from the corresponding encoder output p , as a positive sample.

[0091] Negative sample definition: randomly select features f from the encoder output of other input images n , as negative samples.

[0092] (4.2) Contrastive learning optimization;

[0093] Contrastive loss design: The contrastive constraint module uses the contrastive loss function to optimize the feature distribution between anchor points, positive samples and negative samples, minimizing the distance between anchor points and positive samples while maximizing the distance between anchor points and negative samples. Its core formula is:

[0094]

[0095] Among them, m is the preset margin value, which controls the feature difference between positive and negative samples.

[0096] Feature regularization: In the contrastive learning process, the CCM module regularizes the anchor feature f a Add regularization constraints to prevent the model from overfitting to certain local features:

[0097]

[0098] Joint loss optimization: The contrast loss and the dehazing loss (such as perceptual loss and pixel loss) are weighted combined to form a total loss function:

[0099]

[0100] Among them, λ1 and λ2 are weight coefficients that control the impact of different loss terms.

[0101] (4.3) Feature distance calculation and update;

[0102] Feature distance calculation: In the feature space, the Euclidean distance is used to calculate the distance between the anchor point and the positive sample and the negative sample:

[0103] d pos =||f a -f p ||,d neg =||f a -f n ||, (23)

[0104] Gradient update: The contrastive constraint module optimizes model parameters through back propagation so that the feature distance meets the contrastive learning objective:

[0105] Target:d pos →0,d neg →∞

[0106] (4.4) Enhanced diversity in contrastive learning;

[0107] Multi-view positive sample generation: Use data enhancement (such as viewpoint change, noise addition, etc.) to generate multiple positive samples to enhance the model's adaptability to different scenarios.

[0108] Dynamic negative sample sampling: Prioritize samples that are close to the anchor point features as negative samples to improve the optimization effect of contrast loss.

[0109] (5) Loss function and optimization processing during model training. Specifically, the quality of dehazed images and the robustness of the model are improved through the joint design and optimization of multiple loss functions. The loss function mainly includes the optimization of pixel-level loss, perceptual loss, contrast loss, and total loss. The specific process is as follows:

[0110] (5.1) Pixel-level loss;

[0111] L1 loss: directly calculates the absolute difference between the dehazed image and the clear target image at the pixel level to maintain pixel consistency.

[0112] Structural Similarity Loss (SSIM): Measures the similarity between the dehazed image and the target image in terms of brightness, contrast, and structure to improve the consistency of visual effects.

[0113] (5.2) Perceptual loss;

[0114] Feature perception: Use a pre-trained deep convolutional neural network (such as VGG) to extract high-level features of the dehazed image and the target image, use perceptual loss to measure the feature difference between the two, and emphasize the semantic consistency of the image.

[0115] Multi-scale perception: Calculate the perceptual loss from multiple scales to further improve the clarity and detail of the dehazed image.

[0116] (5.3) contrast loss;

[0117] Relative distance between features: The contrastive learning module (CCM) is used to optimize the feature distance between anchor points, positive samples, and negative samples, thereby enhancing the model’s ability to distinguish fogged images.

[0118] (5.4) Joint optimization of total loss;

[0119] Weighted combination: The pixel-level loss, perceptual loss, and contrast loss are weighted together to form a total loss function:

[0120]

[0121] Among them, λ1, λ2, and λ3 are the weight coefficients of each loss term, which control their impact on model optimization.

[0122] Dynamic weight adjustment: Dynamically adjust the weight of the loss term according to the training stage and the output quality of the dehazed image to ensure that the contribution of each loss term is balanced at different stages.

[0123] (5.5) Optimization strategy;

[0124] Optimizer selection: Adopt Adam optimizer, combined with momentum term and learning rate decay mechanism to accelerate model convergence.

[0125] Learning rate scheduling: Dynamically adjust the learning rate during the training process. The initial learning rate is high and gradually decreases with training iterations to improve the model's ability to optimize details in the later stages.

[0126] (6) The output module transforms the feature map generated by the decoder into the final dehazed image through mapping and post-processing to ensure the visual effect and practicality of the output. The specific process is as follows:

[0127] (6.1) Feature mapping;

[0128] Linear transformation: Use a fully convolutional layer to map the feature map output by the decoder to a clear image consistent with the resolution of the original image:

[0129] I clear =Conv2D(F decoder ,W output ), (25)

[0130] Among them, F decoder is the output feature map of the last layer of the decoder, W output is the weight of the mapping convolution kernel.

[0131] Non-linear activation: Applying an activation function (such as Sigmoid or ReLU) to the output image restricts the pixel values ​​to the range [0, 1] to enhance the visual consistency and contrast of the image.

[0132] (6.2) Artifact removal processing;

[0133] Edge smoothing: Smoothing the edge areas of the generated image to reduce possible artifacts, achieved through a Gaussian filter:

[0134] I smooth = GaussianFilter (I clear ), (26)

[0135] Detail enhancement: Use Laplace filter to enhance the high-frequency details of the image and optimize the detail performance.

[0136] (6.3)Multiple output level support;

[0137] Low-resolution preview: To meet real-time requirements, output a low-resolution clear image preview for quick use.

[0138] High-resolution primary output: Provides a complete high-resolution dehazed image for subsequent target detection, path planning, etc. in autonomous driving tasks.

[0139] (6.4) Quality assessment feedback;

[0140] Image quality evaluation: The quality of the dehazed image is evaluated based on the SSIM (structural similarity) and PSNR (peak signal-to-noise ratio) indicators, and an evaluation report is generated to optimize the training of subsequent models.

[0141] Feedback mechanism: Automatically adjust model parameters (such as learning rate or loss function weight) according to the evaluation results to improve training efficiency.

[0142] The CTFormer-based autonomous driving image defogging method of the present invention has the following significant advantages:

[0143] (1) Effective combination of global and local features: The encoder module achieves balanced modeling of global context information and local details through the combination of T-MHSA and local convolution, significantly improving the feature expression ability of the fogged area.

[0144] (2) Efficient computing and real-time support: The Taylor expansion is used to optimize the attention mechanism, which significantly reduces the computational complexity of the model, enabling it to meet the real-time processing requirements in autonomous driving scenarios.

[0145] (3) Multi-module collaborative optimization: The combination of the decoder module and the CCM module fully utilizes the generalization ability of the contrastive learning enhancement model to ensure the consistency of the dehazing effect in different scenarios and conditions.

[0146] (4) Image detail restoration and clarity improvement: The decoder module restores the image detail information through multi-scale fusion and detail enhancement, generates a high-resolution, clear dehazed image, and supports subsequent target detection and path planning tasks.

[0147] The present invention provides an efficient and reliable image defogging solution for autonomous driving systems and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0148] Figure 1 It is a diagram of the CTFormer model framework of the present invention.

[0149] Figure 2 It is a schematic diagram of multi-scale fusion.

[0150] Figure 3 It is a diagram of the linear Taylor expansion computational complexity of the MSA module.

[0151] Figure 4 Yes x And a schematic diagram of its first-order Taylor expansion curve.

[0152] Figure 5 This is a diagram of the SwinTransformer network architecture. DETAILED DESCRIPTION

[0153] Step 1: Preprocessing of input image

[0154] First, the fogged images collected by the autonomous driving camera system are input into the system. In order to unify the data format and reduce the impact of illumination changes on the model, the images need to be normalized, and the pixel values ​​are standardized from the range of 0-255 to between 0-1. The normalized images are more suitable for subsequent deep learning processing modules. Then, in order to improve the generalization ability of the model, the training data set is expanded through data enhancement techniques, including random cropping, horizontal flipping, rotation, brightness and contrast adjustment, etc. These enhancements can simulate different driving scenarios and camera states, helping the model adapt to complex environments. In addition, adding Gaussian noise and salt and pepper noise can further improve the robustness of the model under camera noise conditions.

[0155] In order to improve the model's ability to perceive spatial information, learnable position encoding is added. Position encoding enhances the representation of spatial features in the image by generating the position information of each pixel. After the position encoding is fused with the image features, an input tensor with enhanced spatial perception is formed, providing a more complete feature representation for the subsequent encoder module.

[0156] Step 2: Feature extraction of encoder module

[0157] The preprocessed image is input to the encoder module of the CTFormer model. The encoder first blocks the image into non-overlapping patches of fixed size. Each patch is fed into a convolutional neural network (CNN).

[0158] Extract local features such as edges, textures, and shapes. These local features provide the model with the detailed information required for the dehazing task.

[0159] Subsequently, the encoder uses an improved Transformer structure to model the global features of the image. Through the multi-head self-attention mechanism (T-MHSA), the model can capture the contextual relationship between long-distance regions, thereby more comprehensively understanding the image content. T-MHSA optimizes the computational complexity of the attention mechanism through Taylor expansion, enabling the model to efficiently process high-resolution images while maintaining the integrity of feature extraction.

[0160] Finally, the encoder fuses local features with global features and generates hierarchical feature representations through multi-scale convolution operations. These multi-level features contain rich details and global semantic information, providing support for the decoder module.

[0161] Step 3: Feature optimization of comparison constraint module

[0162] The features generated by the encoder are further passed to the contrast constraint module (CCM), which optimizes the feature space distribution through contrastive learning. First, the features of anchor samples, positive samples, and negative samples are extracted from the encoder output. The anchor sample is the feature representation of the current input image. The positive sample is obtained by slightly enhancing the anchor image (such as brightness adjustment, rotation), and the negative sample is randomly selected from other images.

[0163] CCM optimizes the feature distance between anchor samples and positive and negative samples through the contrast loss function, making the anchor closer to the positive sample and farther from the negative sample. Through this feature space optimization strategy, the model can more accurately extract the key features of the fogged area and improve the discrimination ability of the defogging task.

[0164] Step 4: Feature reconstruction of decoder module

[0165] The decoder module restores the multi-scale features generated by the encoder to the resolution of the original image through layer-by-layer upsampling and feature fusion. During the upsampling process, the high-resolution features of the encoder are extracted through jump connections and fused with the features of the decoder to retain the detailed information of the image. The fused features are processed by convolution operations and activation functions to further enhance high-frequency information and edge details.

[0166] The decoder also contains a detail enhancement module to refine the fused feature map. This module generates clearer and more natural dehazed images by enhancing information in edges and high-frequency areas.

[0167] Step 5: Dehazed image generation

[0168] The feature map output by the decoder is mapped to a dehazed image, and the pixel values ​​are limited to the range of 0-1 using an activation function (such as Sigmoid). In order to further improve the visual effect of the image, the image is post-processed at the output stage, including edge smoothing, color adjustment, and contrast enhancement. The resulting dehazed image is not only clear, but also retains rich detail information, providing reliable input data for target detection and path planning in autonomous driving scenarios.

[0169] Step 6: Design and optimization of loss function

[0170] During the model training process, the design of the loss function is crucial to the dehazing effect. This method uses a variety of loss functions, including pixel-level loss, perceptual loss, and contrast loss. Among them, the pixel-level loss measures the pixel difference between the dehazed image and the clear image to ensure the brightness and color consistency of the output image; the perceptual loss measures the semantic difference between the two images through the features extracted by the pre-trained model, emphasizing the structure and detail retention of the image; the contrast loss optimizes the feature distribution through contrast learning to improve the model's ability to distinguish fogged areas.

[0171] The weights of all loss functions are dynamically adjusted according to the actual results, ensuring that the model can focus on optimizing key parts at different training stages. Finally, all loss functions are combined into a total loss function, and the model parameters are optimized through back propagation to gradually improve the dehazing effect.

[0172] Step 7: Model testing and deployment

[0173] The trained model is evaluated on an independent test set to verify its performance in real-world scenarios. The quality of the dehazed image is quantified through metrics such as SSIM (Structural Similarity Index) and PSNR (Peak Signal-to-Noise Ratio), and the clarity and detail of the image are ensured through visual observation.

[0174] Table 1 and Table 2 compare the model with other methods through experiments.

[0175] As shown in Table 1, in the experimental results on the indoor haze dataset IHAZE, the CT-Former method outperforms other mainstream dehazing methods in all indicators. Specifically, CT-Former achieved scores of 43.02dB and 0.912 on PSNR and SSIM, respectively, significantly surpassing FFA-Net and TransWeather, demonstrating excellent detail recovery and structure retention capabilities. At the same time, CT-Former's NIQE score is 3.47, slightly better than USC-Former's 3.48, showing good image naturalness. Overall, CT-Former not only achieves high-quality dehazing effects in the IHAZE dataset, but also maintains the naturalness of the image well, verifying its application advantages in complex indoor haze scenes.

[0176] As shown in Table 2, the experimental results on two outdoor haze datasets (Foggy Haze and Dense Haze) show that the CT-Former method shows excellent performance in the dehazing task, which is superior to several existing mainstream dehazing methods, including DCP, AOD-Net, FFA-Net, TransWeather and USC-Former. Specifically, CT-Former achieved a PSNR of 27.33dB and an SSIM value of 0.912 in the FoggyHaze dataset, far exceeding other methods. In particular, in terms of PSNR, CT-Former is 6.08dB and 3.91dB higher than FFA-Net and TransWeather, respectively, fully demonstrating its significant advantages in restoring image details and improving image quality. At the same time, the NIQE score of CT-Former is 3.47, close to 3.48 of USC-Former, indicating that while ensuring the dehazing effect, CT-Former can also maintain a good image naturalness. In the Dense Haze dataset, CT-Former also performed well, once again leading all comparison methods with a PSNR of 19.93dB and an SSIM value of 0.804, especially showing a strong advantage in structural similarity. This performance proves that CT-Former can still effectively remove haze and restore image details and structures under more difficult haze conditions. Although CT-Former is slightly lower than USC-Former in the NIQE indicator, its score of 3.56 still shows a high image naturalness. These results show that CT-Former not only performs well under lighter haze conditions, but also has excellent dehazing capabilities and image quality retention capabilities in more complex Dense Haze environments, demonstrating its potential for wide application in various outdoor haze scenes.

[0177] The optimized model is exported to a lightweight format (such as ONNX or TensorRT) for deployment on the onboard computing platform of the autonomous vehicle. The model is optimized by pruning and quantization before deployment to ensure that its running speed meets the real-time requirements. In actual operation, the system monitors the defogging effect and performance of the model, and further optimizes the defogging method by collecting new data and updating model parameters.

[0178] Table 1: Results of this method and other methods on the indoor dataset IHAZE.

[0179]

[0180] Table 2: Results of this method and other methods on the outdoor datasets Foggy Haze and Dense Haze.

[0181]

Claims

1. A method for defogging an autonomous driving image based on CTFormer, characterized in that: A model combining convolutional feature extraction and improved Transformer architecture is adopted, denoted as CTFormer. Through multi-scale feature modeling, Taylor unrolled attention mechanism (T-MHSA), contrast constraint module (CCM) and detail enhancement module, efficient and accurate dehazed image generation is achieved, significantly improving the quality and processing efficiency of dehazed images. The CTFormer model includes an input preprocessing module, an encoder module, a decoder module, a contrast constraint module (CCM), a loss function and optimization module, and an output module; wherein: Input preprocessing module: normalizes and enhances the input fogged image, and uses learnable position encoding to enhance the spatial information expression capability of the input features, so that the model can adapt to different weather and environmental conditions; Encoder module: The improved Transformer architecture is used to optimize the traditional multi-head self-attention mechanism into a high-efficiency attention mechanism based on Taylor expansion (T-MHSA), which is combined with local convolution operations to capture the multi-scale global features of the image and enhance the model's ability to represent the fogged area. Decoder module: gradually restores the high resolution and detail information of the image through multi-scale upsampling and cross-layer skip connections; combined with the detail enhancement module, it further improves the clarity and edge performance of the image; Contrastive Constraint Module (CCM): It uses contrastive learning strategy to construct the feature distribution of anchor points, positive samples and negative samples, optimizes the relative relationship between features through contrast loss, and improves the discriminative ability of dehazing features; Loss function and optimization module: Design a variety of loss functions, including pixel-level loss, perceptual loss, and contrast loss, and optimize the model training process through dynamic weight combination to ensure the quality and consistency of the final output image; The output module maps the feature map generated by the decoder into a high-definition dehazed image, and combines post-processing techniques such as edge smoothing and detail enhancement to output high-quality dehazing results that meet the needs of autonomous driving.

2. The automatic driving image defogging method according to claim 1, characterized in that: The specific steps are as follows: (1) The preprocessing module preprocesses the foggy image; Specifically, the foggy images collected in the autonomous driving scenario are preprocessed to convert the original images into standardized input tensors that can be processed by the model; (2) The encoder module encodes the preprocessed image, that is, the preprocessed foggy image is sent as the input tensor to the encoder module, and multi-scale features are extracted through the hybrid attention mechanism (T-MHSA) and local convolution; the feature map output by the encoder is used as the input of the subsequent module to provide multi-scale and hierarchical feature representation; (3) The decoder module restores the multi-scale features output by the encoder into a high-quality clear image, and gradually restores the image resolution and details through multi-layer upsampling and detail enhancement modules; (4) The contrast constraint module (CCM) optimizes the distinguishing ability of features and enhances the comprehensive performance of the encoder and decoder through contrastive learning strategy; (5) Loss function and optimization processing during model training. Specifically, the quality of dehazed images and the robustness of the model are improved through the joint design and optimization of multiple loss functions. The loss functions mainly include the optimization of pixel-level loss, perceptual loss, contrast loss, and total loss. (6) The output module transforms the feature map generated by the decoder into the final dehazed image through mapping and post-processing to ensure the visual effect and practicality of the output.

3. The automatic driving image defogging method according to claim 2, characterized in that: The preprocessing module in step (1) preprocesses the foggy image, specifically including: Normalization operation: Convert the pixel value I(x,y) of the input image to the interval [0,1] to eliminate the brightness difference between different input images. The formula is: Among them, min(I) and max(I) are the minimum and maximum values ​​of image pixels, respectively. ′ (x, y) is the normalized pixel value; Data augmentation: Expanding the training dataset by: Random cropping: crop random areas of the image to retain the main information of the image; Flip operation: randomly flip the image horizontally or vertically; Add noise: introduce Gaussian noise or salt and pepper noise to simulate sensor noise in real environment; Lighting adjustment: randomly adjust image brightness, contrast, and saturation to improve the lighting robustness of the model; Learnable position encoding: By adding a learnable position encoding matrix PE, the model can perceive the spatial position information of the image. The calculation formula of the position encoding is: Among them, i represents the position, j represents the dimension index, and d is the encoding dimension; the position encoding matrix will be added to the input image tensor element by element to enhance the model's understanding of spatial relationships.

4. The method for defogging an image for autonomous driving according to claim 3, characterized in that: In step (2), the encoder module extracts multi-scale features of the input image and models global dependencies; combined with the multi-head self-attention mechanism based on Taylor expansion (T-MHSA) and local convolution feature extraction technology, the specific calculation process is: (2.1) Input feature normalization and block operation; Layer Normalization: Normalize the input tensor Perform layer normalization to improve the stability of features. The calculation formula is as follows: Among them, μ is the characteristic mean, σ 2 is the characteristic variance, ∈ is the stability constant, T ′ is the normalized feature tensor; Block operation: normalize the input image features T ′ Divide the window size into N non-overlapping sub-blocks according to P×P The size of each sub-block is P×P×C; (2.2) Global feature modeling based on T-MHSA; Query, key and value generation: For each sub-block B i Generate the query matrix Q, key matrix K, and value matrix V: Q=B i W Q ,K=B i W K ,V=B i W V , (5) in, is a learnable weight matrix, d represents the dimensions of query, key, and value; Attention weight calculation: Taylor expansion is used to approximate the Softmax function to reduce the computational complexity. The attention score calculation formula is as follows: Through Taylor expansion, the Softmax function is approximated as: The calculation formula for attention output OOO is: O=A(Q,K)V, (8) Multi-head attention mechanism: The attention outputs of multiple heads are concatenated and linearly transformed. The calculation formula is: O multi-head =Concat(O1,O2,…,O h )W O , (9) in, is the linear transformation matrix, h is the number of attention heads; (2.3) Local feature extraction enhanced by hybrid convolution; Local convolution operation: for each sub-block B i A convolution operation is performed to extract fine-grained features. The convolution kernel size is 3×3 and the calculation formula is: F conv =Conv2D(B i ,W conv ), (10) Among them, W conv is the convolution kernel parameter; Fusion of local and global features: The local features F extracted by convolution are conv With the global feature O multi-head After addition, the layers are normalized: F fusion =LayerNorm(F conv +O multi-head ), (11) (2.4) Output of encoder; Multi-scale feature representation: The encoder aggregates the processed features through multi-scale convolution to generate feature maps of different resolutions: F multi-scale =Concat(F scale1 ,F scale2 ,…,F scaleN ), (12) Layer stacking: Multiple encoding layers are stacked to further enhance the feature representation capability, and the output feature map F encoder It will be used as input to the decoder.

5. The automatic driving image defogging method according to claim 4, characterized in that: The specific process of step (3) is as follows: (3.1) Cross-layer skip connection; Feature fusion of skip connections: extracting multi-scale features from each layer of the encoder And pass it to the corresponding layer of the decoder through the skip connection: in, is the output of the previous layer of the decoder, is the fused feature; Feature transformation: fusion of features Transformed by 3×3 convolution to enhance the expressiveness of features while reducing redundant information: Among them, W conv is the convolution kernel parameter; (3.2) Multi-scale upsampling Upsampling operation: The decoder uses layer-by-layer upsampling to restore the low-resolution feature map to high resolution, and uses the nearest neighbor interpolation method to achieve upsampling. The formula is: Multi-scale feature fusion: During the upsampling process, a multi-scale fusion module is introduced to combine feature maps of different resolutions to improve detail recovery capabilities: Among them, w i is the weight of each scale, and N represents the number of scales; (3.3) Detail enhancement High-frequency information extraction: Use the refined convolution operation to extract high-frequency information to enhance the clarity of details and edges: F detail =Conv2D(F multi-scale ,W detail ), (17) Among them, W detail To refine the convolution weights; Residual enhancement: Combine the original features F output by the encoder original And the output of the current decoding layer to form a residual connection: (3.4) Decoder output Restore the original resolution: The last layer of the decoder uses a fully connected layer to map the features to the original image size H×W and output a clear image: I clear =Conv2D(F enhanced ,W output ), (19) Dehazing result output: The image pixel values ​​are limited to the range of [0,1] through a nonlinear activation function to generate a dehazed high-definition image.

6. The method for defogging an image for autonomous driving according to claim 5, characterized in that: The specific process of step (4) is as follows: (4.1) Construction of anchor samples and positive and negative samples; Anchor sample definition: Select the global feature f of the current input image from the output of the encoder a As an anchor feature; Positive sample definition: Generate an enhanced version of the same input image by random perturbation and extract its features f from the corresponding encoder output p , as a positive sample; Negative sample definition: randomly select features f from the encoder output of other input images n , as negative samples; (4.2) Contrastive learning optimization; Contrastive loss design: The contrastive constraint module uses the contrastive loss function to optimize the feature distribution between anchor points, positive samples and negative samples, so as to minimize the distance between the anchor points and positive samples and maximize the distance between the anchor points and negative samples. The formula is: Among them, m is the preset boundary value, which controls the feature difference between positive and negative samples; Feature regularization: In the contrastive learning process, the CCM module regularizes the anchor feature f a Add regularization constraints to prevent the model from overfitting to certain local features: Joint loss optimization: The contrast loss and the dehazing loss are weighted together to form a total loss function: Among them, λ1 and λ2 are weight coefficients, which control the influence of different loss terms; (4.3) Feature distance calculation and update; Feature distance calculation: In the feature space, the Euclidean distance is used to calculate the distance between the anchor point and the positive sample and the negative sample: d pos =‖f a -f p ‖,d neg =‖f a -f n ‖, (23) Gradient update: The contrastive constraint module optimizes model parameters through back propagation so that the feature distance meets the contrastive learning objective: Target:d pos →0,d neg →∞ (4.4) Enhanced diversity in contrastive learning; Multi-view positive sample generation: Use data enhancement to generate multiple positive samples to enhance the model's adaptability to different scenarios; Dynamic negative sample sampling: select samples that are close to the anchor point features as negative samples to improve the optimization effect of contrast loss.

7. The method for defogging an image for autonomous driving according to claim 6, characterized in that: The specific process of step (5) is as follows: (5.1) Pixel-level loss; L1 loss: directly calculates the absolute difference between the dehazed image and the clear target image at the pixel level to maintain pixel consistency; Structural Similarity Loss (SSIM): measures the similarity between the dehazed image and the target image in terms of brightness, contrast, and structure to improve the consistency of visual effects; (5.2) Perceptual loss; Feature perception: Use pre-trained deep convolutional neural networks to extract high-level features of the dehazed image and the target image, use perceptual loss to measure the feature difference between the two, and emphasize the semantic consistency of the image; Multi-scale perception: Calculates perceptual loss at multiple scales to further improve the clarity and detail of dehazed images; (5.3) contrast loss; Relative distance between features: The contrastive learning module (CCM) is used to optimize the feature distance between anchor points, positive samples, and negative samples, thus enhancing the model’s ability to distinguish fogged images. (5.4) Joint optimization of total loss; Weighted combination: The pixel-level loss, perceptual loss, and contrast loss are weighted together to form a total loss function: Among them, λ1, λ2, λ3 are the weight coefficients of each loss term, which control their influence on model optimization; Dynamic weight adjustment: Dynamically adjust the weight of loss items according to the training stage and the output quality of the dehazed image to ensure balanced contribution of each loss item at different stages; (5.5) Optimization strategy Optimizer selection: Adopt Adam optimizer, combined with momentum term and learning rate decay mechanism to accelerate model convergence; Learning rate scheduling: Dynamically adjust the learning rate during the training process. The initial learning rate is high and gradually decreases with training iterations to improve the model's ability to optimize details in the later stages.

8. The method for defogging an image for autonomous driving according to claim 7, characterized in that: The specific process of step (6) is as follows: (6.1) Feature mapping; Linear transformation: Use a fully convolutional layer to map the feature map output by the decoder to a clear image consistent with the resolution of the original image: I clear =Conv2D(F decoder ,W output ), (25) Among them, F decoder is the output feature map of the last layer of the decoder, W output is the weight of the mapping convolution kernel; Non-linear activation: Applying an activation function to the output image restricts the pixel value range to [0,1] to enhance the visual consistency and contrast of the image; (6.2) Artifact removal processing; Edge smoothing: Smoothing the edge areas of the generated image to reduce possible artifacts, achieved through a Gaussian filter: I smooth = GaussianFilter (I clear ), (26) Detail enhancement: Use Laplace filter to enhance the high-frequency details of the image and optimize the detail performance; (6.3)Multiple output level support; Low-resolution preview: To meet real-time requirements, a low-resolution clear image preview is output for quick use; High-resolution main output: Provides a complete high-resolution dehazed image for subsequent target detection and path planning in autonomous driving tasks; (6.4) Quality assessment feedback; Image quality evaluation: Evaluate the quality of dehazed images based on SSIM and PSNR indicators, and generate an evaluation report to optimize subsequent model training; Feedback mechanism: Automatically adjust model parameters based on evaluation results to improve training efficiency.

Citation Information

Patent Citations

  • Image defogging method and device, electronic equipment and storage medium

    CN115908159A

  • Target tracking method based on convolution Transform combination

    CN116645625A

  • Defogging method based on foggy day traffic road image

    CN118365558A

  • Swin Transform self-adaptive image fusion method containing perception enhancement module

    CN118644401A

  • Lightweight image defogging method and system

    CN119151828A

Cited By

  • Target detection task-driven image defogging method

    CN120339120A

  • An image dehazing method driven by object detection task

    CN120339120B

  • Mine dust fog image defogging method and system based on three-dimensional inverse interactive intersection

    CN120430984A

  • Real scene image defogging method based on disturbance defense and semantic guidance

    CN120580173A

  • Infrared-guided image restoration method under interference of non-uniform scattering medium

    CN120765508A