An Image Dehazing Method Based on Frequency Domain Attention Mechanism of Discrete Cosine Transform

By adopting a frequency domain attention mechanism based on discrete cosine variation in image defog technology, the basic network module for fusion of DCT and MLP is designed, and the problem of poor image defog removal effect in the prior art is solved, and better image detail information extraction and defog removal effect is achieved.

CN115953305BActive Publication Date: 2025-05-30NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211223828.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-09
Publication Date
2025-05-30
Estimated Expiration
2042-10-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively remove haze in images, affecting image quality and performance of computer vision tasks.

Method used

Using a frequency domain attention mechanism based on discrete cosine variation, a basic network module integrated with DCT and MLP is designed to extract better image detail information by modeling images in the frequency domain.

Benefits of technology

Modeling images in the frequency domain can obtain better image detail information, significantly improving the image defog effect, and better than existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115953305B_ABST
    Figure CN115953305B_ABST
Patent Text Reader

Abstract

The present invention provides an image dehazing method based on a frequency-domain attention mechanism using discrete cosine transform, comprising the following steps: S1, constructing an attention module based on the frequency domain; S2, constructing a basic network module; S3, constructing a dehazing network model; S4, designing a loss function; S5, training the dehazing network model using hazy images and haze-free images to obtain the model parameters of the dehazing network model; S6, importing the model parameters trained in step S3 into the dehazing network model, inputting a hazy image, and outputting a haze-free image. A basic network module integrating discrete cosine transform (DCT) and multi-layer perceptron (MLP) is designed to model the image in the frequency domain, obtaining better image detail information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology of image defogging, and particularly to an image defogging method based on a frequency-domain attention mechanism of discrete cosine transform. Background Art

[0002] Dust and smoke floating particles in the air will absorb and scatter sunlight, resulting in a decline in the image quality in video surveillance and affecting the performance of computer vision tasks such as target detection and image segmentation. Therefore, image defogging is a key image preprocessing technology, which can reduce the interference of haze weather on image quality by restoring the image quality.

[0003] With the great success of deep learning in the field of computer vision, image defogging based on deep learning has gradually become the mainstream, but most methods use convolutional neural networks. The convolutional neural network extracts image features step by step from the underlying edge features through a multi-layer structure, which can be regarded as an implicit modeling of the frequency domain. The discrete cosine transform (DCT) has been widely used in the field of image compression by converting the spatial domain signal to the frequency domain.

[0004] Therefore, the present patent invents an image defogging method based on a frequency-domain attention mechanism of discrete cosine transform, which explicitly models the image in the frequency domain and obtains better image detail information. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide an image defogging method based on a frequency-domain attention mechanism of discrete cosine transform, and a basic network module integrating discrete cosine transform (DCT) and multi-layer perceptron (MLP) is designed, so that the image is modeled in the frequency domain and better image detail information is obtained.

[0006] To solve the above technical problem, an embodiment of the present invention provides an image defogging method based on a frequency-domain attention mechanism of discrete cosine transform, including the following steps:

[0007] S1. Construct an attention module based on the frequency domain;

[0008] S2. Construct a basic network module: composed of several residual modules, an attention module based on the frequency domain, and a 3×3 convolutional layer;

[0009] S3. Construct a defogging network model: composed of a 3×3 convolutional layer, a downsampling layer, several basic network modules, an upsampling layer, and a 3×3 convolutional layer;

[0010] S4. Design a loss function: composed of a smooth L 1 loss function;

[0011] S5. Use the foggy image and the fog-free image to train the defogging network model to obtain the model parameters of the defogging network model;

[0012] S6. Import the model parameters trained in step S3 into the defogging network model, input the foggy image, and output the fog-free image.

[0013] Among them, in step S1, the frequency-domain based attention module includes two 3×3 convolutional layers, two discrete cosine transform network layers DCT, and a multi-layer perceptron MLP with one hidden layer, which are fused through the Sigmoid activation function.

[0014] Among them, the discrete cosine transform network layer transforms the input feature f in , and uses the discrete cosine transform to transform f in to the frequency domain to obtain N groups of frequency signal components Multiply these N groups of frequency components to obtain the output:

[0015]

[0016] Among them, in step S2, the residual module consists of a 3×3 convolutional layer, a ReLU activation function, and a 3×3 convolutional layer.

[0017] Among them, in the defogging network model, there is a short connection between the downsampling layer and the upsampling layer, and there is a short connection between the basic network modules.

[0018] Among them, the number of the basic network modules is set to an odd number to ensure the symmetry of the network model, and introduce the features of the low-level network into the high-level network through short connections to improve the feature expression.

[0019] Among them, in step S4, the smooth L 1 loss function is used to measure the quality of the model prediction, representing the gap between the prediction and the actual data, and is defined as follows:

[0020]

[0021] Among them, N represents the number of training samples, represents the pixel value of the i-th channel of the restored image, and J i (x) represents the pixel value of the i-th channel of the ground truth image.

[0022] Among them, in step S5, the steps for training the defogging network model are:

[0023] (1) Select a sample set of foggy-fog-free image pairs, 12,000 for each, and 1,200 test images;

[0024] (2) The optimization function is Adam, with parameters β 1 = 0.9 and β 2 = 0.99, and the learning rate is set to 0.001;

[0025] (3) Train for 200 epochs, test the results every 20 epochs. The test results are measured using peak signal-to-noise ratio and structural similarity, and the best results are selected to save the model parameters.

[0026] The beneficial effects of the above technical solutions of the present invention are as follows:

[0027] The present invention proposes an image dehazing method based on a frequency-domain attention mechanism of discrete cosine transform, designs a basic network module integrating discrete cosine transform (DCT) and multi-layer perceptron (MLP) for extracting image features, and measures the quality of the test model prediction through a smooth L 1 loss function, enabling image modeling in the frequency domain and obtaining better image detail information. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is the dehazing network model diagram in the present invention;

[0029] Figure 2 It is the structural diagram of the attention module based on the frequency domain in the present invention;

[0030] Figure 3 It is the structural diagram of the residual module in the present invention;

[0031] Figure 4 It is the structural diagram of the basic network module in the present invention;

[0032] Figure 5 It is the comparison diagram of the image before and after dehazing in the first embodiment of the present invention;

[0033] Figure 6 It is the comparison diagram of the image before and after dehazing in the second embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0035] The embodiments of the present invention provide an image dehazing method based on a frequency-domain attention mechanism of discrete cosine transform, including the following steps:

[0036] S1. Construct an attention module based on the frequency domain (see Figure 2 );

[0037] S2. Construct a basic network module (see Figure 4):It consists of several residual modules, a frequency-domain-based attention module, and a 3×3 convolutional layer;

[0038] S3. Construct a defogging network model (see Figure 1 ):It is composed of a 3×3 convolutional layer, a downsampling layer, several basic network modules, an upsampling layer, and a 3×3 convolutional layer;

[0039] S4. Design a loss function: It consists of a smooth L 1 loss function;

[0040] S5. Use foggy images and fog-free images to train the defogging network model to obtain the model parameters of the defogging network model;

[0041] S6. Import the model parameters trained in step S3 into the defogging network model, input foggy images, and output fog-free images.

[0042] The frequency-domain-based attention module includes two 3×3 convolutional layers, two discrete cosine transform network layers DCT, and a multi-layer perceptron MLP with one hidden layer, which are fused through the Sigmoid activation function.

[0043] The discrete cosine transform network layer transforms the input feature f in , and uses the discrete cosine transform to transform f in to the frequency domain to obtain N groups of frequency signal components Multiply these N groups of frequency components to obtain the output:

[0044]

[0045] In step S2, the residual module consists of a 3×3 convolutional layer, a ReLU activation function, and a 3×3 convolutional layer. The structure is as Figure 3 shown.

[0046] The number of the basic network modules is set to an odd number to ensure the symmetry of the network model, and the features of the low-level network are introduced into the high-level network through short connections to improve feature expression.

[0047] The input of the basic network module is the image to be defogged. After being preprocessed by a 3×3 convolutional layer, the feature dimension is reduced by a downsampling to reduce the computational amount, then the features are extracted through several basic network modules in step 2, and finally the feature dimension is restored by an upsampling layer, and a 3×3 convolutional layer is used for post-processing to obtain the restored image.

[0048] In step S4, the smooth L 1 loss function is used to measure the quality of the model prediction and show the gap between the prediction and the actual data, and is defined as follows:

[0049]

[0050] Among them, N represents the number of training samples, represents the pixel value of the i-th channel of the restored image, and J i (x) represents the pixel value of the i-th channel of the ground truth image.

[0051] Among them, in step S5, the steps of training the defogging network model are as follows:

[0052] (1) Select a sample set of foggy-fogless image pairs, 12,000 each, and 1,200 test images;

[0053] (2) The optimization function is Adam, and the parameters β 1 = 0.9 and β 2 = 0.99, and the learning rate is set to 0.001;

[0054] (3) Train for 200 rounds, test the results every 20 rounds. The test results are measured by peak signal-to-noise ratio (PSNR) and structural similarity (SSIM), and select the best results to save the model parameters.

[0055] To verify the innovation of the present invention, tests are carried out. The test results are: measured by peak signal-to-noise ratio (PSNR) and structural similarity (SSIM), on the indoor test set of RESIDE, PSNR = 36.62, SSIM = 0.9901, and on the outdoor test, PSNR = 34.71, SSIM = 0.9904. The results are better than the existing methods. The visualization effect after defogging is as shown in Figure 5 、 Figure 6 shown.

[0056] Among them, Figure 5 a and Figure 6 a are the images before defogging, Figure 5 b and Figure 6 b are the images after defogging.

[0057] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An image dehazing method based on a frequency-domain attention mechanism using discrete cosine transform, characterized in that, it includes the following steps: S1. Construct a frequency-domain based attention module; S2. Construct a basic network module: composed of several residual modules, a frequency-domain based attention module, and a 3×3 convolutional layer; S3. Construct a dehazing network model: sequentially composed of a 3×3 convolutional layer, a downsampling layer, several basic network modules, an upsampling layer, and a 3×3 convolutional layer; S4. Design the loss function: It consists of a smooth L 1 loss function; S5. Use hazy images and haze-free images to train the dehazing network model to obtain the model parameters of the dehazing network model; S6. Import the model parameters trained in step S5 into the dehazing network model, input hazy images, and output haze-free images; In step S1, the frequency-domain based attention module includes two 3×3 convolutional layers, two discrete cosine transform network layers DCT, and a multi-layer perceptron MLP with one hidden layer, and is fused through a Sigmoid activation function.

2. The image dehazing method based on a frequency-domain attention mechanism using discrete cosine transform according to claim 1, characterized in that, The discrete cosine transform network layer takes the input feature f in , and uses the discrete cosine transform to transform f in into the frequency domain to obtain N groups of frequency signal components Multiply these N groups of frequency components, and the output is:

3. The image dehazing method based on a frequency-domain attention mechanism using discrete cosine transform according to claim 1, characterized in that, In step S2, the residual module is composed of a 3×3 convolutional layer, a ReLU activation function, and a 3×3 convolutional layer arranged in sequence.

4. The image dehazing method based on a frequency-domain attention mechanism using discrete cosine transform according to claim 1, characterized in that, In the dehazing network model, a short connection is made between the downsampling layer and the upsampling layer, and short connections are made between the basic network modules.

5. The image dehazing method based on a frequency-domain attention mechanism using discrete cosine transform according to claim 1, characterized in that, The number of the basic network modules is set to be odd.

6. The image dehazing method based on a frequency-domain attention mechanism using discrete cosine transform according to claim 1, characterized in that, In step S4, smooth L 1 The loss function is used to measure the quality of the model prediction, representing the degree of difference between the predicted data and the actual data, and is defined as follows: Among them, N represents the number of training samples, represents the pixel value of the i-th channel of the restored image, and J i (x) represents the pixel value of the i-th channel of the ground truth image.

7. For the image dehazing method based on a frequency-domain attention mechanism using discrete cosine transform according to claim 1, in step S5, the steps of training the dehazing network model are: (1) Select a sample set of hazy-haze-free image pairs, 12,000 for each, and 1,200 test images; (2) The optimization function is Adam, with parameters β 1 = 0.9 and β 2 = 0.99, and the learning rate is set to 0.001; (3) Train for 200 rounds, test the results once every 20 rounds, and measure the test results using peak signal-to-noise ratio and structural similarity, and select the best result to save the model parameters.