Image deblurring method based on frequency domain feature attention mechanism

By introducing the frequency domain feature attention mechanism in the image defuzzy network and adaptively selecting frequency information, the problem that the existing technology is difficult to effectively model the high-frequency features of images is solved, and a better image defuzzy effect is achieved.

CN119941571APending Publication Date: 2025-05-06CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510041463.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing image defuzzing networks are difficult to effectively model the high-frequency features of images, resulting in poor defuzzing effects.

Method used

An image defuzzing method based on frequency domain feature attention mechanism is proposed. The FDAM module adaptively determines which frequency information should be retained, thereby enhancing the global feature extraction capability and capturing of image edge parts.

Benefits of technology

This method not only reduces the computational complexity, but also improves the image debuffering effect, especially in the recovery of the edge parts of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941571A_ABST
    Figure CN119941571A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to an image deblurring method based on a frequency domain feature attention mechanism, which comprises the following steps: acquiring an image deblurring data set, and dividing the data set into a training set and a test set; inputting the data in the training set into the image deblurring model for training, and testing the trained image deblurring model by using the data in the test set; performing deblurring processing on a to-be-processed image by adopting the trained image deblurring model to obtain a deblurred image; wherein the image deblurring model is composed of three encoders and two decoders; according to the method, the calculation complexity is reduced, and the global feature extraction capability and the capture of the image edge part are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and in particular relates to an image deblurring method based on a frequency domain feature attention mechanism. Background Art

[0002] Image deblurring is a classic problem in low-level computer vision, which aims to restore a clear image from a blurred input image. There are many factors that may cause image blur, such as camera shake during shooting, object motion, and optical defocus. Image deblurring is a key problem in the underlying visual task. If it can be well solved, it will greatly promote the research of other visual tasks, such as semantic segmentation, object detection, etc.

[0003] Thanks to the rapid development of deep learning, more and more scholars have begun to apply deep learning technology to image deblurring tasks and have proposed a large number of deblurring networks. Among various deblurring networks, networks that use attention mechanisms (such as Restormer and Uformer) can often better restore blurred images. Uformer is a Transformer based on UNet, which uses self-attention based on non-overlapping windows to deblur a single image. Although the use of window segmentation strategy reduces the computational cost, coarse segmentation cannot fully explore the information of each patch and cannot model the long-range information of the image. Restormer greatly reduces the computational complexity of self-attention by calculating the channel attention of features, but this global attention of the feature map channel can only model the low-frequency features of the image, and cannot model the high-frequency features of the image well. Summary of the invention

[0004] In order to solve the problems existing in the above-mentioned prior art, the present invention proposes an image deblurring method based on frequency domain feature attention mechanism, the method comprising: obtaining an image deblurring dataset, dividing the dataset into a training set and a test set; inputting the data in the training set into an image deblurring model for training, and using the data in the test set to test the trained image deblurring model; using the trained image deblurring model to deblur the image to be processed to obtain a deblurred image; wherein the image deblurring model is composed of three encoders and two decoders.

[0005] Beneficial effects of the present invention:

[0006] Since not all low-frequency information and high-frequency information are helpful for potential clear image restoration, the present invention develops an FDAM module that can adaptively decide which frequency information should be retained. Since each frequency component represents the global information of the feature, this module not only reduces the computational complexity compared to ordinary attention operations, but also enhances the global feature extraction capability and the capture of image edge parts. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 It is the overall flow chart of the present invention;

[0008] Figure 2 Schematic diagram of the image deblurring model based on the frequency domain feature attention mechanism proposed in the present invention;

[0009] Figure 3 A schematic diagram of the encoder and decoder modules proposed in the present invention;

[0010] Figure 4 Schematic diagram of the frequency domain feature attention module proposed in the present invention;

[0011] Figure 5 This is a schematic diagram of the GEGLU activation function proposed in the present invention;

[0012] Figure 6 An example diagram of using the GoPro dataset. DETAILED DESCRIPTION

[0013] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0014] An image deblurring method based on frequency domain feature attention mechanism, such as Figure 1 As shown, the method includes: obtaining an image deblurring data set, dividing the data set into a training set and a test set; inputting the data in the training set into an image deblurring model for training, and using the data in the test set to test the trained image deblurring model; using the trained image deblurring model to deblur the image to be processed to obtain a deblurred image; wherein the image deblurring model is composed of three encoders and two decoders.

[0015] In this embodiment, training the image deblurring model includes:

[0016] Step 1: preprocess the data in the training set and use the preprocessed image as the first image; sample the first image to obtain a second image of 1 / 4 size and a third image of 1 / 16 size;

[0017] Step 2: Input the first image into a first encoder for feature extraction and downsampling operations to obtain a first feature map;

[0018] Step 3: splice the first feature map with the second image, input the spliced ​​feature map into the convolution layer to extract common abstract features; input the common abstract features into the second encoder for downsampling operation to obtain a second feature map;

[0019] Step 4: splice the second feature map with the second image, input the spliced ​​feature map into the convolution layer to extract common abstract features; input the common abstract features into the third encoder to obtain a third feature map;

[0020] Step 5, upsampling the first feature map and the second feature map, and upsampling the third feature map twice; performing 1×1 convolution on the upsampled first feature map, the second feature map, and the third feature map, and then adding the pixels, and inputting the added feature map into the first decoder; adding the image output by the first decoder to the blurred image pixel by pixel to obtain a first clear image;

[0021] Step 6, down-sample the first feature map, and up-sample the second feature map and the third feature map; perform 1×1 convolution on the down-sampled first feature map and the down-sampled second feature map and the third feature map, and then perform a pixel-by-pixel addition operation, and input the added feature map into the second encoder to obtain a second clear image;

[0022] Step 7: Calculate the loss function of the model based on the input image, the first clear image, and the second clear image, adjust the model parameters, and complete the model training when the loss function converges.

[0023] In this embodiment, the GoPro dataset is selected as the image deblurring dataset, and the dataset is divided into a training set and a test set in a ratio of 7:3. The test set is not changed, and the training set is subjected to data preprocessing: the image is cut into 8 384×384 images at equal intervals and randomly flipped horizontally up and down to obtain the final training set.

[0024] Normalize the blurred images in the training set: normalize the pixel values ​​0-255 to the range 0-1 and input them into the frequency domain feature attention network, such as Figure 2 . By calculating the difference between the network output and the clear image, the loss is calculated and the gradient is returned to update the parameters in the frequency domain feature attention network. The loss function L is set to L1 Loss + Fre Loss. Let the network output be Y and the clear image be X. Since the network proposed in the present invention is a multi-input multi-output network, the network has 3 outputs Y1, Y2, Y3 and 3 clear images of different sizes X1, X2, X3.

[0025] L1 Loss:

[0026]

[0027] Frequency Loss:

[0028]

[0029] The final loss function is:

[0030] L total =L1+α×Fre

[0031] Among them, L1 is the L1 distance between the deblurred image and the corresponding clear image, also known as the Manhattan distance. Constructing the L1 loss function requires three pairs of deblurred-clear image pairs, and finally taking the average. α is a hyperparameter in the network, which is used to adjust the proportion of Fre frequency domain loss in the entire loss function. Y i represents the output image of the network, i.e., the deblurred image. Since the network is a multi-scale network, there are output images of three scales; X i represents the corresponding clear image. Function F(·) represents the fast Fourier transform (FFT), F(Y i ), F(X i ) means transforming the output image and the corresponding clear image into the frequency domain.

[0032] Preferably, the hyperparameter α is set to 0.01.

[0033] In this embodiment, a multi-input multi-output U-net network is used. The network is composed of encoders and decoders at different levels, and the internal structures of the encoders and decoders are the same. The information exchange between the feature map encoder and the decoder has been clearly introduced in the technical solution and the main points of the invention. Now, the flow process of the feature map in the encoder is explained.

[0034] The encoder consists of a layer normalization module, a convolution layer with a convolution kernel size of 1*1, a depth-separable convolution layer with a convolution kernel size of 3*3, a frequency domain attention module, a convolution layer with a convolution kernel size of 1*1, a layer normalization module, a convolution layer with a convolution kernel size of 1*1, a GEGLU module, and a convolution layer with a convolution kernel size of 3*3; each module is connected in sequence. When the image feature dimension input to the encoder is C×H×W, the feature map is processed serially in the order of each module.

[0035] The frequency domain feature attention module includes an 8×8 matrix W with all initialization elements set to 1. The parameters are optimized and updated as the network is trained. A rearrange function is used to divide the feature map into multiple 8×8 image blocks. The process of the feature map passing through the frequency domain feature attention module is as follows: first, the rearrange function is applied to the feature map to divide it into multiple 8×8 image blocks, then each image block is multiplied by the corresponding elements of the matrix W, and finally, the rearrange function is used to merge the multiple 8×8 image blocks into the original size.

[0036] The GEGLU module includes a GELU function: GELU(x) = x*Φ(x), where Φ(x) is the cumulative probability distribution of the Gaussian distribution. The GEGLU module first divides the feature map into two equal parts according to the channel dimension, denoted as x1 and x2, applies the GELU function to the x1 feature map, performs nonlinear activation on the feature map, and then multiplies the activated x1 and x2 accordingly to screen important features.

[0037] like Figure 3 As shown, the image in the 0-1 interval is input into the first encoder, denoted as x_input. Assuming the size of the feature map is C×H×W, it is first normalized through the LayerNorm layer, and the size of the feature map remains unchanged, denoted as x1. Then x1 is input into the 1×1, conv layer for dimensionality increase operation, and the convolution kernel size is 1×1, which doubles the channel dimension of the feature map to obtain a feature map x2 with a size of 2C×H×W. Then x2 is input into the 3×3, dconv layer for depth-separable convolution operation to obtain a feature map x3. The size of x3 is 2C×H×W, and the convolution kernel size is 3×3. The use of depth-separable convolution can improve computational efficiency. Then x3 is input into the frequency domain feature attention module proposed in the present invention.

[0038] x1=LN(x_input)

[0039] x2=conv(x1)

[0040] x3=dconv(x2)

[0041] In this embodiment, if Figure 4As shown, the frequency domain feature attention module includes a feature map fast Fourier forward and inverse transform and two rearrangement operations. First, a quantization matrix W is generated according to the size of x3, the size is 2C×1×1×patch_size×patch_size, and the initial element value is set to 1. The patch_size of the present invention is set to 8, and different values ​​can be set according to different data sets. The feature map rearrangement operation is now explained: x3 is divided into feature maps of a size of patch_size×patch_size. Taking a feature map of size H×W and a channel number of 1 as an example, the size of the divided feature map is H / 8×W / 8×8×8. The 2C feature maps are divided in turn, and the size of the feature map x4 is 2C×H / 8×W / 8×8×8. Then, x4 is subjected to FFT operation to obtain each frequency component, and the obtained frequency component map is recorded as x5. At this time, the size of x5 is 2C×H / 8×W / 8×8×8 and each element value is a complex number. Next, the quantization matrix W is multiplied with x5, that is, the elements in the same position are multiplied. The multiplication operation means the degree of attention given to different frequency components. After the dot multiplication operation is completed, the IFFT operation is performed to convert x5 from the complex domain to the real domain. Under the FFT-dot multiplication-IFFT operation process, the network completes the frequency domain attention operation. Since the initial element value of the quantization matrix W is set to 1, the attention degree of all frequency components is the same at this time. In order to make the quantization matrix W continuously update the parameters, W needs to be registered to the image deblurring network. After registration with the nn.Parameter function, the quantization matrix W becomes part of the network, and the parameters of the quantization matrix W are updated every time the loss is calculated and the gradient is returned. Suppose feature map x5 passes through the frequency domain feature attention module to obtain feature map x6, and then input x6 into the 1×1, conv layer for dimensionality reduction, with a convolution kernel size of 1×1, to obtain feature map x7 of size C×H×W. In order to process the feature map more finely, it is necessary to perform nonlinear activation function processing on the feature map. The operation of adding the corresponding elements of x_input and x7 is used as the input of the next LayerNorm layer, and y1 is recorded as x_input+x7. y1 is input into the LayerNorm layer for normalization to obtain y2. Similarly, y2 is passed through the 1×1, conv layer for dimensionality increase, with a size of 2C×H×W, and then passes through the GEGLU activation function.

[0042] x4=rearrange(x3)

[0043] x5=FFT(x4)

[0044] x6=IFFT(W⊙x5)

[0045] x7=conv(x6)

[0046] y1=x_input+x7

[0047] y2=LN(y1)

[0048] Among them, IFFT is the inverse Fourier transform, FFT is the Fourier transform, rearrange is a rearrangement function, which divides the feature map into multiple 8×8 image blocks for subsequent processing. x3 is the feature map after depth-separable convolution, x4 is the feature map after block division, x6 is the feature map obtained after the frequency domain feature attention module, W is the attention coefficient matrix, ⊙ is the point multiplication operation of the corresponding elements, x7 is the feature of the feature map x6 after feature extraction, x_input is the input feature of the encoder, y1 is the feature obtained after a round of residual learning, and y1 is processed by layer normalization LN to obtain y2 as the input of the subsequent module.

[0049] The GEGLU module proposed in the present invention is as follows Figure 5 As shown, the specific formula is:

[0050] GEGLU(x)=GELU(xW)×(xV)

[0051] Among them, GELU stands for the GELU activation function, and the specific formula is: GELU(x) = x*Φ(x), where Φ(x) is the cumulative probability distribution of the Gaussian distribution. To simplify the explanation, the formulas here do not include bias terms. GEGLU controls the output through a gating mechanism, which can be seen as the selection of important features like Attention. Its advantage is that it not only has the nonlinearity of a general activation function, but also has a linear channel when backpropagating the gradient, similar to the addition operation in the ResNet residual network to transfer the gradient, which can alleviate the gradient disappearance problem.

[0052] In order to achieve the effect of GEGLU, the present invention sets a 1×1,conv layer after the LayerNorm layer, changes the channel of y2 from C to 2C, and then divides the feature map into two according to the number of channels, set to z1 and z2, as shown in Figure 5 As shown. At this time, z1=xW, z2=xV, that is, the weights of W and V are realized through a convolutional layer. Then z1 is applied to the GELU function to realize the GELU(xW) part, and then z1 and z2 are multiplied by points, which is regarded as the selection of important features. Finally, the result is passed through a 3×3, conv layer for feature extraction. Finally, the final output of the encoder is obtained by adding y1 through a jump link. So far, the internal structure of the encoder has been introduced. By constructing two residual blocks in an encoder, the network can learn features and representations more deeply and accurately. The jump connection allows the network to directly transmit information, avoiding the feature degradation problem of superimposing multiple encoders and decoders. The output of the encoder is represented by x_output.

[0053] concat(z1,z2)=conv(y2)

[0054] GEGLU(conv(y2))=GELU(z1)⊙z2

[0055] x_output=conv(GEGLU(conv(y2)))+y1

[0056] Among them, the concat function is used to concatenate the feature maps z1 and z2 according to the channel dimension, y2 is the feature map after the previous layer normalization LN processing, GEGLU represents the GEGLU activation function module, GELU represents the activation function, and x_output represents the final output of the encoder. The final image deblurring result is as follows Figure 6 shown.

[0057] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation modes of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. An image deblurring method based on frequency domain feature attention mechanism, characterized in that: include: Obtain an image deblurring dataset and divide the dataset into a training set and a test set; The data in the training set are input into the image deblurring model for training, and the trained image deblurring model is tested using the data in the test set; the image to be processed is deblurred using the trained image deblurring model to obtain a deblurred image; wherein the image deblurring model is composed of three encoders and two decoders.

2. The image deblurring method based on frequency domain feature attention mechanism according to claim 1, characterized in that: The structure of the encoder is the same as that of the decoder.

3. The image deblurring method based on frequency domain feature attention mechanism according to claim 1, characterized in that: Training an image deblurring model involves: Step 1: preprocess the data in the training set and use the preprocessed image as the first image; sample the first image to obtain a second image of 1 / 4 size and a third image of 1 / 16 size; Step 2: Input the first image into a first encoder for feature extraction and downsampling operations to obtain a first feature map; Step 3: splice the first feature map with the second image, input the spliced ​​feature map into the convolution layer to extract common abstract features; input the common abstract features into the second encoder for downsampling operation to obtain a second feature map; Step 4: splice the second feature map with the second image, input the spliced ​​feature map into the convolution layer to extract common abstract features; input the common abstract features into the third encoder to obtain a third feature map; Step 5, upsampling the first feature map and the second feature map, and upsampling the third feature map twice; performing 1×1 convolution on the upsampled first feature map, the second feature map, and the third feature map, and then adding the pixels, and inputting the added feature map into the first decoder; adding the image output by the first decoder to the blurred image pixel by pixel to obtain a first clear image; Step 6, down-sample the first feature map, and up-sample the second feature map and the third feature map; perform 1×1 convolution on the down-sampled first feature map and the down-sampled second feature map and the third feature map, and then perform a pixel-by-pixel addition operation, and input the added feature map into the second encoder to obtain a second clear image; Step 7: Calculate the loss function of the model based on the input image, the first clear image, and the second clear image, adjust the model parameters, and complete the model training when the loss function converges.

4. The image deblurring method based on frequency domain feature attention mechanism according to claim 3, characterized in that: The preprocessing of the data in the training set includes: cutting the image into 8 images of 384×384 at equal distances and randomly flipping them horizontally up and down; normalizing the flipped image to obtain the first image.

5. The image deblurring method based on frequency domain feature attention mechanism according to claim 3, characterized in that: The encoder consists of a layer normalization module, a convolution layer with a convolution kernel size of 1*1, a depth-wise separable convolution layer with a convolution kernel size of 3*3, a frequency domain attention module, a convolution layer with a convolution kernel size of 1*1, a layer normalization module, a convolution layer with a convolution kernel size of 1*1, a GEGLU module, and a convolution layer with a convolution kernel size of 3*3; each module is connected in sequence.

6. The image deblurring method based on frequency domain feature attention mechanism according to claim 5, characterized in that: The frequency domain feature attention module includes an 8×8 matrix W with all initialization elements set to 1 and a rearrangement function rearrange; The frequency domain feature attention module processes the input feature map by: using the rearrange function to divide the input feature map into multiple 8×8 image blocks; Each image block is multiplied by the matrix W between corresponding elements; finally, multiple 8×8 image blocks are merged into the original size through the rearrange function.

7. The image deblurring method based on frequency domain feature attention mechanism according to claim 5, characterized in that: The GEGLU module processes the feature map as follows: the GEGLU module first divides the feature map into two equal parts according to the channel dimension, recorded as x1 and x2; applies the GELU function to the x1 feature map, performs nonlinear activation on the feature map, and then multiplies the activated x1 and x2 accordingly to screen important features.

8. The image deblurring method based on frequency domain feature attention mechanism according to claim 3, characterized in that: The loss function of the model is: L total =L1+α×Fre Among them, L1 is the L1 distance between the deblurred image and the corresponding clear image, α is the hyperparameter in the network, Fre is the frequency domain loss, and Y i is the output image of the network, X i is the corresponding clear image, F(·) is the fast Fourier transform, F(Y i ) is the output image transformed into the frequency domain, F(X i ) is the corresponding clear image transformed into the frequency domain.

Citation Information

Cited By

  • Machine tool spindle thermal error prediction method and system based on regional division thermal image

    CN120802832A