Image restoration method

By adopting U-Net architecture and transformer blocks in the image restoration network, combining compact and cross-window attention, the problems of large computing volume and insufficient global modeling capabilities are solved, and efficient image restoration effect is achieved.

CN120031759BActive Publication Date: 2025-07-08DONGHAI LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510457433.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-08
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing transformer-based image restoration method has a large amount of calculation and insufficient global modeling capabilities, resulting in poor image restoration effect.

Method used

The image restoration network using U-Net architecture and transformer blocks combines compact attention and cross-window attention to perform self-attention calculations in the channel domain and spatial domain respectively. Information is extracted through step-by-step deformation convolution and cross-window attention, reducing the calculation amount and enhancing global modeling capabilities.

Benefits of technology

It effectively reduces the amount of calculation, improves the efficiency and quality of image restoration, enhances the global modeling ability, and improves the effect of image restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031759B_ABST
    Figure CN120031759B_ABST
Patent Text Reader

Abstract

Image Restoration Method An image restoration method, which solves the problems of large computational complexity and insufficient global modeling ability of existing transformer-based image restoration methods, belongs to the field of image processing technology. The present invention includes: inputting the image to be restored into an image restoration network, and the image restoration network outputs the restored image; the image restoration network is implemented using a U-Net architecture and transformer blocks; the U-Net architecture includes four layers of encoding and four layers of decoding. Each of the first three layers of encoding and the last three layers of decoding uses a transformer block, and all use compact attention to extract channel-domain information. The fourth layer of encoding and the first layer of decoding share a transformer block, and cross-window attention is used to extract spatial-domain information. The present invention performs self-attention calculations in the channel domain and the spatial domain respectively, leveraging their respective advantages and reducing the computational complexity. The cross-window attention method makes up for the deficiency of the global modeling ability of window segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image restoration method, belonging to the technical field of image processing. Background Art

[0002] The image restoration task aims to restore clear and detailed results from degraded images caused by various environmental and device conditions. In reality, atmospheric particles such as fog and snow scatter light, resulting in phenomena such as reduced contrast, color deviation, and blurred details in the image, while imaging under low-light conditions has problems such as high noise and missing details in dark areas due to insufficient light. Currently, although the transformer-based image restoration method has achieved certain results, it still has limitations in the following three aspects: 1) The computational complexity of the transformer increases quadratically with the growth of the resolution. 2) Existing transformer methods use window partitioning to reduce the computational overhead, but this weakens the global modeling ability of the transformer. 3) The transformer usually calculates all information, which may include invalid information, and using this invalid information may have a negative effect on image restoration. Summary of the Invention

[0003] Aiming at the problems of large computational complexity and insufficient global modeling ability of the existing transformer-based image restoration method, the present invention provides an image restoration method.

[0004] An image restoration method of the present invention includes:

[0005] Input the image to be restored into the image restoration network, and the image restoration network outputs the restored image;

[0006] The image restoration network is implemented by using a U-Net architecture and transformer blocks; the U-Net architecture includes four encoding layers and four decoding layers. Each of the first three encoding layers and the last three decoding layers uses a transformer block, and the fourth encoding layer and the first decoding layer share a transformer block. Among them, the transformer blocks of the first three encoding layers and the last three decoding layers use compact attention to extract channel domain information, and the fourth encoding layer and the first decoding layer use cross-window attention to extract spatial domain information.

[0007] Preferably, the method of using compact attention to extract channel domain information is:

[0008] Use strided deformable convolution to extract and compress the information of the input image features, that is, in the form where the convolution kernel size and the stride of the strided deformable convolution are equal, to generate a position bias of size 2×H×W, where H and W represent the height and width of the image to be restored.

[0009] Preferably, the method for extracting spatial domain information using cross-window attention is as follows:

[0010] First, the input image features are evenly divided into two parts Y1 and Y2 in the channel dimension;

[0011] Y1 is evenly window-divided to obtain multiple small windows, and parallel window attention calculations are performed on each small window respectively;

[0012] Y2 is evenly window-divided to obtain multiple small windows, and the information at the corresponding positions of each small window is extracted to form a tensor T, and the token attention of each tensor T is calculated in parallel;

[0013] The calculated window attention and token attention are fused using ordinary convolution as the extracted spatial domain information.

[0014] Preferably, the image restoration network includes 3 shallow layers and 7 transformer blocks:

[0015] For the image to be restored Perform scale transformation to obtain 3 different-scale images of the image to be restored;

[0016] Denote the tensor set, the superscript represents the dimension, H is the height of the image to be restored, W is the width of the image to be restored, and C represents the number of channels;

[0017] The image to be restored is input to the first transformer block after 3×3 convolution, and the output of the first transformer block is downsampled;

[0018] For the 3 scale images sorted from high to low, the image to be restored at the first scale is input to the first shallow layer, and the first shallow layer outputs the feature image;

[0019] The downsampling result of the output of the first transformer block and the output of the first shallow layer are concatenated, and the concatenated result is input to the second transformer block after 3×3 convolution, and the output of the second transformer block is downsampled;

[0020] The image to be restored at the second scale is input to the second shallow layer, and the second shallow layer outputs the feature image;

[0021] The downsampling result of the output of the second transformer block and the output of the second shallow layer are concatenated, and the concatenated result is input to the third transformer block after 3×3 convolution, and the output of the third transformer block is downsampled;

[0022] The image to be restored at the third scale is input into the third shallow layer, and the third shallow layer outputs a feature image;

[0023] The downsampling result of the output of the third transformer block and the output of the third shallow layer are concatenated. After passing through a 3×3 convolution, the concatenated result is input into the fourth transformer block, and the output of the fourth transformer block is upsampled;

[0024] The upsampling result of the output of the fourth transformer block and the output of the third transformer block are concatenated. After passing through a 1×1 convolution, the concatenated result is input into the fifth transformer block, and the output of the fifth transformer block is upsampled;

[0025] The upsampling result of the output of the fifth transformer block and the output of the second transformer block are concatenated. After passing through a 1×1 convolution, the concatenated result is input into the sixth transformer block, and the output of the sixth transformer block is upsampled;

[0026] The upsampling result of the output of the sixth transformer block and the output of the first transformer block are concatenated. After passing through a 1×1 convolution, the concatenated result is input into the seventh transformer block;

[0027] The output of the seventh transformer block passes through a 3×3 convolution and is added element-wise to the image to be restored, and the addition result is the restored image.

[0028] Preferably, the transformer block includes a compact deformable attention module, a cross-window attention module, and a gated depth convolutional feed-forward network;

[0029] The input image features enter the compact deformable attention module or the cross-window attention module after layer normalization,

[0030] The transformer blocks in the first three layers of encoding and the last three layers of decoding use the compact deformable attention module to extract channel-domain information, and the fourth layer of encoding and the first layer of decoding use cross-window attention to extract spatial-domain information;

[0031] Element-wise addition is performed on the output of the compact deformable attention module or cross-window attention module and the input image features. The result of the addition is subjected to layer normalization and then enters the gated depth convolutional feed-forward network. Element-wise addition is performed on the output of the gated depth convolutional feed-forward network and the result of the addition, and the result of the addition is the output of the transformer block.

[0032] Preferably, the compact deformable attention module includes a 1×1 convolution, an offset generator, and two depthwise separable convolutions DWConv;

[0033] The input image features are X, The image features X are subjected to layer normalization and then enter the 1×1 convolution. The result X' after the 1×1 convolution enters the offset generator, and the offset generator outputs a position offset Vp, Vp ∈? 2×H×W The first depthwise separable convolution DWConv is set to have the same convolution kernel size and stride size. The first depthwise separable convolution DWConv selects the relevant information of the query Query and the key Key in the result X' through the position offset Vp, completes the encoding of the query Query and the key Key, and obtains the encoding matrices Q' and K'. After the encoding matrices Q' and K' are respectively reshaped, the encoding matrices Q and K are obtained. After the encoding matrices Q and K are matrix-multiplied, a softmax operation is performed to obtain where k represents the convolution kernel size of the first depthwise separable convolution DWConv;

[0034] The second depthwise separable convolution DWConv is a 3×3 convolution, selects the relevant information of the value Value in the result X', completes the encoding of the value Value, and obtains the encoding matrix V'. After the encoding matrix V' is reshaped, the encoding matrix V is obtained.

[0035] The obtained is matrix-multiplied with the encoding matrix V and then reshaped to obtain X attn , X attn is the channel domain information extracted by the compact deformable attention module.

[0036] Advantages of the present invention: The present invention performs self-attention calculations in the channel domain and the spatial domain respectively, gives full play to their respective advantages, and reduces the amount of calculation. In this network, 1) The present invention studies the strided deformable convolution module, reduces the size of the position bias generated by the deformable convolution, and improves the efficiency of the deformable convolution. 2) The present invention proposes a cross-window attention method, which makes up for the deficiency of the global modeling ability of window segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a schematic diagram of the principle of the image restoration network of the present invention;

[0038] Figure 2 It is a schematic diagram of the principle of the transformer block;

[0039] Figure 3 It is a schematic diagram of the principle of the compact deformable attention module;

[0040] Figure 4 It is a schematic diagram of the principle of the cross-window attention module;

[0041] Figure 5 It is a schematic diagram of the principle of the shallow layer;

[0042] Figure 6 It is a schematic diagram of the principle of the offset generator.

[0043] Figure 7 It is the effect diagram of the cross-window attention module;

[0044] Figure 8 It is the effect diagrams of two groups of images before and after image restoration. Detailed implementation manners

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0047] Next, the present invention will be further described in conjunction with the accompanying drawings and specific embodiments, but it is not a limitation of the present invention.

[0048] The image restoration method of this embodiment includes:

[0049] Step 1, establish an image restoration network;

[0050] The images processed in this embodiment can be 3D data It is represented that H and W represent the height and width of the image, and C represents the number of channels of the image. In a preferred embodiment, the image restoration network of this embodiment is implemented using a U-Net architecture and transformer blocks; the U-Net architecture can capture multi-scale features and contain more underlying information. Different self-attention methods (spatial domain / channel domain) are used in different transformer blocks. When the feature resolution is relatively large, this embodiment calculates attention in the channel domain, and when the resolution is relatively small, it calculates attention in the spatial domain, and the image restoration work is carried out in the form of combining the spatial domain and the channel domain.

[0051] Specifically, the U-Net architecture of this embodiment includes four layers of encoding and four layers of decoding. Each of the first three layers of encoding and the last three layers of decoding uses a transformer block, and all use compact attention to extract channel domain information. The fourth layer of encoding and the first layer of decoding share a transformer block, and cross-window attention is used to extract spatial domain information.

[0052] In this embodiment, the method of using compact attention to extract channel domain information is as follows:

[0053] The information of the input image features is extracted and compressed using strided deformable convolution. The convolution kernel size and stride of the strided deformable convolution are in an equal form to generate a position bias of size 2×H×W, where H and W represent the height and width of the image to be restored;

[0054] In the channel domain, starting from the query Query and the key Key, this embodiment uses strided deformable convolution to share the position bias between them and extract the most relevant information to achieve efficient extraction of channel domain information.

[0055] In this embodiment, the method of using cross-window attention to extract spatial domain information is as follows:

[0056] First, the input image features are evenly divided into two parts Y1 and Y2 in the channel dimension; Y1 is evenly windowed to obtain multiple small windows, and parallel window attention calculations are performed on each small window respectively. After windowing, the attention is only calculated within the small windows, lacking cross-window connections. In this embodiment, Y2 is evenly windowed to obtain multiple small windows, the information corresponding to each small window is extracted to form a new tensor T, and the token attention of each tensor is calculated in parallel; then ordinary convolution is used to fuse the calculated window attention and token attention as the extracted spatial domain information..

[0057] The cross-window attention extraction in this embodiment not only captures the relationship between pixels within a single window, but also models the cross-window dependency through the window position information to capture the relationship between pixels in different windows, making up for the lack of global modeling ability of windowing.

[0058] Specifically, as Figure 1 shown, the image restoration network of this embodiment includes 3 shallow layers and 7 Transformer blocks: the first 4 Transformer blocks are encoders, and the last 3 Transformer blocks are decoders;

[0059] The image to be restored is subjected to scale transformation to obtain 3 different scale images of the image to be restored;

[0060] represents a set of tensors, the superscript represents the dimension, H is the height of the image to be restored, W is the width of the image to be restored, and C represents the number of channels;

[0061] The image to be restored is input into the first Transformer block after a 3×3 convolution, and the output of the first Transformer block is downsampled;

[0062] For the 3 scale images sorted from high to low, the image to be restored at the first scale is input into the first shallow layer, and the first shallow layer outputs a feature image;

[0063] The downsampling result of the output of the first Transformer block and the output of the first shallow layer are concatenated, and the concatenated result is input into the second Transformer block after a 3×3 convolution, and the output of the second Transformer block is downsampled;

[0064] The image to be restored at the second scale is input into the second shallow layer, and the second shallow layer outputs a feature image;

[0065] The downsampling result of the output of the second Transformer block and the output of the second shallow layer are concatenated, and the concatenated result is input into the third Transformer block after a 3×3 convolution, and the output of the third Transformer block is downsampled;

[0066] The image to be restored at the third scale is input into the third shallow layer, and the third shallow layer outputs a feature image;

[0067] The downsampling result of the output of the third Transformer block and the output of the third shallow layer are concatenated, and the concatenated result is input into the fourth Transformer block after a 3×3 convolution, and the output of the fourth Transformer block is upsampled;

[0068] The upsampling result of the output of the 4th transformer block is concatenated with the output of the 3rd transformer block. The concatenated result is input to the 5th transformer block after a 1×1 convolution, and the output of the 5th transformer block is upsampled;

[0069] The upsampling result of the output of the 5th transformer block is concatenated with the output of the 2nd transformer block. The concatenated result is input to the 6th transformer block after a 1×1 convolution, and the output of the 6th transformer block is upsampled;

[0070] The upsampling result of the output of the 6th transformer block is concatenated with the output of the 1st transformer block. The concatenated result is input to the 7th transformer block after a 1×1 convolution;

[0071] The output of the 7th transformer block is subjected to a 3×3 convolution and then element-wise added to the image to be restored, and the added result is the restored image.

[0072] This embodiment uses a U-Net architecture with four layers for deep feature extraction. During the upsampling and downsampling processes, inverse pixel rearrangement and pixel rearrangement operations are applied to aggregate the information in the encoder and decoder, and then a 3×3 convolution is used to reduce the number of channels and output the image. This embodiment adopts a multi-input multi-output strategy, and the loss function is constrained during training to capture the multi-scale information of the image. Among them, the shallow layer is used to extract the multi-scale information of the input image. In this embodiment, the shallow layer can be implemented by sequentially connecting 3×3 convolutions, 1×1 convolutions, 3×3 convolutions, and 1×1 convolutions.

[0073] Specifically, as Figure 2 shown, the transformer block includes a compact deformable attention module, a cross-window attention module, and a gated depth convolutional feed-forward network;

[0074] The input image features enter the compact deformable attention module or the cross-window attention module after layer normalization,

[0075] The transformer blocks in the first three layers of encoding and the last three layers of decoding use the compact deformable attention module to extract channel domain information, and the fourth layer of encoding and the first layer of decoding use cross-window attention to extract spatial domain information;

[0076] Element-wise addition is performed on the output of the compact deformable attention module or cross-window attention module and the input image features. The result of the addition is normalized by layer normalization and then enters the gated depth convolutional feed-forward network. Element-wise addition is performed on the output of the gated depth convolutional feed-forward network and the result of the addition, and the result of the addition is the output of the transformer block.

[0077] In the channel domain, the compact deformable attention module designed in this embodiment has good information extraction ability. The compact deformable attention module can selectively extract the effective information of features and compress this information, providing a reliable information source for subsequent image restoration. Its structure is as Figure 3 shown. For the feature map Y ∈ R C×H×W The value Value (V) is encoded using ordinary convolution. To make the information for attention calculation as relevant as possible, strided deformable convolution is used to encode Query (Q) and Key (K) and downsample to reduce the computational load. After downsampling, Q, Through matrix multiplication QK T ∈ R C×C A common-sized Transformer attention map can be obtained, and then through: softmax(QK T ) · V ∈ R C×HW The output is obtained. Specifically, the compact deformable attention module of this embodiment includes a 1×1 convolution, an offset generator, and two depthwise separable convolutions DWConv;

[0078] The input image feature is X, The image feature X is normalized by layer normalization and then enters the 1×1 convolution. The result X' after the 1×1 convolution enters the offset generator, and the offset generator outputs the position offset Vp, Vp ∈? 2×H×W , the first depthwise separable convolution DWConv is set to have the same convolution kernel size and stride size. The first depthwise separable convolution DWConv selects the relevant information of Query and Key in the result X' by sharing the position offset Vp, completing the encoding of Query and Key, and obtaining the encoded matrices Q' and K', After reshaping the encoded matrices Q' and K' respectively, the encoded matrices Q and K are obtained, After matrix multiplication of the encoded matrices Q and K, a softmax operation is performed to obtain where k represents the convolution kernel size of the first depthwise separable convolution DWConv;

[0079] The second depthwise separable convolution DWConv is a 3×3 convolution, which selects the relevant information of Value in the result X' to complete the encoding of Value and obtain the encoded matrix V', After reshaping the encoding matrix V', the encoding matrix V is obtained.

[0080] The obtained After performing matrix multiplication with the encoding matrix V and then reshaping, X is obtained. attn , X attn is the channel domain information extracted by the compact deformable attention module. The offset generator in this embodiment includes a k×k convolution and a 1×1 convolution connected in sequence.

[0081] In the spatial domain, the cross-window attention module designed in this embodiment has good spatial information aggregation ability. The cross-window attention module considers both global and local information, enhancing the global modeling ability of the network. As Figure 7 shown, the cross-window attention module includes window attention and token attention between windows, which are used to extract local and global information respectively. In window attention, the feature map is partitioned into windows, and then token attention is calculated for each small window. In token attention between windows, the windows are also partitioned. At this stage, the information at the same position of each window is extracted to form a new window, and token attention is calculated for it. Then, the results of the two stages are sent into a 3×3 convolution for merging.

[0082] Step 2: Train the established image restoration network:

[0083] During the training process of the entire model, this embodiment uses two loss functions to constrain the optimal objective of the model. They are the per-pixel loss and the frequency domain loss. The per-pixel loss is used to measure the element difference between the clean image and the high-quality image restored by the model. The frequency domain loss is used to calculate the difference between the frequency domain information corresponding to the clean image and the restored high-quality image after Fourier transform. The formulaic expressions of the above two loss functions are as follows:

[0084] L pixel = ||I gt - D(I bad )||1

[0085] L f = ||F(I gt ) - F(I bad )||1

[0086] In the formula, I gt represents the clean image, and I bad represents the high-quality image containing noise. D(g) represents the calculation process of the image restoration network, and F(g) represents the Fourier transform.

[0087] Step 3: Input the image to be restored into the trained image restoration network, and the image restoration network outputs the restored image;

[0088] The quantitative results on the defogging and desnowing datasets are shown in Table 1:

[0089] Table 1 Test results of the method of this embodiment on the defogging and desnowing datasets

[0090]

[0091] The quantitative results on the low-light image enhancement dataset are shown in Table 2:

[0092] Table 2 Test results of the method of this embodiment on the low-light image enhancement dataset

[0093] LOL-v1 L0L-v2real LOL-v2syn PSNR 43.40 43.12 43.21 SSIM 0.992 0.991 0.991

[0094] The images of the image restoration results of the image restoration network in this embodiment are as Figure 8 shown in the two groups of images before and after restoration in (a) and (b).

[0095] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not depart from the spirit and scope of the present invention as defined by the appended claims. It should be understood that different dependent claims and the features in this document can be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a single embodiment can be used in other embodiments.

Claims

1. An image restoration method, characterized in that Including: Input the image to be restored into the image restoration network, and the image restoration network outputs the restored image; The image restoration network is implemented using a U-Net architecture and transformer blocks; the U-Net architecture includes four layers of encoding and four layers of decoding. Each of the first three layers of encoding and the last three layers of decoding uses a transformer block, and the fourth layer of encoding and the first layer of decoding share a transformer block. Among them, the transformer blocks in the first three layers of encoding and the last three layers of decoding use compact attention to extract channel domain information, and the fourth layer of encoding and the first layer of decoding use cross-window attention to extract spatial domain information; The method of using compact attention to extract channel domain information is: Extract and compress the information of the input image features using strided deformable convolution, that is, in the form where the convolution kernel size and the stride of the strided deformable convolution are equal, and generate a position offset of size , and represent the height and width of the image to be restored; The method of using cross-window attention to extract spatial domain information is: First, the input image features are evenly divided into two parts in the channel dimension ; Pair Perform uniform window segmentation to obtain multiple small windows, and perform parallel window attention calculations on each small window respectively; Pair Perform uniform window partitioning to obtain multiple small windows, and extract the information at the corresponding positions of each small window to form a tensor , and calculate each tensor in parallel 's token attention; Use ordinary convolution to fuse the calculated window attention and token attention as the extracted spatial domain information.

2. The image restoration method according to claim 1, wherein The image restoration network includes 3 shallow layers and 7 transformer blocks: Treat the restored image Perform scale transformation to obtain three different scale images of the restored image; Denotes a set of tensors, with superscripts indicating dimensions, is the height of the image to be restored, is the width of the image to be restored, denotes the number of channels; The image to be restored is input into the first transformer block after passing through a 3×3 convolution, and the output of the first transformer block is downsampled; Three scale images sorted from high to low. The image to be restored at the first scale is input into the first shallow layer, and the first shallow layer outputs the feature image; The downsampled result of the output of the first transformer block and the output of the first shallow layer are concatenated, and the concatenated result is input into the second transformer block after passing through a 3×3 convolution, and the output of the second transformer block is downsampled; The image to be restored at the second scale is input into the second shallow layer, and the second shallow layer outputs the feature image; The downsampled result of the output of the second transformer block and the output of the second shallow layer are concatenated, and the concatenated result is input into the third transformer block after passing through a 3×3 convolution, and the output of the third transformer block is downsampled; The image to be restored at the third scale is input to the third shallow layer, and the third shallow layer outputs the feature image; The downsampled result of the output of the third transformer block and the output of the third shallow layer are concatenated, and the concatenated result is input into the fourth transformer block after passing through a 3×3 convolution, and the output of the fourth transformer block is upsampled; The upsampled result of the output of the fourth transformer block and the output of the third transformer block are concatenated, and the concatenated result is input into the fifth transformer block after passing through a 1×1 convolution, and the output of the fifth transformer block is upsampled; The upsampled result of the output of the fifth transformer block and the output of the second transformer block are concatenated, and the concatenated result is input into the sixth transformer block after passing through a 1×1 convolution, and the output of the sixth transformer block is upsampled; The upsampled result of the output of the sixth transformer block and the output of the first transformer block are concatenated, and the concatenated result is input into the seventh transformer block after passing through a 1×1 convolution; The output of the 7th transformer block is convolved with a 3×3 convolution and then added element-wise to the image to be restored, and the result of the addition is the restored image.

3. The image restoration method according to claim 2, wherein The transformer block includes a compact deformable attention module, a cross-window attention module, and a gated depthwise convolutional feed-forward network; The input image features are normalized by layer normalization and then enter the compact deformable attention module or the cross-window attention module. The transformer blocks in the first three encoding layers and the last three decoding layers use the compact deformable attention module to extract channel-domain information, and the fourth encoding layer and the first decoding layer use cross-window attention to extract spatial-domain information; The output of the compact deformable attention module or the cross-window attention module is added element-wise to the input image features, and the result of the addition is normalized by layer normalization and then enters the gated depthwise convolutional feed-forward network. The output of the gated depthwise convolutional feed-forward network and the result of the addition are added element-wise, and the result of the addition is the output of the transformer block.

4. The image restoration method according to claim 3, wherein The compact deformable attention module includes a 1×1 convolution, an offset generator, and two depthwise separable convolutions DWConv; The input image features are , , image features After layer normalization, enter a 1×1 convolution. The result after the 1×1 convolution enters the offset generator, and the offset generator outputs a position offset , The first depthwise separable convolution DWConv is set to have the same convolution kernel size and stride size. The first depthwise separable convolution DWConv selects the relevant information of the query Query and the key Key from the result to complete the encoding of the query Query and the key Key, obtaining the encoding matrices and and , , , the encoding matrices and are respectively reshaped to obtain the encoding matrices and , , , the encoding matrices and are matrix-multiplied and then a softmax operation is performed to obtain , where represents the convolution kernel size of the first depthwise separable convolution DWConv; The second depthwise separable convolution DWConv is a convolution that selects relevant information of the values in the result Value to complete the encoding of the values Value and obtain the encoded matrix , . After reshaping the encoded matrix , the encoded matrix is obtained, ; Obtained After performing matrix multiplication with the encoding matrix and reshaping, obtain , , which is the channel domain information extracted by the compact deformable attention module.

5. The image restoration method according to claim 4, characterized in that The offset generator includes a sequential connection k × k convolution and 1×1 convolution.

6. The image restoration method according to claim 2, wherein The shallow layer includes a 3×3 convolution, a 1×1 convolution, a 3×3 convolution, and a 1×1 convolution connected in sequence.

Citation Information

Patent Citations

  • Image processing method and device based on depth dynamic self-adjustment

    CN118864503A

  • Synthetic aperture optical image restoration method based on local-global feature enhancement

    CN119205581A