Transform-based image restoration method
By introducing the Transformer-based local-region-global perceptual attention basic module in the image recovery method, the problem of difficult decoupling architecture and basic modules in the prior art independently contribute to performance improvement is solved, and an efficient image recovery effect is achieved under a variety of bad weather conditions.
Patent Information
- Application Number
- CN202411864850.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively decouple the independent contribution of architecture and basic modules to performance improvement, making it difficult to evaluate and optimize the two separately, which in turn affects the effect of image recovery, especially in a variety of harsh weather conditions.
The image recovery method based on Transformer is adopted to enhance feature extraction and fusion through the local-region-global perceptual attention basic module (LRG), combining multi-head multi-scale fusion attention and channel space dual attention feedforward network to enhance feature extraction, expression and robustness.
The image recovery effect of the model under a variety of harsh weather conditions is improved, feature extraction and expression capabilities are enhanced, robustness is improved, and the generalization performance of image recovery is improved.
Smart Images

Figure CN120013814A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision image technology, and in particular to a Transformer-based image restoration method. Background Art
[0002] Due to the hardware limitations of image acquisition equipment and the influence of extremely bad weather (such as raindrops, rain lines, haze, and snowflakes), outdoor image acquisition systems will collect low-quality images. Low-quality images not only affect the subjective visual experience of the human eye and easily lead to operational errors, but also interfere with subsequent intelligent processing, increasing the risk of misjudgment and accidents. The imaging mechanism under different weather conditions is different, and its impact on the image also varies with the severity of the weather, extending from local impact to global impact. For example, snow particles mainly cause local occlusion effects; rainy days have both global rain line effects and may also cause local occlusions due to raindrops on the lens; haze weather usually causes image blur on a global scale.
[0003] In order to improve the generalization of the model in processing degraded images under various severe weather conditions, it is necessary to comprehensively consider the diverse impacts of different weather conditions on images. Existing image restoration methods that combine network architecture and basic module design (such as CNN-based AirNet, MPRNet, and Transformer-based IPT, etc.) are difficult to decouple the independent contributions of the architecture and basic modules to performance improvement, making it difficult to evaluate and optimize the two separately. Since the model is built by stacking multiple basic modules, the design of the basic modules has a decisive influence on the overall performance. To this end, it is necessary to optimize the basic modules to improve the model's feature extraction ability, expression ability, and robustness, thereby improving the effect of image restoration. Summary of the invention
[0004] The purpose of the embodiment of the present invention is to provide a method for coping with the image degradation problem caused by various severe weather conditions.
[0005] In order to achieve the above-mentioned purpose, an embodiment of the present invention provides a Transformer-based image restoration method, which includes: obtaining image features after preliminary processing of an input image; inputting the image features into a local-regional-global perception attention basic module LRG with different network depths for feature extraction and fusion to obtain fused image features; performing a downsampling operation on the fused image features to obtain image features of different dimensions; after fusing the image features of different dimensions layer by layer according to the corresponding feature sizes, upsampling the fused image features of different dimensions to the feature size of the image features; and concatenating the upsampled features with the fused image features, and integrating them into the image features obtained after preliminary processing of the input image to obtain output image features, and converting the output image features into an output image.
[0006] Optionally, the local-regional-global perception attention basic module LRG includes multi-head multi-scale fusion attention and a channel space dual attention feedforward network; wherein, the multi-head multi-scale fusion attention includes a local information extraction module and a global information extraction module.
[0007] Optionally, the local information extraction module includes rotational equivariant convolution and an alternating local self-attention mechanism; the global information extraction module includes a frequency domain adaptive attention mechanism and a query-aware global sparse adaptive attention mechanism.
[0008] Optionally, the image features are input into a local-regional-global perception attention basic module LRG of different network depths for feature extraction and fusion, including: normalizing the image feature layer and performing linear projection so that the image features are mapped to multiple attention heads of the multi-head multi-scale fusion attention; using the local information extraction module to obtain local feature information of the image features, and adding part of the image features and part of the local feature information and inputting them into the global information extraction module to obtain the global feature information of the image features; concatenating the local feature information and the global feature information to obtain combined features, and normalizing the combined features and inputting them into the channel space dual attention feedforward network; and using the channel space dual attention feedforward network to obtain features of receptive fields of different sizes of the combined features, and fusing the features of the receptive fields of different sizes in spatial dimensions and channel dimensions.
[0009] Optionally, the use of the local information extraction module to obtain the local feature information of the image feature includes: using the filter of the rotational equivariant convolution to perform a convolution operation on the input image feature at a preset angle, obtaining rotation parameters of different angles to update the filter, and using the filter to perform a convolution operation on the image feature to generate an output feature map; using the alternating local self-attention mechanism to perform a blocking operation on the input image feature, dividing the image feature into square, horizontal and vertical bar areas, performing regional block self-attention, horizontal self-attention and vertical self-attention in the square, horizontal and vertical bar areas respectively, and combining the outputs after the self-attention through a concatenate operation and outputting them; and adding the feature map output by the rotational equivariant convolution and the output of the alternating local self-attention mechanism and outputting them as the local feature information of the image feature obtained by the local information extraction module.
[0010] Optionally, obtaining the global feature information of the image features includes: using the frequency domain adaptive attention mechanism to convert the input image features into frequency domain features, performing operations on the frequency domain features to obtain attention features, and adding the attention features to the output of the local information extraction module and inputting them into the query-aware global sparse adaptive attention mechanism; the query-aware global sparse adaptive attention mechanism uses a query image block to filter the key-value pair areas in the input image features to obtain the most relevant key-value pair areas, and aggregating the most relevant key-value pair areas and calculating to obtain sparse global features.
[0011] Optionally, the method of using the channel-space dual-attention feedforward network to obtain features of receptive fields of different sizes of the combined features includes: inputting the combined features into a linear mapping layer and an activation function layer for processing, and dividing the processed combined features into two parts of features according to channels, and performing deep convolution operations and dilated convolution operations on the two parts of features, respectively, to obtain features of receptive fields of different sizes.
[0012] In a second aspect, the present invention provides an image restoration system based on Transformer, the system comprising: a preprocessing module, used to obtain image features after performing preliminary processing on an input image; a multi-head multi-scale fusion attention module, used to extract local feature information and global feature information of the image features, and fuse the two to obtain multi-scale feature information; a channel space double attention feedforward network, used to obtain the spatial dimension and channel dimension feature information of the multi-scale feature information, and fuse them to obtain the fused image features; a downsampling module, used to perform a downsampling operation on the fused image features to obtain image features of different dimensions; a summing module, used to fuse the image features of different dimensions layer by layer according to the corresponding feature sizes; an upsampling module, used to fuse the fused image features of different dimensions layer by layer and upsample them to the feature size of the image features; a feature information fusion module, used to concatenate the upsampled features with the fused image features, and integrate them into the image features obtained after preliminary processing of the input image to obtain output image features; and a feature conversion module, used to convert the output image features into an output image.
[0013] Optionally, the multi-head multi-scale fusion attention module includes a local information extraction module and a global information extraction module; wherein, the local information extraction module includes rotational equivariant convolution and alternating local self-attention mechanism, and the global information extraction module includes frequency domain adaptive attention mechanism and query-aware global sparse adaptive attention mechanism.
[0014] In a third aspect, the present invention provides a processor for running a program, wherein the program, when run, is used to execute the Transformer-based image restoration method described in any one of the present applications.
[0015] Through the above technical solution, the present invention first converts the input image into image features after preliminary processing, and uses the local-regional-global perception attention basic module LRG to obtain the key features in the local, regional and global ranges of the image features, and combines them to achieve the purpose of fusing multi-scale feature information. And by fusing the combined multi-scale feature information in the channel dimension and the spatial dimension, the information interaction between channels and the spatial modeling ability are enhanced.
[0016] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings:
[0018] Figure 1 is a flowchart of a Transformer-based image restoration method provided by an embodiment of the present disclosure;
[0019] Figure 2 It is a schematic diagram of the architecture of a Transformer-based image restoration method equipped with a basic module provided by an embodiment of the present disclosure;
[0020] Figure 3 It is a structural diagram of a local-regional-global perception attention basic module LRG provided by an embodiment of the present disclosure;
[0021] Figure 4 is a structural schematic diagram of a local information extraction module provided by an embodiment of the present disclosure;
[0022] Figure 5 is a structural diagram of a global information extraction module provided by an embodiment of the present disclosure;
[0023] Figure 6 is a schematic diagram of a channel space dual attention feedforward network provided by an embodiment of the present disclosure;
[0024] Figure 7 A Transformer-based image restoration system is provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] The specific implementation of the embodiment of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the embodiment of the present invention, and is not used to limit the embodiment of the present invention.
[0026] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application are in compliance with the relevant provisions of national laws and regulations. In the embodiments of this application, some existing solutions in the industry such as certain software, components, and models may be mentioned, which should be considered as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of this application, but it does not mean that the applicant has or will necessarily use the solution.
[0027] Figure 1 is a flowchart of a Transformer-based image restoration method provided by an embodiment of the present disclosure, Figure 2 Schematic diagram of the architecture of a Transformer-based image restoration method equipped with a basic module provided by an embodiment of the present disclosure. Figure 1As shown, the method includes: obtaining image features after preliminary processing of the input image; inputting the image features into the local-regional-global perception attention basic module LRG of different network depths for feature extraction and fusion to obtain fused image features; performing a downsampling operation on the fused image features to obtain image features of different dimensions; after fusing the image features of different dimensions layer by layer according to the corresponding feature sizes, upsampling the fused image features of different dimensions to the feature size of the image features; and concatenating the upsampled features with the fused image features, and integrating them into the image features obtained after the preliminary processing of the input image to obtain output image features, and converting the output image features into an output image.
[0028] Understandably, Figure 1 The Transformer-based image restoration method provided in the embodiment of the present disclosure is based on Figure 2 The architecture of the basic module provided in the embodiment of the present disclosure is realized. The architecture of the basic module provided in the embodiment of the present disclosure adopts U-Net as the network architecture, and the whole includes two parts, an encoder and a decoder, and the structures of the two are completely symmetrical. It can be understood that the encoder and decoder mentioned here include the local-regional-global perception attention basic module LRG proposed in the present disclosure. During the training process, the embodiment of the present disclosure adopts a four-layer network depth. First, the input image is preprocessed, and the image is converted into image features and input into the network architecture of the present disclosure. The local-regional-global perception attention basic module LRG is used to extract and fuse the image features to obtain the fused image features. As the network depth increases, the downsampling module of the encoder in the embodiment of the present disclosure performs downsampling operations layer by layer, and gradually extracts and compresses the feature information to obtain image features of different dimensions. At the same time, the number of attention heads increases according to the network depth, and is set to 8, 16, 32, and 64 in sequence. Subsequently, the upsampling module in the decoder performs upsampling operations step by step, and restores the feature information to high resolution layer by layer. In addition, the encoder and the decoder establish a residual connection between the corresponding levels, and directly pass the output of each level of the encoder to the corresponding decoder part, which is used to concatenate the image features of different dimensions obtained by downsampling and the features after upsampling, and integrate the image features of the input image to obtain the output image features, further enriching the feature representation. The output features are converted into an output image through a feature conversion module, and the output image is the image to be restored in the present disclosure. The specific operating parameters are shown in Tables 1 to 4.
[0029] Table 1 Parameters of Stage 1 unit in codec
[0030]
[0031] Table 2 Parameters of Stage 2 unit in codec
[0032]
[0033] Table 3 Parameters of Stage 3 unit in codec
[0034]
[0035] Table 4 Parameters of Stage 4 unit in codec
[0036]
[0037]
[0038] Figure 3 Schematic diagram of the structure of a local-regional-global perception attention basic module LRG provided by an embodiment of the present disclosure. Figure 3 As shown, the local-regional-global perception attention basic module LRG includes multi-head multi-scale fusion attention and channel space dual attention feedforward network; wherein, the multi-head multi-scale fusion attention includes a local information extraction module and a global information extraction module.
[0039] Figure 4 is a structural diagram of a local information extraction module provided by an embodiment of the present disclosure, Figure 5 Schematic diagram of a global information extraction module provided by an embodiment of the present disclosure. Figure 4 and Figure 5 As shown, the local information extraction module includes rotational equivariant convolution and alternating local self-attention mechanism; the global information extraction module includes frequency domain adaptive attention mechanism and query-aware global sparse adaptive attention mechanism.
[0040] Specifically, the local-regional-global perception attention basic module LRG provided in the embodiment of the present disclosure can extract local feature information and global feature information of the input image features, and combine the two, and output the combined feature information after fusion of the channel dimension and the spatial dimension. The LRG module mainly includes a multi-head multi-scale fusion attention and a channel space dual attention feedforward network. First, a local and global information extraction module is designed in the multi-head multi-scale fusion attention, in which the local information extraction module can effectively solve the problem of local information sharing, especially capturing features with similar structures at different rotation angles; the global information extraction module greatly reduces resource consumption by optimizing the global attention mechanism, and solves the defect that global sparse attention cannot be adaptive. Secondly, the channel space dual attention feedforward network expands the receptive field while keeping the computational cost unchanged, thereby enhancing the detail capture capability in the spatial dimension and the feature interactivity in the channel dimension in the feedforward network process. Finally, the embodiment of the present disclosure is universal, can be equipped with various algorithm architectures and deployed in various outdoor image acquisition systems, and has the ability to restore image quality in real time and adaptively under various severe weather conditions.
[0041] In some embodiments, the image features are input into a local-regional-global perception attention basic module LRG of different network depths for feature extraction and fusion, including: normalizing the image feature layer and performing linear projection so that the image features are mapped to multiple attention heads of the multi-head multi-scale fusion attention; using the local information extraction module to obtain local feature information of the image features, and adding part of the image features and part of the local feature information and inputting them into the global information extraction module to obtain the global feature information of the image features; concatenating the local feature information and the global feature information to obtain combined features, and normalizing the combined features and inputting them into the channel space dual attention feedforward network; and using the channel space dual attention feedforward network to obtain features of receptive fields of different sizes of the combined features, and fusing the features of the receptive fields of different sizes in spatial dimensions and channel dimensions.
[0042] Specifically, the multi-head multi-scale fusion attention combines rotational equivariant convolution and alternating local self-attention mechanism as a local information extraction module, proposes a frequency domain adaptive attention mechanism and a query-aware global sparse adaptive attention mechanism, and combines the two as a global information extraction module. The attention range from small to large receptive fields is gradually added and fused along the channel, and finally the spatial and frequency domain information content captured in the local, regional and global ranges are aggregated to achieve the purpose of capturing and fusing multi-scale information. The channel space dual attention feedforward network enhances the information interaction and spatial modeling capabilities between channels by combining deep convolution, dilated convolution and element multiplication operations on the basis of MLP. By training the network, it can adaptively identify image degradation features under different severe weather conditions and achieve effective restoration, while ensuring real-time inference speed, thereby meeting the efficient processing requirements in practical applications.
[0043] It can be understood that the multi-head multi-scale fusion attention first focuses on the layer-normalized feature X∈R (H ×W)×C A linear projection is performed, where H represents the height of the input image, W represents the width of the image, and C represents the number of channels of the image. It is mapped to h attention heads. Subsequently, these heads are evenly divided into four branches (each branch contains h / 4 heads), processing local, regional, and global information respectively.
[0044] In some embodiments, the use of the local information extraction module to obtain the local feature information of the image feature includes: using the rotation equivariant convolution filter to perform a convolution operation on the input image feature at a preset angle, obtaining rotation parameters of different angles to update the filter, and using the filter to perform a convolution operation on the image feature to generate an output feature map; using the alternating local self-attention mechanism to block the input image feature, dividing the image feature into square, horizontal and vertical bar areas, performing regional block self-attention, horizontal self-attention and vertical self-attention in the square, horizontal and vertical bar areas respectively, and combining the outputs after the self-attention through a concatenate operation and outputting them; and adding the feature map output by the rotation equivariant convolution and the output of the alternating local self-attention mechanism and outputting them as the local feature information of the image feature obtained by the local information extraction module.
[0045] Specifically, the background content of low-quality images often has some parts with similar structures at different rotation angles. However, standard convolutional networks are usually unable to predict the feature mapping relationship in the same local rotation structure information, and it is difficult to share the expression. In contrast, rotation equivariant convolution combines the rotation and circular shift within the channel to make the rotation operation of the convolution filter consistent with the rotation operation of the input image, so that the change of the feature map at different rotation angles becomes predictable. Specifically, the embodiment of the present disclosure rotates the input image by 2kπ / t degrees (k=1,2,3,4), where the equivariant t is set to 4, that is, each rotation is 90 degrees. The filter of the rotation equivariant convolution performs a convolution operation according to the preset 2kπ / t angle, and shares the same parameters between different rotation angles, thereby realizing the reuse of parameters four times at four rotation angles, significantly reducing the number of parameters of the network when processing the rotation input. At the same time, the rotation equivariant design enables the convolutional neural network to effectively identify the feature expression of similar objects at different rotation angles, and enhances the robustness to the rotation angle on the basis of retaining the translation invariance of the convolution module.
[0046] like Figure 4 As shown in Figure 5, the rotation equivariant convolution includes a shallow feature extraction network, a deep feature extraction network and an output. The shallow feature extraction network and output of the rotation equivariant convolution each contain an Fconv_PCA, and the deep feature extraction network contains 4 groups of residual networks that perform rotation operations respectively. Each group of residual networks contains an Fconv_PCA and an activation function ReLU. Fconv_PCA first initializes the PCA basis function, weights, biases, and expansion multiples of the convolution kernel. Then the PCA basis transformation weight calculation is performed, and the basis function and weights are linearly transformed to obtain the transformed weights. The complete convolution kernel size_P is formed according to the expansion multiple, and finally a 2D convolution operation is performed on the input to generate an output feature map. The specific operation parameters are shown in Table 5.
[0047] Table 5 Operation parameters of rotational equivariant convolution
[0048]
[0049] Specifically, the alternating local self-attention mechanism aims to enhance the expressive power of regional features. The features of natural images usually appear in an anisotropic manner, and the global isotropic attention is redundant for capturing anisotropic image features. Self-attention processing is performed within the anisotropic stripes, and an approximate global attention effect can be achieved in a manually processed static prior manner. The present invention is designed to gradually increase the local block size by sequentially executing 4 LRG basic modules in each layer of network depth. The set block size is [4, 8, 16, 32]. The local attention in each basic module includes three shapes: square, horizontal bar and vertical bar. And self-attention operation is performed in each block, which effectively improves the ability to capture local and regional features.
[0050] The local-region-global perception attention base module LRG is executed 4 times in each network depth of the encoder-decoder. The alternating local self-attention mechanism performs regional block self-attention, horizontal self-attention and vertical self-attention in each network depth. The size of the block increases with the increase of network depth, and the block side length is set to 4, 8, 16, and 32 respectively. Figure 4 As shown, the input features The blocks are designed as square X L , horizontal bar X R and vertical bar X C , as shown in the formula:
[0051]
[0052] The dimension of the block shape is Number of blocks Where H, W, C are the height, width and number of channels of the input image, and sh, sw are the height and width of the block. This design ensures that the model can effectively extract important information in the area block by flexibly adjusting the block size. In the self-attention calculation, assuming that the projected query, key and value of the mth feature head all have the dimension of dm, the self-attention of the area square, horizontal strip and vertical strip is:
[0053]
[0054] in, denote the projection matrices of q, k and v of the mth head respectively, is the output of the mth feature head and the nth self-attention, and Attn(·) is the self-attention calculation, as shown in the formula: Among them, Q is the query matrix, K is the key matrix, V is the value matrix, K Tis the transposed matrix of matrix K, and d represents the distance between Q and K. The outputs of n attentions are combined through concatenation as the output of the self-attention feature head group of the regional square, horizontal bar and vertical bar: The output dimension of the self-attention feature head group of the area square, horizontal bar and vertical bar is Y L ∈R (H×W)×C / 4 ,Y R ∈R (H×W)×C / 4 ,Y C ∈R (H ×W)×C / 4 The self-attention of the regional square, horizontal bar and vertical bar is combined through concatenation, and the output of the final local attention branch is: X attn2 =Local Attention(X 2 )=concatenate(Y L ,Y R ,Y C ), the specific operation parameters are shown in Table 6. Through the above operations, the receptive field is gradually increased, which can improve the module's ability to extract local and regional features. The overall operation process formula of the local information capture module is 2 =X attn1 +X attn2 , where X attn1 is the output of the rotational equivariant convolution operation branch, X attn2 is the output of the local attention operation branch.
[0055] Table 6 Alternating local self-attention mechanism operation parameters
[0056]
[0057] In some embodiments, the obtaining of global feature information of the image features includes: using the frequency domain adaptive attention mechanism to convert the input image features into frequency domain features, performing operations on the frequency domain features to obtain attention features, and adding the attention features to the output of the local information extraction module and inputting them into the query-aware global sparse adaptive attention mechanism; the query-aware global sparse adaptive attention mechanism uses query image blocks to filter the key-value pair areas in the input image features to obtain the most relevant key-value pair areas, and aggregating the most relevant key-value pair areas and calculating to obtain sparse global features.
[0058] Specifically, Figure 5As shown, the global information extraction module includes a frequency domain adaptive attention mechanism and a query-aware global sparse adaptive attention mechanism. This module can adaptively remove noise according to different positions of the image and focus on the most relevant global information, thereby more effectively understanding and restoring the global content of the image.
[0059] Global attention calculations often consume a lot of computing resources, and through frequency domain transformation, the main features in the image can be concentrated in fewer frequency components, thereby promoting efficient feature extraction. In addition, the frequency domain information expression can more clearly highlight the boundary between bad weather (such as rain lines, raindrops, snowflakes) and the image background, improve the noise filtering effect, and ensure the integrity of the background information. The disclosed embodiment globally converts the processing of image feature information from the spatial domain to the frequency domain level. QK in attention operation T The calculation of is obtained by taking the inner product of each element in the query matrix Q and the key matrix K: <Q,K t >= i ,k j >.q i and k j They are feature F q and F k Inspired by the convolution theorem, the frequency domain adaptive attention mechanism in the disclosed embodiment uses the fast Fourier transform method to convert the spatial domain features F q and F k Convert to frequency domain features Q, K, and effectively estimate the attention map by element-wise multiplication in the frequency domain, instead of calculating QK in the spatial domain T This is a more complicated calculation method of matrix multiplication.
[0060] Specifically, first, the input feature X 3 Perform a linear layer operation to ensure that the input X 3 Keep a stable distribution, and then process it through three 1×1 convolutional layers and 3×3 deep convolutional layers to obtain the feature F q , F k and F v Then, for feature F q and F k Apply the Fast Fourier Transform (FFT) to estimate F q and F k Correlation in the frequency domain A: Where F(·) represents the fast Fourier transform FFT, F -1 (·) represents inverse FFT, represents the conjugate transpose operation. We get F q and F k After the correlation matrix A in the frequency domain is obtained, the correlation matrix A is further processed using a 3×3 deep convolution to better aggregate information and conduct a deeper study of the frequency domain correlation. v Perform point multiplication to obtain the final attention feature V attn :V attn =Dwconv 3 (LN(A))F v . Where Dwconv 3 (·) is a 3×3 deep convolutional layer, and LN(·) is a layer normalization operation. Finally, the attention output of the frequency domain adaptive attention mechanism branch is obtained by estimating the aggregated features: attn3 =X 3 +Conv(V attn ), the specific operation parameters are shown in Table 7. The output result of the frequency domain adaptive attention mechanism branch is added to the output result of the local information extraction module as part of the feature input for the next operation: Y 3 =Y 2 +X attn3 .
[0061] Table 7 Frequency domain adaptive attention mechanism operation parameters
[0062]
[0063]
[0064] The query-aware global sparse adaptive attention mechanism in the second branch of the global information extraction module in the disclosed embodiment is an adaptive, query-aware global sparse attention mechanism. This attention mechanism can adaptively and sparsely focus on the most relevant parts of the content in different images globally. By dividing the image into multiple patches, most of the irrelevant key-value (KV) regions are first coarsely filtered out globally according to the image blocks of the query Q, and only the k key-value regions that are most relevant to the query are retained. Then the k key-value regions are clustered together and fine-grained sparse global attention is performed with the query Q.
[0065] Specifically, in the last branch of the multi-head multi-scale fusion attention, the embodiment of the present disclosure inputs the feature X 4 ∈R H×W×C Divided into s×s non-overlapping blocks, and then linearly project them to obtain Q, K, and V. Use the Top-k operator to derive k correlation matrices I∈N s×s×k , get the k blocks most relevant to the query block: I = topk (QK T). The i-th row of the correlation matrix I contains the k regions that are most relevant to the i-th region. After obtaining the correlation matrix I, further fine-grained inter-block attention is performed. For each query Q, the key-value pairs in all the k regions found are aggregated together to obtain the aggregated key-value pair K g and V g , and then calculate the attention score. Finally, a 3×3 deep convolution layer and a 5×5 dilated convolution layer are used to enhance the local context features again. The purpose of using a 3×3 deep convolution layer and a 5×5 dilated convolution layer is to expand the receptive field without increasing the number of parameters, and weighted with the attention score of the Top-k global selection: X attn4 =Attn(Q,K g ,V g )+Dwconv 3 (V)+Dlconv 5 (V) Dwconv 3 Denotes a 3×3 depth-wise separable convolution operation with a receptive field size, Dlconv 5 represents the dilated convolution operation with a receptive field size of 5×5. The specific operation parameters are shown in Table 8. The overall operation process of the global information capture module is as follows: 4 =Y 3 +X attn4 .
[0066] The above operation extracts features of different scales in the channel dimension, and expands the receptive field from small to large by gradually adding them. Finally, the focus content of the four channel branches is aggregated to achieve comprehensive capture of multi-scale information. Through this design, the model can extract image features at multiple scales, including local, regional, and global, significantly improving the processing capability and generalization performance of complex image information.
[0067] Table 8 Query-aware global sparse adaptive attention mechanism operation parameters
[0068]
[0069] The operations performed by the local information extraction module and the global information extraction module in the disclosed embodiment are to extract features of different scales in the channel dimension, and expand the receptive field from small to large in sequence by gradually adding them. Finally, the focus content of the four channel branches is aggregated to achieve comprehensive capture of multi-scale information. This design can extract image features at multiple scales, locally, regionally, and globally, significantly improving the processing capability and generalization performance of complex image information.
[0070] In some embodiments, the use of the channel-space dual-attention feedforward network to obtain features of receptive fields of different sizes of the combined features includes: inputting the combined features into a linear mapping layer and an activation function layer for processing, and dividing the processed combined features into two parts of features according to the channel, and performing deep convolution operations and dilated convolution operations on the two parts of features, respectively, to obtain features of receptive fields of different sizes.
[0071] Figure 6 Schematic diagram of a channel space dual attention feedforward network provided by an embodiment of the present disclosure. Figure 6 As shown in the figure, the channel space dual attention feedforward network takes the overall output of multi-head multi-scale fusion attention as the input X of this module, which first passes through a linear mapping layer W 1 (·) and GELU activation function σ(·) to obtain X′: X′=σ(W 1 X). Then X′ is divided into X by channel. 1 ′ ,X 2 ′ The two parts are respectively subjected to depth convolution and dilated convolution operations in the two channels. The depthwise separable convolution branch sets the convolution kernel size to 3×3 and the receptive field to 3×3; the dilated convolution branch sets the dilation factor to 2 and the convolution kernel size to 3×3. The dilated convolution increases the receptive field to 5×5 without increasing the number of parameters. This enables the network to capture a wider range of contextual information and understand the image content of larger regional features. The specific operation parameters are shown in Table 9. The features processed by the two different receptive fields are fused by element-by-element multiplication, and then passed through a linear mapping layer W. 2 (·) Restore to the original channel dimension: Y = W 2 [(Dwconv 3 (X′ 1 ))⊙(Dlconv 5 X 2 ′ )]. Where Dwconv 3 Denotes a 3×3 depth-wise separable convolution operation with a receptive field size, Dlconv 5 represents the dilated convolution operation with a receptive field size of 5×5, and ⊙ represents element-wise multiplication.
[0072] Table 9 Channel-space dual attention feedforward network parameters
[0073]
[0074] It can be understood that the channel-space dual-attention feedforward network provided in the disclosed embodiment combines the advantages of deep convolution and dilated convolution, effectively expands the receptive field without increasing the number of parameters, not only better captures detail information in the spatial dimension, but also enhances the interaction of channel features through element multiplication between channels, significantly improving the feature expression ability of the network. Compared with a simple fully connected layer, this part of the design can capture nonlinear spatial information and reduce the channel redundancy of the fully connected layer.
[0075] It should be noted that during the training process, the architecture of the Transformer-based image restoration method equipped with the basic module provided by the embodiment of the present disclosure is trained from scratch in an end-to-end manner using the ADAM optimizer. For all input weather data sets, the unified operation is performed, the image is cropped to a size of 256×256 according to the center, the number of training iterations is set to 100, the batch size is set to 16, and the model is supervised using L1 loss and perceptual loss Lpercep. The initial value of the learning rate is 0.0002, and a learning rate scheduler ReduceLROnPlateau is introduced. The scheduler takes the optimizer as input. When the monitored indicator stops improving, the learning rate will be reduced to half of the original, further optimizing the training process and avoiding premature stopping or overtraining.
[0076] like Figure 7 As shown, the embodiment of the present disclosure provides an image restoration system based on Transformer, and the system includes: a preprocessing module, which is used to obtain image features after preliminary processing of the input image; a multi-head multi-scale fusion attention module, which is used to extract local feature information and global feature information of the image features, and fuse the two to obtain multi-scale feature information; a channel space double attention feedforward network, which is used to obtain the spatial dimension and channel dimension feature information of the multi-scale feature information, and fuse them to obtain the fused image features; a downsampling module, which is used to perform a downsampling operation on the fused image features to obtain image features of different dimensions; a summing module, which is used to fuse the image features of different dimensions layer by layer according to the corresponding feature sizes; an upsampling module, which is used to upsample the fused image features of different dimensions to the feature size of the image features; a feature information fusion module, which is used to concatenate the upsampled features with the fused image features, and fuse them into the image features obtained after the input image is preliminarily processed to obtain output image features; and a feature conversion module, which is used to convert the output image features into an output image.
[0077] In some embodiments, the multi-head multi-scale fusion attention module includes a local information extraction module and a global information extraction module; wherein, the local information extraction module includes rotational equivariant convolution and alternating local self-attention mechanism, and the global information extraction module includes frequency domain adaptive attention mechanism and query-aware global sparse adaptive attention mechanism.
[0078] An embodiment of the present invention provides a processor, which is used to run a program, wherein the Transformer-based image restoration method is executed when the program is run.
[0079] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0080] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0081] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0082] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0083] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0084] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0085] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0086] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0087] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A Transformer-based image restoration method, characterized in that: The method comprises: After preliminary processing of the input image, image features are obtained; The image features are input into the local-region-global perception attention basic module LRG with different network depths for feature extraction and fusion to obtain fused image features; Performing a downsampling operation on the fused image features to obtain image features of different dimensions; Fusing the image features of different dimensions layer by layer according to the corresponding feature sizes, and upsampling the fused image features of different dimensions to the feature size of the image features; and The upsampled features are concatenated with the fused image features, and are integrated with the image features obtained after the input image is preliminarily processed to obtain output image features, and the output image features are converted into an output image.
2. The image restoration method according to claim 1, characterized in that: The local-regional-global perception attention basic module LRG includes a multi-head multi-scale fusion attention and a channel space dual attention feedforward network; Among them, the multi-head multi-scale fusion attention includes a local information extraction module and a global information extraction module.
3. The image restoration method according to claim 2, characterized in that: The local information extraction module includes rotational equivariant convolution and alternating local self-attention mechanism; The global information extraction module includes a frequency domain adaptive attention mechanism and a query-aware global sparse adaptive attention mechanism.
4. The image restoration method according to claim 3, characterized in that: The image features are input into the local-regional-global perception attention basic module LRG of different network depths for feature extraction and fusion, including: Normalizing the image feature layer and then performing linear projection so that the image features are mapped to multiple attention heads of the multi-head multi-scale fusion attention; Using the local information extraction module to obtain local feature information of the image feature, and adding part of the image feature and part of the local feature information and inputting the sum into the global information extraction module to obtain global feature information of the image feature; Concatenate the local feature information with the global feature information to obtain a combined feature, and normalize the combined feature and input it into the channel space dual attention feedforward network; and The channel-space dual-attention feedforward network is used to obtain the features of the receptive fields of different sizes of the combined features, and the features of the receptive fields of different sizes are fused in the spatial dimension and the channel dimension.
5. The image restoration method according to claim 4, characterized in that: The using the local information extraction module to obtain the local feature information of the image feature includes: Using the rotation equivariant convolution filter to perform a convolution operation on the input image feature at a preset angle, obtaining rotation parameters of different angles to update the filter, and using the filter to perform a convolution operation on the image feature to generate an output feature map; Using the alternating local self-attention mechanism, the input image features are divided into square, horizontal and vertical bar regions, and regional block self-attention, horizontal self-attention and vertical self-attention are performed in the square, horizontal and vertical bar regions respectively, and the outputs after the self-attention are combined and outputted through a concatenate operation; and The feature map output by the rotational equivariant convolution and the output of the alternating local self-attention mechanism are added and output as the local feature information of the image feature acquired by the local information extraction module.
6. The image restoration method according to claim 4, characterized in that: The acquiring of the global feature information of the image features comprises: The input image features are converted into frequency domain features by using the frequency domain adaptive attention mechanism, the frequency domain features are operated to obtain attention features, and the attention features are added to the output of the local information extraction module and then input into the query-aware global sparse adaptive attention mechanism; The query-aware global sparse adaptive attention mechanism uses the query image block to filter the key-value pair areas in the input image features to obtain the most relevant key-value pair areas, and obtains the sparse global features by aggregating and calculating the most relevant key-value pair areas.
7. The image restoration method according to claim 4, characterized in that: The method of using the channel space double attention feedforward network to obtain features of receptive fields of different sizes of the combined features includes: The combined features are input into the linear mapping layer and the activation function layer for processing, and the processed combined features are divided into two parts of features according to the channel, and the two parts of features are respectively subjected to deep convolution operation and dilated convolution operation to obtain features of receptive fields of different sizes.
8. A Transformer-based image restoration system, characterized in that: The system comprises: A preprocessing module is used to obtain image features after preliminary processing of the input image; A multi-head multi-scale fusion attention module is used to extract local feature information and global feature information of the image features, and fuse the two to obtain multi-scale feature information; A channel-space dual-attention feedforward network is used to obtain the spatial dimension of the multi-scale feature information and the feature information of the channel dimension, and fuse them to obtain fused image features; A downsampling module, used for performing a downsampling operation on the fused image features to obtain image features of different dimensions; An addition module, used for fusing the image features of different dimensions layer by layer according to the corresponding feature sizes; An upsampling module, used for upsampling the fused image features of different dimensions to the feature size of the image features; a feature information fusion module, configured to concatenate the upsampled features with the fused image features, and integrate the image features obtained after preliminary processing of the input image to obtain output image features; and The feature conversion module is used to convert the output image features into the output image.
9. The image restoration system according to claim 8, characterized in that: The multi-head multi-scale fusion attention module includes a local information extraction module and a global information extraction module; Among them, the local information extraction module includes rotational equivariant convolution and alternating local self-attention mechanism, and the global information extraction module includes frequency domain adaptive attention mechanism and query-aware global sparse adaptive attention mechanism.
10. A processor, characterized in that: Used to run a program, wherein the program, when run, is used to execute: the Transformer-based image restoration method according to any one of claims 1 to 7.
Citation Information
Cited By
Rape pod segmentation network model and system and pod counting method
CN120182608A