Image restoration method fusing linear attention mechanism

By integrating the linear attention mechanism to model the window and channel dimensions of the image feature map, the problem of high computational complexity of the visual Transformer image recovery method is solved, and efficient image recovery effect is achieved, which is suitable for computer vision tasks.

CN120355632AActive Publication Date: 2025-07-22JIANGSU HAOHAN INFORMATION TECH

Patent Information

Application Number
CN202510855077.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-22
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

The existing image recovery method based on visual Transformer has high computational complexity, resulting in inefficiency, especially when processing large-scale image data, and is difficult to process and train in real time, and the feature processing efficiency is inefficient at the window level and channel level, limiting its application potential in resource-constrained environments.

Method used

The fusion linear attention mechanism is adopted to optimize the image recovery model by performing attention modeling of window dimensions and channel dimensions on the image feature map, combining convolutional operations and residual connections, so as to optimize the image recovery model, reduce the computational complexity and improve the model efficiency.

Benefits of technology

It realizes that while reducing the computational complexity, improves the efficiency and applicability of the model, enhances image recovery quality and performance, and is suitable for a variety of computer vision tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355632A_ABST
    Figure CN120355632A_ABST
Patent Text Reader

Abstract

The invention discloses an image restoration method fusing a linear attention mechanism, and relates to the technical field of image processing, and the method comprises the steps: obtaining an image sample set, enabling a to-be-restored image and a corresponding target image to form a sample pair, and inputting the sample pair into an image restoration model; performing convolution preprocessing on the to-be-restored image, and extracting a first feature map with the size of W * H * C; window and channel dimension attention is established based on a linear attention mechanism, fusion attention output features are acquired, and K times of repetition are performed to output a final feature map; performing convolution post-processing on the final feature map, and performing residual connection with the to-be-restored image to generate a restored image; and constructing a loss function based on the target image, and optimizing image restoration model parameters through a gradient descent method. Therefore, the technical effects of reducing the calculation complexity, improving the model efficiency, enhancing the model applicability and performance, and maintaining or improving the image restoration quality are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image restoration method integrating a linear attention mechanism. Background Art

[0002] Currently, image restoration technology plays a key role in many fields, such as medical imaging, satellite image processing, and digital image restoration. With the rise of deep learning, a variety of advanced image restoration methods have been proposed, among which methods based on visual Transformers, such as Restormer and CAT, have gradually become mainstream. These methods can effectively capture long-distance pixel dependencies in images by introducing attention mechanisms, thereby achieving significant performance improvements in image restoration tasks.

[0003] However, existing image restoration methods based on visual transformers have a major problem in practical applications: the attention mechanism they use leads to a significant increase in computational complexity. Especially when processing large-scale image data, this increase in computational complexity directly reduces the efficiency of the model, making real-time processing and large-scale data training difficult. In addition, these methods do not perform targeted optimization when processing feature maps of different sizes and numbers of channels, resulting in inefficient feature processing at the window level and channel level, further increasing the computational burden of the model and limiting its application potential in resource-constrained environments. Summary of the invention

[0004] The present invention provides an image restoration method integrating a linear attention mechanism to solve the technical problems of high computational complexity and low efficiency in the prior art, thereby achieving the technical effect of reducing computational complexity, improving model efficiency, enhancing model applicability and performance, while maintaining or improving image restoration quality.

[0005] The present invention provides an image restoration method integrating a linear attention mechanism, comprising: An image sample set is obtained, and a sample pair is formed by the image to be restored and the corresponding target image, which is input into an image restoration model, wherein the output of the image restoration model is a restored image.

[0006] The image to be restored is preprocessed, and a convolution operation is used to extract a first feature map, wherein the size of the first feature map is W×H×C, W and H are the image width and height respectively, C is the number of convolution channels, and W, H, and C are all positive integers greater than 0.

[0007] Based on the linear attention mechanism, window dimension attention and channel dimension attention are established for the first feature map, and the fused attention output feature is obtained.

[0008] Repeat K times to obtain K of the fused attention output features, and the output is the final feature map.

[0009] Perform a post-processing operation based on convolution on the final feature map, and perform a residual connection between the convolution result and the image to be restored to obtain the restored image.

[0010] Construct a loss function based on the target image, and use the gradient descent method to optimize and update the parameters of the image restoration model.

[0011] In a feasible implementation, based on the linear attention mechanism, establish window dimension attention and channel dimension attention for the first feature map, and obtain the fused attention output feature, including: Perform window partitioning on the first feature map to obtain a set of sub-feature maps, and calculate sub-window attention features for each sub-feature map, and integrate them to obtain window attention features.

[0012] Partition the first feature map into a set of channel sub-blocks in the channel dimension, and calculate sub-channel attention features for each channel sub-block, and integrate them to obtain channel attention features.

[0013] Calculate the fused attention feature map according to the first feature map, the window attention feature and the channel attention feature.

[0014] In a feasible implementation, obtaining the window attention feature includes: Based on the sliding window method, partition the first feature map into windows of size M×M to form the set of sub-feature maps.

[0015] Perform a linear transformation on each sub-feature map to generate a first intermediate feature map, where the size of the intermediate feature map is W×H×3C.

[0016] Trisect the first intermediate feature map in the channel dimension, and based on the linear attention mechanism, calculate the first sub-linear attention feature in combination with the trisecting result, and further calculate the window attention feature.

[0017] In a feasible implementation, obtaining the channel attention feature includes: Partition the first feature map into L sub-blocks in the channel dimension, each sub-block includes C / L groups of features, and the L sub-blocks generate a set of channel sub-blocks.

[0018] Perform a linear transformation on each channel sub-block to generate a second intermediate feature map.

[0019] Trisect the second intermediate feature map in the channel dimension, and based on the linear attention mechanism, calculate the second sub-linear attention feature by combining the trisection result, and further calculate the channel attention feature.

[0020] In a feasible implementation, performing a linear transformation on each sub-feature map includes: For each of the sub-feature maps, perform the linear transformation of, where is the weight matrix, is the bias vector, is the first intermediate feature map, and .

[0021] In a feasible implementation, performing a linear transformation on each channel sub-block includes: For each of the sub-feature maps, perform the linear transformation of, where is the weight matrix, is the bias vector, is the second intermediate feature map, and .

[0022] In a feasible implementation, the linear attention function in the linear attention mechanism is defined as: ; where represents the linear attention feature of the nth set element; , , denotes taking the dth power of bit by bit; , , are the trisection results.

[0023] In a feasible implementation, calculating the first sub-linear attention feature, the operation formula is: ; where represents the ith first sub-linear attention feature. , , are the weight matrices. , , are the bias vectors, is the calculation result of the ith linear attention mechanism.

[0024] In a feasible implementation, calculating the channel attention feature includes: Concatenate the second sub-linear attention features output by each of the channel sub-blocks to obtain the channel attention feature with the number of channels restored to C.

[0025] In a feasible implementation manner, the loss function is defined as: ; where represents the restored image, is the target image, and N is the number of training samples.

[0026] The present invention discloses an image restoration method integrating a linear attention mechanism, including: obtaining an image sample set, forming a sample pair with the image to be restored and the corresponding target image, and inputting it into an image restoration model, where the output of the image restoration model is the restored image; preprocessing the image to be restored, and using a convolution operation to extract a first feature map, where the size of the first feature map is W××H×C, W and H are the image width and height respectively, C is the number of convolution channels, and W, H, and C are all positive integers greater than 0; establishing window-dimensional attention and channel-dimensional attention for the first feature map based on the linear attention mechanism, and obtaining a fused attention output feature; repeating K times to obtain K fused attention output features, and the output is the final feature map; performing post-processing operations based on convolution on the final feature map, and performing residual connection on the convolution result and the image to be restored to obtain the restored image; constructing a loss function based on the target image, and using the gradient descent method to optimize and update the parameters of the image restoration model. The image restoration method integrating a linear attention mechanism disclosed by the present invention solves the technical problems of high computational complexity and low efficiency, and realizes the technical effects of reducing computational complexity, improving model efficiency, enhancing model applicability and performance, and at the same time maintaining or improving the image restoration quality. Description of the Drawings

[0027] Figure 1 It is a flowchart of an image restoration method integrating a linear attention mechanism according to the present invention.

[0028] Figure 2 It is a flowchart in the image restoration method integrating a linear attention mechanism according to the present invention. Detailed Embodiments

[0029] The above technical solutions will be described in detail below in combination with the specification drawings and specific embodiments to better understand the above technical solutions. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited to the example embodiments only used to explain the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings rather than all of them.

[0030] Embodiment Figure 1 is a flowchart of an image restoration method integrating a linear attention mechanism according to the present invention. Among them, the image restoration method integrating a linear attention mechanism includes: S100: Obtain an image sample set, form a sample pair by combining the image to be restored and the corresponding target image, and input it into the image restoration model, where the output of the image restoration model is the restored image.

[0031] Specifically, the image sample set refers to a data set containing a large number of image pairs, and each image pair is composed of an image to be restored and a target image. Among them, the image to be restored refers to an image affected by degradation such as noise, blur, and damage, while the target image refers to the corresponding complete, clear, and non-degraded image, which can be used as a reference standard.

[0032] Among them, the image restoration model is a deep learning-based model, and its core function is to transform the image to be restored into a restored image by learning the features and patterns of the image. The restored image refers to the output image that is as close as possible to the target image after being processed by the model.

[0033] Specifically, first, collect the images to be restored with various degradation situations and the corresponding clear target images. These images can come from different application scenarios, such as satellite images with cloud occlusion and clear images without occlusion. Then, pair each image to be restored with the corresponding target image to form a sample pair. For example, for a group of images affected by Gaussian noise, pair them with the original noise-free images one by one to form sample pairs. These sample pairs are input into the image restoration model, and the model gradually adjusts its own parameters by learning the mapping relationship between the sample pairs to achieve effective restoration of the image to be restored.

[0034] By obtaining an image sample set and forming sample pairs to train the input model, two main technical effects can be achieved: First, the model can learn the image restoration rules under different degradation types and degrees, improving the generality and adaptability of the model; Second, through a large number of sample pair trainings, the accuracy and quality of the model's restored images can be improved, making them closer to the real target images, thereby improving the overall performance of image restoration.

[0035] S200: Preprocess the image to be restored, and use a convolution operation to extract the first feature map. Among them, the size of the first feature map is W×H×C, where W and H are the image width and height respectively, C is the number of convolution channels, and W, H, and C are all positive integers greater than 0.

[0036] Specifically, preprocessing is a series of preliminary operations performed on an image to improve the image quality, extract useful information, or reduce the data complexity. By sliding a convolution kernel on the preprocessed image and performing element-wise multiplication and summation operations, local features of the image, such as edges and textures, can be extracted.

[0037] Specifically, a feature map refers to a two-dimensional or three-dimensional data structure representing image features obtained after a convolution operation. It contains various local feature information of the image, such as edge contours and texture details, providing a basic feature representation for subsequent linear attention mechanism processing. The value at each position represents the intensity of a certain feature in the corresponding image region. The image width (W) and height (H) are the number of pixels in the horizontal and vertical directions of the image, and the number of convolution channels (C) usually corresponds to the number of convolution kernels, representing the dimensions of different features.

[0038] Specifically, taking a grayscale image to be restored with a size of 256×256 pixels as an example, a 3×3 convolution kernel is used for convolution operation: First, place the convolution kernel at the upper left corner of the image and perform element-wise multiplication with the 3×3 pixel values at the corresponding positions of the image, and then sum all the products to obtain the first output value. Then, the convolution kernel moves one pixel to the right and repeats the above operation at the new position until it covers the entire image width. Similarly, the convolution kernel moves one pixel down and continues the convolution operation until it covers the entire image height. During this process, each convolution kernel corresponds to one channel. Assuming 16 different convolution kernels are used, the size of the first feature map obtained is 254×254×16 (because the convolution operation reduces the width and height of the image, assuming no padding operation here).

[0039] Extracting the first feature map through the convolution operation in preprocessing can effectively extract the local features of the image, highlight the key structures and texture information in the image, and provide an important feature basis for subsequent image restoration tasks; at the same time, it can also reduce the scale of the data, lower the complexity of model processing, while retaining the key feature information, which helps to improve the running efficiency and restoration performance of the model.

[0040] S300: Perform window-dimension attention establishment and channel-dimension attention establishment on the first feature map based on the linear attention mechanism, and obtain the fused attention output feature.

[0041] Specifically, the linear attention mechanism efficiently calculates the attention weights through linear transformation. Compared with the traditional dot-product-based attention mechanism, it can reduce the computational complexity.

[0042] Among them, window-dimension attention establishment refers to constructing an attention relationship in the spatial dimension of the feature map (i.e., the local region of the image, with windows as units), highlighting the importance of local features in restoration. Channel-dimension attention establishment focuses on the channel direction of the feature map, judges the contribution degree of different channel features to the image restoration task, and enhances the expression of key channel features.

[0043] Specifically, the obtained fused attention output feature refers to integrating the attention calculation results of the window dimension and the channel dimension to form a feature representation that combines spatial and channel information for subsequent image restoration processing.

[0044] In some embodiments, as Figure 2 shown, performing window-dimension attention establishment and channel-dimension attention establishment on the first feature map based on the linear attention mechanism, and obtaining the fused attention output feature includes: Perform window partitioning on the first feature map to obtain a set of sub-feature maps, and calculate the sub-window attention features for each sub-feature map, and integrate them to obtain the window attention feature.

[0045] Partition the first feature map into a set of channel sub-blocks in the channel dimension, and calculate the sub-channel attention features for each channel sub-block, and integrate them to obtain the channel attention feature.

[0046] Calculate the fused attention feature map according to the first feature map, the window attention feature, and the channel attention feature.

[0047] Specifically, first, the first feature map is window-divided to obtain a set of sub-feature maps. Assume the size of the first feature map is 64×64×32. The image is divided into windows of 4×4, so the size of each sub-feature map is 4×4×32, and the entire set of sub-feature maps has (64 / 4)×(64 / 4)=16×16 sub-feature maps. Then, for each sub-feature map, the sub-window attention feature is calculated. This can be achieved by obtaining query, key, and value vectors through linear transformation, then calculating the attention weights using the linear attention formula, and finally combining the value vectors with the weights to obtain the attention feature of each sub-window. Integrating these sub-window features yields the window attention feature, which is of size 64×64×1 (W×H). Meanwhile, in the channel dimension, the 32 channels of the first feature map are divided into 4 channel sub-blocks, each containing 8 channels, forming a set of channel sub-blocks. The global feature of each channel sub-block is calculated, such as obtaining the representative feature of the channel sub-block through average pooling or max pooling, then obtaining the sub-channel attention feature through operations such as linear transformation and activation function, and finally integrating these sub-channel features to obtain the channel attention feature, whose size is 64×64×32.

[0048] Furthermore, the fused attention feature map is calculated based on the first feature map, the window attention feature, and the channel attention feature. This can be done by concatenating the three features. For example, concatenating them in the channel dimension to form a feature of 64×64×(32 + 32 + 1)=64×64×65. This fusion method enables the feature map to contain both the original features, spatial attention information, and channel attention information simultaneously, enhancing the expressive power of the features.

[0049] Through the above step-by-step attention calculation and fusion process in the window and channel dimensions, the following technical effects can be achieved: First, window division and sub-window attention calculation can capture the pixel relationships within the local regions of the image, highlighting local features and improving the model's ability to process image details; Second, channel division and sub-channel attention calculation can effectively judge the importance of different channel features, enhance the feature expression of key channels, and reduce the interference of redundant information; Third, fusing the attention features in the two dimensions of window and channel with the original features can form a feature representation that combines spatial and channel information, making the feature map more discriminative and targeted, thereby improving the performance of the image restoration model and the quality of the restored image. At the same time, the application of the linear attention mechanism also reduces the computational complexity and improves the running efficiency of the model.

[0050] In some embodiments, obtaining the window attention feature includes: Based on the sliding window method, the first feature map is divided into windows of size M×M to form the set of sub-feature maps.

[0051] Perform a linear transformation on each sub - feature map to generate a first intermediate feature map, where the size of the intermediate feature map is \(W\times H\times3C\).

[0052] Trisect the first intermediate feature map in the channel dimension, and based on the linear attention mechanism, calculate the first sub - linear attention feature by combining the trisecting results, and further calculate the window attention feature.

[0053] In some implementation manners, performing a linear transformation on each sub - feature map includes: For each of the sub - feature maps, execute the linear transformation of, where is the weight matrix, is the bias vector, is the first intermediate feature map, and .

[0054] Specifically, first, based on the sliding window strategy, divide the input first feature map into several local windows of size M × M to form a sub - feature map set , where N is the number of windows. Then, perform a linear transformation on each sub - feature map to obtain the first intermediate feature map , and its calculation method is: ; where, is the weight matrix, is the bias vector.

[0055] Specifically, then, trisect the first intermediate feature map in the channel dimension, and use them as the Query, Key, and Value inputs in the attention mechanism respectively to calculate the linear attention.

[0056] Specifically, further, fuse the linear attention result and the trisecting result to calculate the multiple first sub - linear attention features corresponding to multiple sub - feature maps in the sub - feature map set.

[0057] In some implementation manners, the first sub - linear attention feature is calculated by the operation formula: ; where, represents the \(i\) - th first sub - linear attention feature. , , are weight matrices. , , is the bias vector, Compute the result for the i-th linear attention mechanism.

[0058] Specifically, finally, the obtained multiple first sub-linear attention features are concatenated to obtain the window attention feature, where the window attention feature is size.

[0059] In some embodiments, obtaining a channel attention feature includes: The first feature map is divided into L sub-blocks in the channel dimension, each sub-block includes C / L groups of features, and the L sub-blocks generate a channel sub-block set.

[0060] Perform a linear transformation on each channel sub-block to generate the second intermediate feature map.

[0061] The second intermediate feature map is divided into three equal parts in the channel dimension, and based on the linear attention mechanism, the second sub-linear attention feature is calculated in combination with the three-division result, and the channel attention feature is further calculated.

[0062] Specifically, first, the first feature map Divided into two parts in the channel dimension: L channel sub-blocks, each of which contains C / L channels, forming a channel sub-block set; illustratively, the channel sub-block set is recorded as ; Then, for each channel sub-block Perform a linear transformation to obtain the second intermediate feature map .

[0063] In some implementations, performing a linear transformation on each channel sub-block includes: For each of the sub-feature graphs, perform The linear transformation of is the weight matrix, is the bias vector, is the second intermediate feature map, and .

[0064] Specifically, further, the second intermediate feature map , divided into three equal parts in the channel dimension, as Query, Key, and Value respectively, and through the same method and principle as the above calculation of the first sub-linear attention feature, multiple second sub-linear attention features are calculated based on the linear attention mechanism.

[0065] In some implementations, calculating the channel attention feature includes: The second sub-linear attention features output by each of the channel sub-blocks are concatenated to obtain the channel attention features with the number of channels restored to C.

[0066] Specifically, finally, the multiple second sub-linear attention features output by multiple channel sub-blocks are concatenated, so that the channel dimension is restored to the original number of channels C, and a channel attention feature map is obtained.

[0067] Optionally, the original first feature map, the window attention feature map, and the channel attention feature map are fused to obtain a fused attention feature map, where the fusion method includes but is not limited to weighted summation, convolution after concatenation, attention weighted fusion, etc.

[0068] Through the above dual linear attention modeling of the window dimension and the channel dimension, the local spatial relationship and the inter-channel dependence of the image features can be captured more effectively; the feature expression ability can be improved, and the model's attention ability to the target area can be enhanced; at the same time, compared with the traditional attention mechanism, the linear attention mechanism has higher computational efficiency and scalability; it is applicable to various computer vision task scenarios such as image recognition, object detection, and image segmentation.

[0069] S400: Repeat K times to obtain K of the fused attention output features, and the output is the final feature map.

[0070] Specifically, repeating K times here means that the processing module based on the linear attention mechanism (including the establishment of attention in the window dimension and the channel dimension and the fusion process) is cyclically applied K times in the feature map processing flow. Each cycle will generate a new fused attention output feature, and these features will sequentially transmit and accumulate information.

[0071] Exemplarily, execute K times the above fused attention calculation process, and obtain K fused attention output feature maps respectively. Then, these K feature maps can be cascaded, weighted summed, or other integration methods to generate the final feature map. The final feature map integrates the results of multiple attention calculations, contains richer and deeper feature information, and is used for subsequent image restoration processing.

[0072] S500: Perform a post-processing operation based on convolution on the final feature map, and perform a residual connection between the convolution result and the image to be restored to obtain the restored image.

[0073] Specifically, the post-processing operation based on convolution refers to applying a convolutional layer on the final feature map to further extract and transform features to generate a representation closer to the target image. The residual connection is an operation of directly adding the previous feature to the subsequent feature, which is used to retain the original information and alleviate the problem of gradient disappearance. The restored image is the final output image after convolution post-processing and residual connection, aiming to be as close as possible to the target image and restore the quality of the original image.

[0074] Exemplarily, first, a convolution operation is performed on the final feature map (assuming a size of 64×64×32). For example, a 3×3 convolution kernel is used. After convolution, the size of the feature map may become 62×62×16 (assuming the number of output channels is 16). Then, a residual connection is made between this convolution result and the original image to be restored (assuming a size of 64×64×1 if it is a grayscale image). Since the sizes may not be consistent, it may be necessary to perform interpolation or other operations on the convolution result to make its size consistent with the image to be restored. For example, the convolution result is upsampled to 64×64×1, and then the convolution result is added to the image to be restored to obtain the restored image.

[0075] The above-mentioned residual connection method allows the model to learn the residual information between the restored image and the image to be restored, that is, the part that needs to be corrected, rather than directly learning the complete image, making the learning process more efficient and easier to converge.

[0076] S600: Construct a loss function based on the target image, and use the gradient descent method to optimize and update the parameters of the image restoration model.

[0077] Feasibly, the loss function may include but is not limited to one or more of the following: pixel-level loss (such as L1, L2 loss); perceptual loss, based on the feature differences extracted by a pre-trained network (such as VGG); adversarial loss (if using a GAN structure) to enhance the realism of the image; structural similarity loss (SSIM) to maintain the structural consistency of the image.

[0078] In some embodiments, the loss function is defined as: ; where, represents the restored image, is the target image, and N is the number of training samples.

[0079] In some embodiments, the linear attention function in the linear attention mechanism is defined as: ; where, characterizes the linear attention feature of the nth set element; , , represents taking the dth power of bit by bit; , , are the results of trisecting.

[0080] Specifically, , , The result of trisection respectively corresponds to the above-mentioned Query, Key, and Value; t is the original input vector (for example , a certain vector element in ; the vector t is the intermediate vector result obtained after the input vector ReLU passes through the activation function; represents the Euclidean norm of the vector and is used for scaling; represents taking the th power of each element in the vector d bit by bit; d is a positive integer hyperparameter used to control the non-linearity degree of feature mapping.

[0081] The above-mentioned linear attention function enhances the non-linear expression ability of feature mapping by introducing bit-by-bit power transformation and normalization operation. Compared with the traditional linear attention mechanism, this method improves the model's ability to model high-order features while maintaining a low computational complexity.

[0082] Exemplarily, tests were conducted on the image deraining datasets Rain200L and Rain200H, the image deblurring dataset HIDE, the image dehazing dataset OTS, and the image desnowing dataset CSD. The results are shown in the following table. Taking PSNR and SSIM as evaluation metrics, the higher the value, the better the image restoration effect. The number of parameters and Flops reflect the size and computational amount of the model, and the smaller the value, the better.

[0083] Table 1 Performance results on different datasets exemplarily

[0084] It can be seen from the results that although the PSNR and SSIM of this method are slightly reduced, the number of parameters and Flops are much smaller than those of other methods, proving that this method is more suitable for running on embedded devices.

[0085] In summary, the image restoration method integrating a linear attention mechanism provided by the present invention has the following technical effects: By obtaining an image sample set, forming a sample pair of the image to be restored and the corresponding target image, and inputting it into an image restoration model, where the output of the image restoration model is the restored image; preprocessing the image to be restored, and using a convolution operation to extract a first feature map, where the size of the first feature map is W××H×C, W and H are the image width and height respectively, C is the number of convolution channels, and W, H, and C are all positive integers greater than 0; establishing window - dimension attention and channel - dimension attention for the first feature map based on a linear attention mechanism, and obtaining a fused attention output feature; repeating K times to obtain K fused attention output features, and the output is the final feature map; performing post - processing operations based on convolution on the final feature map, and performing residual connection between the convolution result and the image to be restored to obtain the restored image; constructing a loss function based on the target image, and using the gradient descent method to optimize and update the parameters of the image restoration model. Furthermore, the technical effects of reducing the computational complexity, improving the model efficiency, enhancing the model applicability and performance, and at the same time maintaining or improving the image restoration quality are achieved.

[0086] It should be understood that the disclosed embodiments of the present invention and the above descriptions enable those skilled in the art to implement the present invention using the present invention. At the same time, the present invention is not limited to the part of the embodiments mentioned above. It should be understood that ordinary skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. An image restoration method integrating a linear attention mechanism, characterized in that, Including: Obtain an image sample set, form a sample pair by combining the image to be restored and the corresponding target image, and input it into the image restoration model, where the output of the image restoration model is the restored image; Preprocess the image to be restored, and extract a first feature map using a convolution operation, where the size of the first feature map is W×H×C, W and H are the image width and height respectively, C is the number of convolution channels, and W, H, and C are all positive integers greater than 0; Based on the linear attention mechanism, establish window-dimensional attention and channel-dimensional attention for the first feature map, and obtain a fused attention output feature; Repeat K times to obtain K fused attention output features, and the output is the final feature map; Perform post-processing operations based on convolution on the final feature map, and perform residual connection on the convolution result and the image to be restored to obtain the restored image; Construct a loss function based on the target image, and use the gradient descent method to optimize and update the parameters of the image restoration model.

2. The image restoration method integrating a linear attention mechanism according to claim 1, wherein Based on the linear attention mechanism, establish window-dimensional attention and channel-dimensional attention for the first feature map, and obtain a fused attention output feature, including: Perform window partitioning on the first feature map to obtain a set of sub-feature maps, calculate sub-window attention features for each sub-feature map, and integrate them to obtain window attention features; Partition the first feature map into a set of channel sub-blocks in the channel dimension, calculate sub-channel attention features for each channel sub-block, and integrate them to obtain channel attention features; Calculate the fused attention feature map according to the first feature map, the window attention feature, and the channel attention feature.

3. The image restoration method integrating a linear attention mechanism according to claim 2, wherein Obtain the window attention feature, including: Based on the sliding window method, partition the first feature map into windows of size M×M to form the set of sub-feature maps; Perform a linear transformation on each sub-feature map to generate a first intermediate feature map, where the size of the intermediate feature map is W×H×3C; Trisect the first intermediate feature map in the channel dimension, and based on the linear attention mechanism, calculate the first sub-linear attention feature in combination with the trisection result, and further calculate the window attention feature.

4. The image restoration method integrating a linear attention mechanism according to claim 2, characterized in that Obtain the channel attention feature, including: Partition the first feature map into L sub-blocks in the channel dimension, each sub-block includes C / L groups of features, and the L sub-blocks generate a set of channel sub-blocks; Perform a linear transformation on each channel sub-block to generate a second intermediate feature map; Trisect the second intermediate feature map in the channel dimension, and based on the linear attention mechanism, calculate the second sub-linear attention feature in combination with the trisection result, and further calculate the channel attention feature.

5. The image restoration method integrating a linear attention mechanism according to claim 3, wherein, Performing a linear transformation on each sub-feature map includes: For each of the sub-feature maps, perform a linear transformation, where is a weight matrix, is a bias vector, is a first intermediate feature map, and .

6. The image restoration method integrating a linear attention mechanism according to claim 4, wherein Performing a linear transformation on each channel sub-block includes: For each of the sub-feature maps, perform a linear transformation, where is a weight matrix, is a bias vector, is the second intermediate feature map, and .

7. The image restoration method integrating a linear attention mechanism according to claim 1, wherein, The linear attention function in the linear attention mechanism is defined as: ; Among them, characterize the linear attention features of the nth set element; , , denote taking the d-th power of bit by bit; , , are the trisection results.

8. The image restoration method integrating a linear attention mechanism according to claim 3, wherein Calculate the first sub-linear attention feature, and the operation formula is: ; Among them, represents the i-th first sub-linear attention feature; , , are weight matrices; , , are bias vectors, is the calculation result of the i-th linear attention mechanism.

9. The image restoration method integrating a linear attention mechanism according to claim 4, characterized in that Calculating the channel attention feature includes: Concatenate the second sub-linear attention features output by each channel sub-block to obtain the channel attention feature with the number of channels restored to C.

10. The image restoration method integrating a linear attention mechanism according to claim 1, characterized in that The loss function is defined as: ; Among them, represents the restored image, is the target image, and N is the number of training samples.

Citation Information

Patent Citations

  • Image super-resolution reconstruction method based on fused attention mechanism residual network

    CN111192200A

  • Neural network image defogging method based on multi-level feature fusion and attention guidance

    CN111915531A

  • Infrared-visible light fused deep neural network and modeling method thereof

    CN111985625A

  • Image restoration method and system based on multiple attention mechanisms

    CN119130863A

  • Image reconstruction method based on multi-window cross feature fusion attention mechanism

    CN119168860A

Cited By

  • Image correction and recovery method and system based on image-text multi-mode large model

    CN121616476A