Image Reconstruction Method Based on Multi-Window Cross-Feature Fusion Attention Mechanism

The method addresses computational inefficiencies and local attention issues in Transformer-based image super-resolution by employing multiple window cross-attention fusion and attention mechanisms, enhancing feature extraction and fusion to improve image reconstruction precision and efficiency.

CN119168860BActive Publication Date: 2025-07-15HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411115445.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-14
Publication Date
2025-07-15
Estimated Expiration
2044-08-14

AI Technical Summary

Technical Problem

Existing image super-resolution methods have high demand for computing resources and insufficient attention to local areas, resulting in performance degradation.

Method used

The multi-window cross-feature fusion attention mechanism is adopted, and the traditional window attention mechanism is further divided into overlapping small windows of multiple shapes, combining the spatial and channel attention mechanism, reducing the calculation amount and improving the feature extraction accuracy.

Benefits of technology

While reducing the computational amount, it improves the accuracy of image super-resolution and the lightweight of the model, which can better capture local and global features of the image and generate high-quality super-resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119168860B_ABST
    Figure CN119168860B_ABST
Patent Text Reader

Abstract

The present invention discloses an image reconstruction method based on a multi-window cross-feature fusion attention mechanism. The method reduces the resolution of the original image through bicubic interpolation to obtain model training data, and uses the original image as a sample label. For the training data, shallow features are first extracted through blueprint separable convolution. Then, the efficient separable distillation module is improved to obtain multiple attention fusion distillation modules for extracting deep features. A feature fusion module based on multi-window cross-attention is used to perform operations such as multiple segmentations, attention calculations, multi-window cross-attention calculations, and recombinations on the deep features to achieve feature enhancement. Then, it is fused with the shallow features to output a reconstructed high-resolution image. According to the different downsampling ratios of the training data, models with different reconstruction ratios can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and relates to image feature extraction and image super-resolution technology, and specifically relates to an image reconstruction method based on a multi-window cross-feature fusion attention mechanism. Background Art

[0002] In traditional image processing, due to the limitations of device resolution and acquisition conditions, images are often blurred and distorted. To improve image clarity and quality, super-resolution technology is introduced, and its goal is to restore low-resolution and degraded images to corresponding high-resolution images. Super-resolution technology can not only improve image quality and enhance people's visual experience, but also assist the object detection task of the model and improve its accuracy. In addition, super-resolution technology can achieve the same effect while reducing costs, break through the limitations of optical imaging systems, and obtain high-resolution images. With the continuous innovation of digital image processing technology, super-resolution has become one of the research directions that has attracted much attention in the field of image processing, and image super-resolution technology is widely used in fields such as video surveillance and medical imaging.

[0003] In the process of image super-resolution, it is necessary to extract and fuse deep features, and while maintaining the accuracy of the original information of the image, generate a reconstructed image close to the original high-resolution image according to the extracted features. A relatively common image super-resolution method is the Transformer model using the attention mechanism, which captures the global information of the image according to the image features and can better understand the relationship between different regions in the image, especially tasks involving long-distance relationships. However, the defects of these methods are: one is that a large amount of computing resources are required, and the other is that the attention to local regions is insufficient, which reduces the performance of image super-resolution. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention proposes an image reconstruction method based on a multi-window cross-feature fusion attention mechanism. The square window of the traditional window attention mechanism is further divided into overlapping small windows of various shapes, which not only fully considers the context information of local regions, but also further reduces the amount of calculation, thereby realizing the lightweight and high-precision of the model.

[0005] The image reconstruction method based on the multi-window cross-feature fusion attention mechanism specifically includes the following steps:

[0006] Step 1: Data collection and processing

[0007] Step 1.1: Collect high-resolution images of different scenes, including building images, face images, comic images, etc., reduce the image resolution by the method of bicubic interpolation, and form an image pair with the original high-resolution image to obtain a diverse data set.

[0008] Step 1.2: Crop the image pairs obtained in Step 1.1 into multiple slices of the same size to achieve data augmentation and dataset expansion, and then convert the slices into Tensor vectors. Use the high-resolution image in an image pair as the label and the low-resolution image as the training sample.

[0009] Step 2: Shallow feature extraction

[0010] First, extract the shallow features of the sample images. The specific steps are as follows:

[0011] Step 2.1: Expand the channels of the Tensor vectors of the training samples in Step 1.2, copy and expand the number of channels from 3 channels to 12 channels to obtain the corresponding feature image tensors.

[0012] Step 2.2: Use Blueprint Separable Convolutions (BSConv) to extract features from the feature image tensors in Step 2.1. First, through the depth convolution layer, perform independent convolution operations on 12 channels, independently calculate the convolution results to obtain a tensor of the same size as the input, and then through the pointwise convolution layer, use a 1*1 convolution kernel to integrate the information of each channel of the tensor to obtain the shallow feature F in , and the shallow feature preserves the information of the image mapping from low-dimensional features to high-dimensional features.

[0013] Step 3: Deep feature extraction

[0014] After Steps 1 and 2, the shallow feature F in the low-resolution image is initially extracted in , and by improving the Effcient Separable Distillation Block (ESDB), multiple Attention Fuse Distillation Blocks (AFDB) are obtained for extracting deep information. The attention fusion distillation module includes a distillation layer, a condensation layer, and an enhancement layer. The specific steps are as follows:

[0015] Step 3.1: Feature refinement

[0016] Send the shallow feature F obtained in Step 2 in into the distillation layer. The distillation layer includes L levels, and each level includes two branches for extracting and integrating the information learned by the model.

[0017] In the l-th layer, the first branch passes through an ordinary convolution layer and a Gaussian Error Linear Unit (GELU) activation function layer to learn complex patterns and features to obtain Fd1 The second branch passes through the blueprint separable convolutional layer, GELU activation layer and residual connection of the Blueprint Separable Residual Block (BSRB) to introduce deeper feature representations while maintaining the original information, resulting in F r1 :

[0018] (1)

[0019] (2)

[0020] (3)

[0021] (4)

[0022] Among them, , when l = 1, is the shallow feature F in , is the feature map of the input activation layer.

[0023] Step 3.2, Feature Aggregation

[0024] Concatenate the L outputs of the distillation layer on the channel dimension to obtain the concatenated feature map F ct Then, feature aggregation is achieved through a convolutional layer to obtain F con :

[0025] (5)

[0026] Among them, CONCAT( ) represents channel concatenation, and Conv1( ) represents a 1*1 convolutional layer.

[0027] Step 3.3, Feature Enhancement

[0028] The enhancement layer performs feature enhancement on the feature map obtained in Step 3.2 through three attention layers. By adjusting the structures of the three attention layers, the three attention layers can compensate for each other's defects to obtain accurate image features.

[0029] First, input the feature map F con into the ECA (Efficient Channel Attention) attention module to focus on the information in the channel dimension, consider the correlation between different channels, and improve the efficiency of the network in learning specific features. First, obtain the global information of each channel through one-dimensional global average pooling (GAP) , and then obtain the learnable weights of each channel through a fully connected layer The output F is calculated based on the weights of each channel eca :

[0030] (6)

[0031] (7)

[0032] Where 、 、 are the length, width, and number of channels of the feature map F con respectively, represents the pixel value at the position (i, j) in the c-th channel of the feature map F con , σ represents the activation function, represents the convolution operation, represents the pixel value at the position (i, j) in the c-th channel of the output feature map F eca .

[0033] Meanwhile, the feature map F con is input into the ESA (Efficient Spatial Attention) attention module. This module mainly focuses on the information in the spatial dimension, improves the perception ability of local information by learning the relationships between positions, and obtains the output F esa :

[0034] (8)

[0035] (9)

[0036] Where represents the spatial information at the position (i, j) in the feature map F con , is the weight matrix obtained through learning, represents the pixel value at the position (i, j) in the c-th channel of the output feature map F esa .

[0037] Finally, F eca and F esa are concatenated on the channel dimension and then input into the CCA (Compact Channel Attention) attention module. The channel relationships are modeled using low-rank matrix factorization, reducing the number of parameters and computational complexity, and obtaining the output F cca . Finally, another feature enhancement output F is obtained through a 1*1 convolution eo1 .

[0038] Step 3.4: Use the enhanced output F eo1 as the input of the next attention fusion distillation module. After repeating this process multiple times, concatenate the output features of each attention fusion distillation module along the channel dimension, then output the integrated aggregated features through a convolutional layer, and finally obtain the deep features F in the low-resolution image through GELU activation ac .

[0039] Step 4: Multi-window cross-attention

[0040] Design a feature fusion module MWT based on multi-window cross-attention to further enhance the deep features F ac .

[0041] Step 4.1: First, input the deep features F ac into the ESA attention module, and divide the feature map output by the ESA attention module into a group of non-overlapping feature maps with c channels, size h*w each

[0042] Step 4.2: Input the segmented feature maps into the MHA (Multi-Head Attention) module to obtain the attention map, and then generate three tensors Q (Query), K (Key), and V (Value) through a linear layer

[0043] Step 4.3: Restore the tensors Q, K, and V output by the linear layer to the image shape, changing the dimension from (b,n,h*w,c / / n) to (b,n,h,w,c / / n), where b is the batch size and n is the number of attention heads of the multi-head attention module (MHA). Perform the following three window slicing operations on the image shapes restored from the tensors Q, K, and V simultaneously

[0044] ① Divide the image evenly into t blocks along the vertical direction to obtain t blocks of (Q h , K h , V h )

[0045] ② Divide the image evenly into t blocks along the horizontal direction to obtain t blocks of (Q w , K w , V w )

[0046] ③ Divide the whole image evenly into t blocks to obtain t blocks of (Q c , K c , V c )

[0047] Perform position encoding on each sliced block, add the result of position encoding to the corresponding block, and obtain 3×t groups of (q, k, v). Calculate the attention value O for each block

[0048] (10)

[0049] Among them, d k is the dimension size of k after adding the position encoding.

[0050] Step 4.4: Then, in the way of window slicing, reorganize the attention calculation results to obtain the attention feature maps of three windows. After merging these three attention feature maps on the channels, then integrate the feature information on the channels of the three windows through a convolution once to obtain an output feature map of size h*w. Reorganize multiple output feature maps at the corresponding positions of the windows to obtain an output with the same size as the deep feature F ac and then perform the window sliding (Shift Window, SWIND) operation.

[0051] Step 4.5: Repeat Steps 4.2 to 4.4, and then pass through an ESA attention module to complete the feature enhancement part to obtain the output F wo . After performing a blueprint separable convolution on the output F wo , perform a residual connection with the shallow feature F in and then output a high-resolution image after upsampling.

[0052] Step 5: Image super-resolution

[0053] Set the parameters of bicubic interpolation and train models at different magnification ratios. Select the model with the corresponding magnification ratio according to the requirements, input the low-resolution image into the model, and obtain the super-resolved image through the feature extraction and enhancement in Steps 2 to 4.

[0054] The present invention has the following beneficial effects:

[0055] 1. During the deep image feature extraction process, in order to capture the image features in the spatial and channel dimensions and their associated features, based on the spatial attention mechanism and the channel attention mechanism, a new feature enhancement method is introduced. By dispersing the information in the ESA channel dimension and the information in the ECA spatial dimension to each channel, and then using the channel attention mechanism CCA to integrate and enhance the information, by adjusting the positions of the module structures, the advantages of efficient capture of features in the space and channels of each module can be effectively utilized, and at the same time, the deficiencies such as the lack of association of image features in the channels and space of other structural attention modules are solved, further enhancing the feature information contained in the feature image after feature extraction.

[0056] 2. During the feature fusion process, in order to clearly identify the features of each part and obtain complete and important feature information, an improved feature fusion module based on multi-window cross is proposed. First, the attention of features on different channels is adjusted through channel attention. Then, the feature image is divided into windows, and each window is divided into three parts: vertical, horizontal, and rectangular to calculate the attention. After fusing the three parts, the feature enhancement of the feature fusion part is achieved through the channel attention module, so that it can be noticed which features are more important during the process of upsampling the image. Because the multi-window mechanism is introduced, the computational amount is dispersed to each window, and the features of multiple windows with different shapes can be captured. Therefore, while reducing the computational amount, the accuracy of the model super-resolution is improved, providing an important reference for the generation of the final super-resolution image. Description of the Drawings

[0057] Figure 1 Flowchart of the image reconstruction method based on the multi-window cross feature fusion attention mechanism;

[0058] Figure 2 Original high-resolution image used in the embodiment;

[0059] Figure 3 Low-resolution image obtained by bicubic interpolation in the embodiment;

[0060] Figure 4 Schematic diagram of the image reconstruction model based on the multi-window cross feature fusion attention mechanism;

[0061] Figure 5 Schematic diagram of the attention fusion and distillation module structure;

[0062] Figure 6 Schematic diagram of the multi-window cross feature fusion attention mechanism;

[0063] Figure 7 Schematic diagram of the multi-window cross MWA module;

[0064] Figure 8 Super-resolution image output in the embodiment. Detailed Implementation Manner

[0065] The present invention will be further explained below with reference to the accompanying drawings;

[0066] As Figure 1 shown, the image reconstruction method based on the multi-window cross feature fusion attention mechanism is as follows:

[0067] Step 1. Data collection and processing

[0068] This embodiment uses 2,650 images from Flickr2K and 800 images from DIV2K as the training data set. For Figure 2 the high-resolution images with a size of 256×256 shown, they are downsampled by a factor of 3 using bicubic interpolation to obtain images with a size of 85×85, and then bicubic interpolation is performed again to restore the size to 256×256, obtaining the low-resolution images as shown in Figure 3 . The two 256×256 images in the image pair are cut into 32 64×64 images, forming 16 groups of image pairs to increase the generalization of the model training images. The cut low-resolution images are used as the model training samples, and the high-resolution images are used as the corresponding sample labels.

[0069] Step 2: Shallow feature extraction

[0070] As shown in Figure 4 , the training samples are converted from 64×64 images into tensors with a size of (1, 3, 64, 64), and they are copied to 12 channels and merged and expanded on the channels using Concat to obtain tensors with a size of (1, 12, 64, 64), realizing the enhancement of the color information of the low-resolution images and the capture of the details of the images to reduce the loss of image information. Finally, shallow feature extraction is performed through a blueprint separable convolutional layer to increase the model complexity and enhance the feature representation, obtaining a shallow feature map F with a size of (1, 64, 64, 64) in .

[0071] Step 3: Deep image feature extraction

[0072] For the shallow feature map F in obtained in step 2, it passes through 8 consecutive attention fusion and distillation modules AFDB, and then the results of the 8 times are concatenated in the channel dimension to obtain a deep feature image F with a size of (1, 512, 64, 64) ac .

[0073] As shown in Figure 5 , the attention fusion and distillation module AFDB includes 3 layers of distillation layers, a condensation layer, and an enhancement layer. Each layer of the distillation layer includes 2 branches. The first branch passes through an ordinary convolutional layer and a Gaussian error activation function layer to output a feature map F d1 . The second branch outputs a feature map F r1 through a blueprint shallow residual module and inputs it into the two branches of the next layer. Each layer can continuously refine the features of the feature image. The first layer will reduce the number of channels of the feature image by half, obtaining a feature map with a size of (1, 32, 64, 64). The sizes of the feature maps output by the latter two layers will not change, both being (1, 32, 64, 64). The output feature map F r3After passing through a convolutional layer of size 3*3, it is merged with the feature maps F obtained from the previous three layers d1 to obtain a feature map tensor of size (1, 128, 64, 64). Then, through a common convolutional layer, the number of channels is restored to a feature image of size (1, 64, 64, 64) to complete feature condensation. Finally, in the enhancement layer, deep feature extraction of the image is completed by enhancing the features in the spatial and channel dimensions through three attention layers, ECA, ESA, and CCA. These three attention layers do not change the size of the image

[0074] Step 4: Image Feature Fusion and Reconstruction

[0075] As Figure 6 shown, the deep feature image F ac is input into the convolutional layer and the ESA attention layer for processing to obtain a feature image of size (1, 64, 64, 64). For this feature map, a multi-window cross-feature fusion attention mechanism is adopted. First, the feature map is divided into 4 windows with a length and width of 16 according to the window to obtain 4 feature maps of size (1, 64, 16, 16). After each divided window passes through the MHA and linear modules, it is sent to the MWA (Multi-Window Attention). As Figure 7 shown, in the MWA, by splitting q, k, and v, each window is respectively divided into 4 windows of size (8, 8), 4 windows of size (4, 16), and 4 windows of size (16, 4). Finally, the attention in the three directions is fused to obtain a feature image of size (1, 64, 64, 64) after one-time attention. To enable the attention to notice adjacent features, the window is translated, and then the translated image is processed by the multi-window cross-attention mechanism, and then the final feature image is obtained through one ESA attention layer. Finally, it is connected to the initial low-resolution image by residual connection and then upsampled once to obtain the reconstructed high-resolution image, as Figure 8 shown

[0076] Calculate the computational complexity of dividing using a fixed-size window and dividing using the multi-window cross-feature fusion in this method 、 :

[0077] (12)

[0078] (13)

[0079] where h, w, and c are the height, width, and number of channels of the window, and n is the number of heads of the multi-head attention. Obviously, in this embodiment, the computational complexity of this method is reduced by 1 / 4 compared with the traditional fixed-size window.

[0080] Set the loss function as:

[0081] (14)

[0082] Set the learning rate to 1×10 -3 , set betas to 0.9 and 0.999, and set the parameter to decrease by 1×10 every 100 rounds -5 , and train for a total of 500,000 rounds, saving the model with the best training effect.

[0083] To verify the enhancement effect of the proposed attention mechanism of multi-window cross-feature fusion on the model, relevant comparative experiments were conducted. The specific experimental results are shown in the following table:

[0084]

[0085] The super-resolution magnification is achieved by changing the downsampling multiple of bicubic interpolation. According to the experimental results, the PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity) indicators of this method under different super-resolution magnifications (×2, ×3, ×4) are better than those of the baseline models BSRN, RFDN, and LAPAR-A. Especially at the ×2 and ×3 magnifications, the improvement of this method relative to the baseline model is obvious and more prominent. In addition, this method shows high PSNR and SSIM values on different datasets (Set5, Set14, BSD100, Urban100, Manga109), indicating that it can achieve significant improvements in various scenarios and image types, especially performing very well on the Urban100 and Manga109 datasets.

[0086] On the Urban100 dataset, for the three-fold super-resolution model obtained by this method, the PSNR is increased by 1.48dB. At the same magnification, Manga109 is increased by 0.6dB. The two datasets are composed of 100 urban images and 109 comic images respectively. Among them, the images in the Manga109 dataset have a certain artistic style, and the color information of the comics is richer than that of real-world images, indicating that the method can better handle color relationships and ensure that the images are not too distorted. The performance of this method on the Urban100 dataset shows that this method can effectively capture the relationship between houses in the city, handle the structural information of the houses well, and has an advantage in structure.

[0087] In addition, in terms of parameters and computational complexity, this method is relatively low, showing a high balance between performance and efficiency. Compared with RFDN and LAPAR-A, while improving performance, this method has relatively more reasonable resource requirements. From the consistency of PSNR and SSIM at each magnification, this method shows relatively stable performance and has good adaptability to images at different super-resolution magnifications.

[0088] The above experimental results can prove that this method has achieved comprehensive advantages in the super-resolution task, with better image reconstruction effects, a higher balance between performance and efficiency, broader adaptability, and better stability.

Claims

1. An image reconstruction method based on a multi-window cross-feature fusion attention mechanism, characterized in that: Specifically, it includes the following steps: Step 1: Data collection and processing Collect high-resolution images in different scenarios, reduce the image resolution by bicubic interpolation, and form an image pair with the original high-resolution image; convert the image into a Tensor vector, where the high-resolution image is used as the label and the low-resolution image is used as the training sample; Step 2: Shallow feature extraction Expand the Tensor vector of the training samples to 12 channels, and then input it into the blueprint separable convolution to obtain the shallow feature F in ; Step 3: Deep feature extraction Use the attention fusion distillation module to extract deep information; the attention fusion distillation module includes a distillation layer, a condensation layer, and an enhancement layer. The specific steps are as follows: Step 3.1: Feature refinement Feed the shallow features F obtained in step 2 in into the distillation layer, where the distillation layer consists of L levels, and each level includes two branches. In the l-th level, the first branch passes through a common convolutional layer and a Gaussian error activation function layer GELU to learn complex patterns and features to obtain F d1 ; the second branch passes through a blueprint separable convolutional layer, a GELU activation layer, and a residual connection to introduce deeper feature representations while maintaining the original information to obtain F r1 : F dl = DL l (F r(l-1) ) (1) F rl = RL l (F r(l-1) ) (2) where l = 1…L-1, when l = 1, F r(l-1) = F in , and x is the feature map input to the activation layer; Step 3.2: Feature condensation Concatenate the L outputs of the distillation layer on the channel dimension, and then achieve feature condensation through a convolutional layer to obtain the feature map F con ; Step 3.3: Feature enhancement Meanwhile, the feature map F con is input into the efficient channel attention module ECA and the efficient spatial attention module ESA, and then the feature maps F eca and F esa output by these two attention modules are concatenated on the channel dimension and then input into the compact channel attention module CCA to obtain the output feature map F cca . Finally, another feature enhancement output F eo1 is obtained through a 1×1 convolution operation; Step 3.4: Take the output F after the first feature enhancement eo1 as the input of the next attention fusion and distillation module. After repeating this multiple times, concatenate the output features of each attention fusion and distillation module along the channel dimension, then output the integrated aggregated features through a convolutional layer, and finally obtain the deep features F of the low-resolution image through GELU activation ac ; Step 4: Multi-window cross-attention Design a feature fusion module MWT based on multi-window cross-attention to further enhance the deep feature F according to the following steps ac : Step 4.1: Input the deep feature F ac into the efficient spatial attention module ESA, and divide the output feature map into a group of non-overlapping feature maps with c channels, size h*w each Step 4.2: Input the segmented feature map into the multi-head attention module MHA and the linear layer to generate three tensors Q, K, and V; Step 4.3: Restore the tensors Q, K, and V to the image shape, and then perform the following three window slicing operations: ①Divide it into t blocks evenly along the longitudinal direction of the image to obtain t blocks (Q h , K h , V h ); ②Divide it evenly into t blocks along the horizontal direction of the image, obtaining t blocks (Q w , K w , V w ); ③Divide the entire image evenly into t blocks to obtain t blocks (Q c , K c , V c ); Perform position encoding on each sliced block, add the result of the position encoding to the corresponding block, and obtain 3×t groups of (q, k, v). Calculate the attention value O for each block respectively: where d k is the dimension size of k after adding the positional encoding; Step 4.4: Reorganize the attention calculation results in the window slicing manner to obtain the attention feature maps of three windows. After merging these three attention feature maps in the channel dimension, further integrate the feature information on the channels of the three windows through one convolution operation to obtain an output feature map of size h*w. Reorganize multiple output feature maps at the corresponding positions of the windows to obtain the same output as the deep feature F, and then perform a window sliding operation. ac Same output, and then perform a window sliding operation; Step 4.5: Repeat steps 4.2 to 4.

4. After passing through an ESA attention module and one instance of blueprint separable convolution, perform a residual connection with the shallow feature F in and then perform an upsampling operation to output a high-resolution image; Step 5: Image super-resolution Set the parameters of bicubic interpolation, and train the model at different magnification ratios; select the model with the corresponding magnification ratio according to the requirement, input the low-resolution image into the model, and obtain the super-resolved image through the feature extraction and enhancement in Steps 2 to 4.

2. The image reconstruction method based on the multi-window cross-feature fusion attention mechanism according to claim 1, wherein: Crop the image pair into multiple slices of the same size to achieve data augmentation and dataset expansion.

3. The image reconstruction method based on the multi-window cross-feature fusion attention mechanism according to claim 1, wherein: The blueprint separable convolution first performs independent convolution operations on the input according to channels through the depth convolution layer, independently calculates the convolution results, and obtains a tensor of the same size as the input; then through the pointwise convolution layer, uses a 1*1 convolution kernel to integrate the information of each channel.

4. The image reconstruction method based on the multi-window cross-feature fusion attention mechanism according to claim 1, wherein: The efficient channel attention module ECA first obtains the global information g of each channel through one-dimensional global average pooling GAP, and then obtains the learnable weight W of each channel through a fully connected layer g , and calculates the output F according to the weight of each channel eca : Yc cij = X cij ·(σ(W g *g)) c (7) Where H, W, and c are the length, width, and number of channels of the feature map F con respectively, X cij represents the pixel value at the position (i, j) in the c-th channel of the feature map F con , σ represents the activation function, * represents the convolution operation, and Yc cij represents the pixel value at the position (i, j) in the c-th channel of the output feature map F eca .

5. The image reconstruction method based on the multi-window cross-feature fusion attention mechanism according to claim 1, wherein: The efficient spatial attention module ESA improves the perception ability of local information by learning the relationships between positions and obtains the output F esa : A ij = σ(W s * X ij ) (8) Ys cij = X cij · A ij (9) Where X ij represents the spatial information at the position (i, j) in the feature map F con , W s is the weight matrix obtained through learning, and Ys cij represents the pixel value at the position (i, j) in the c-th channel of the output feature map F esa .

6. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-lead electrocardiogram classification and identification method based on convolution and self-attention mechanism

    CN115470828A

  • Remote sensing image classification method based on CNN-self-attention mechanism hybrid architecture

    CN115641473A