Image super-resolution method
The convolutional layer and residual double multi-attention module can be separated through the blueprint to extract low-resolution image features, solving the problems of insufficient generalization capabilities and data dependence of deep learning models in the field of image super-resolution, and achieving high-quality image recovery and generalization capabilities.
Patent Information
- Application Number
- CN202510430125.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-08
AI Technical Summary
Existing deep learning models have insufficient generalization capabilities in the field of image super-resolution, and the training process requires a large amount of high-resolution image data, which is costly to obtain and label.
The blueprint can be used to separate the convolutional layer and extract the shallow features of low-resolution images, and the residual double multi-attention module RDMAB including channel attention, spatial attention and self-attention mechanisms are extracted to extract intermediate features and combine convolutional operations to reconstruct high-resolution images.
It improves the quality and generalization ability of image recovery, can effectively fuse features on images of different scales and types, and reduces the dependence on high-resolution image data.
Smart Images

Figure CN120278884A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more specifically, to an image super-resolution method. Background Art
[0002] In recent years, deep learning technology has made breakthrough progress in the field of super-resolution. Due to its powerful feature extraction and non-linear mapping capabilities, the convolutional neural network (CNN) has become the mainstream method for super-resolution research. For example, SRCNN (Super-Resolution Convolutional Neural Network) pioneered the use of CNN for super-resolution. Through end-to-end training, it directly learns the mapping from low-resolution images to high-resolution images, significantly improving the super-resolution effect. Many subsequent improved models such as FSRCNN (Fast Super-Resolution Convolutional Neural Network) further optimized the network structure and improved the computational efficiency.
[0003] In addition to CNN, generative adversarial networks (GANs) have also been introduced into the super-resolution task. Through the adversarial training of the generator and discriminator, GANs can generate more realistic and high-resolution images with rich details. For example, ESRGAN (Enhanced Super-Resolution Generative Adversarial Networks) introduced residual dense blocks in the generator and adopted a relative discriminator in the discriminator, further improving the visual quality of the images.
[0004] Although deep learning has achieved remarkable results in the field of super-resolution, it still faces some challenges. For example, the generalization ability of the model is insufficient, and the effect may vary greatly when processing different types of images; the training process requires a large amount of high-resolution image data, and the cost of data acquisition and annotation is relatively high. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide an image super-resolution method with sufficient generalization ability of the model.
[0006] To achieve the above purpose, the present invention provides the following technical solution: An image super-resolution method, comprising the following steps: Step 1, obtain a low-resolution LR image; Step 2, extract the shallow features of the low-resolution LR image through a blueprint separable convolutional layer; Step 3: The shallow features extracted in Step 2 are simultaneously input into the Residual Dual Multi-Attention Block (RDMAB) that includes channel attention, spatial attention, and self-attention mechanisms to extract intermediate features in each dimension. The intermediate feature information is input into a convolutional layer and smoothed to obtain deep features; Step 4: The shallow and deep feature information is fused through convolutional operations to reconstruct the low-resolution (LR) image into a high-resolution (HR) image, and the high-resolution image HR is output.
[0007] As a further improvement of the present invention, the specific steps of extracting the shallow features of the low-resolution LR image through the blueprint separable convolutional layer in Step 2 are as follows: Through the convolutional layer, the shallow features of the to-be-processed low-resolution LR image are extracted according to the following formula: where, represents the shallow features of the to-be-processed low-resolution image, represents the to-be-processed low-resolution image, is the operation of extracting the shallow features in using the blueprint separable convolution.
[0008] As a further improvement of the present invention, the specific steps of obtaining the deep features in Step 3 are as follows: Step 3-1: Calculate the channel attention feature AC-CBCA through the adaptive channel contrast blueprint separable convolution BSConv, and achieve the intra-block aggregation of the AC-CBCA feature block through the adaptive interaction module AIM and the channel gating BSConv feedforward network CGBFN; Step 3-2: Calculate the spatial attention feature AS-BSA through the adaptive spatial BSConv, and achieve the intra-block aggregation of the AS-BSA feature block through the adaptive interaction module AIM and the spatial gating BSConv feedforward network SGBFN; Step 3-3: Aggregate the inter-block features between the channel and spatial dimensions by alternately organizing AC-CBCA and AS-BSA; Step 3-4: Extract intermediate features in each dimension through multiple directly connected RDMABs, input the intermediate feature information into a convolutional layer, and smooth to obtain deep features.
[0009] As a further improvement of the present invention, the specific steps of calculating the channel attention feature through the adaptive channel contrast blueprint separable convolution and achieving the intra-block aggregation of the AC-CBCA feature block through the adaptive interaction module and the channel gating BSConv feedforward network in Step 3-1 are as follows: Along the channel dimension, the feature map is divided into two parts: convolution and multiplication bypass. For the input of a given CGBFN module , the calculation of CGBFN is as follows: , where and represent linear projections, represents the GELU activation function, are the learnable parameters of the BSConv convolution, and are both in space. Among them, represents the hidden dimension in CGBFN, represents the CGBFN operation function, H and W are the height and width of the input image respectively, and C is the number of channels of the input image; After that, the following formula is used to extract the AC-CBCA features of the shallow features of the LR image to be processed: where are the shallow features of the LR image to be processed, represents the LayerNorm module, indicating element-wise multiplication, are the input features of AC-CBCA, are the AC-CBCA features output from the AC-CBCA sub-module, represents the CBCA module, represents the AIM module, represents the CGBFN module, represents the BSConv convolution, represents element-wise addition.
[0010] As a further improvement of the present invention, the specific method for realizing the intra-block aggregation of the AS-BSA feature by calculating the spatial attention feature AS-BSA through the adaptive spatial BSConv and the adaptive interaction module AIM and the spatial gating BSConv feed-forward network SGBFN in step 32 is as follows: The BSConv spatial attention BSA in the AS-BSA module is used to generate a spatial attention feature map by utilizing the relationship inside the feature map space, and the BSConv and The Conv is used to aggregate feature maps to generate spatial descriptors, and the AIM is used to fully fuse the feature maps obtained from the two branches according to the type of self-attention mechanism. The calculation method of BSConv is as follows: For the input X of a given SGBFN module, the calculation of SGBFN is as follows: Where, represents the GELU function, 、 and are the learnable parameters of 3 BSConv convolutions respectively; After that, the following formula is used to extract the AS-BSA features of the shallow features of the LR image to be processed: Where, is the AC-CBCA feature output from the AC-CBCA sub-module and serves as the input of the AS-BSA sub-module, represents the BSA module, represents the AIM module, represents the CGBFN module, represents the BSConv convolution, represents element-wise addition, is the AS-BSA feature output from the AS-BSA sub-module.
[0011] As a further improvement of the present invention, in steps three and four, the intermediate features in each dimension are extracted through multiple directly connected RDMABs, and the intermediate feature information is input into the convolutional layer. The specific method for obtaining the deep features through smoothing is as follows: Each RDMAB includes k dual-channel spatial Transformer blocks DCSTB, followed by a 3*3 convolutional layer for connection. The specific calculation method for the DCSTB to extract features is as follows: Where, represents the SW-MSA module, is the shallow feature of the LR image to be processed, is the input feature of AC-CBCA and SW-MSA, is the AS-BSA feature output from the AS-BSA sub-module, represents the LayerNorm module, representing element-wise multiplication, represents the multi-layer perceptron MLP module, represents element-wise addition. When the DCSTB initially receives the feature When they are input into the first-layer normalization LN operation, subsequently, DCSAM parallel to the SW-MSA module is introduced; After that, for the specific calculation of the self-attention module, given the input feature size of , specifically as follows: First, the input features are reconstructed into in size, which is divided into non-overlapping local windows of size M and the number of , then calculate the self-attention in each window. For the local window feature , calculate the vectors of query (Q), key (K), and value (V) through linear mapping: where, , , is a linear projection without bias. The window-based self-attention can be expressed as: where d is the dimension of the query and the key, and E is the relative position encoding. In addition, the window shift division method is used to establish the connection between adjacent non-overlapping windows, and the shift size is set to half of the window size. After that, given the input feature of the i-th RDMAB, extract the intermediate feature using k DCSTBs; the specific formula for this extraction is: where i is the number of residual blocks in the RDMAB for intermediate feature extraction, j is the number of layers of DCSTB in the RDMAB, represents the operation of the j-th DCSTB in the i-th RDMAB. Next, the intermediate feature extraction module consists of multiple directly connected RDMABs, and this process can be expressed as: where n is the number of intermediate feature extraction blocks, represents the input feature and output feature of the i-th RDMAB. We use the BSConv layer to smooth the intermediate feature to obtain the deep feature .
[0012] As a further improvement of the present invention, the specific steps of fusing the shallow and deep feature information through convolution operation in step four, reconstructing the low-resolution LR image to obtain the high-resolution HR image, and outputting the high-resolution image HR are as follows: Reconstruct the image through the image reconstruction module. The image reconstruction module It consists of a 3*3 standard convolution with pixel shuffling, and global residual connections are added during the reconstruction process to fuse shallow and deep features, as follows: Among them, represents the high-resolution image finally reconstructed through the above steps.
[0013] The beneficial effects of the present invention are as follows. Through the settings of Step 1 to Step 4, the low-resolution image can be passed through the blueprint separable convolution layer to extract the shallow features of the image; the shallow features are simultaneously input into the residual double multi-attention module containing channel attention, spatial attention, and self-attention mechanisms to extract intermediate features in each dimension, and the intermediate feature information is input into the convolution layer to extract deep features through aggregation; the shallow and deep feature information is fused through convolution operations to reconstruct the low-resolution image into a high-resolution image and output the high-resolution image. The present invention can effectively perform feature fusion and selection between different scales and different types of information, thereby improving the quality and generalization ability of image restoration. Description of the Drawings
[0014] Figure 1 is a flowchart of the image super-resolution method of the present invention; Figure 2 In (a) is the structure diagram of AC-CBCA, (b) is the structure diagram of AS-BSA, (c) is the structure diagram of S-I, (d) is the structure diagram of C-I and the structure diagram of AIM, and (e) is the interpretation of the corresponding symbols in figures (a) to (d); Figure 3 In (a) is the structure diagram of CGBFN, and (b) is the structure diagram of SGBFN; Figure 4 is the structure diagram of DCSAM; Figure 5 is the overall structure diagram of the system carried by the image super-resolution method of the present invention. Detailed Embodiment
[0015] The following will further elaborate on the present invention in combination with the embodiments given in the drawings.
[0016] Referring to Figure 1 as shown, an image super-resolution method of this embodiment includes: S1: Obtain a low-resolution (LR) image; S2: Extract the shallow features of the LR image through the blueprint separable convolution layer; S3: The shallow features are simultaneously input into a residual dual multi-attention module (RDMAB) that includes channel attention, spatial attention, and self-attention mechanisms to extract intermediate features in each dimension. The intermediate feature information is input into a convolutional layer and smoothed to obtain deep features; S4: The shallow and deep feature information is fused through convolutional operations to reconstruct the LR image into a high-resolution (HR) image, and the HR image is output.
[0017] Further, the specific steps of step S2 are as follows: Through a convolutional layer, according to the following formula, the shallow features of the LR image to be processed are extracted: Where, represents the shallow features of the low-resolution image to be processed, represents the low-resolution image to be processed, is the operation of extracting the shallow features in using blueprint separable convolution.
[0018] Further, the specific steps in step S3 of extracting intermediate features in each dimension according to the shallow features being simultaneously input into the RDMAB that includes channel attention, spatial attention, and self-attention mechanisms, and inputting the intermediate feature information into a convolutional layer and smoothing to obtain deep features are as follows: S301: Calculate the channel attention feature (AC-CBCA) through an adaptive channel contrast blueprint separable convolution (the blueprint separable convolution is denoted as: BSConv), and achieve intra-block aggregation within the AC-CBCA feature block through an adaptive interaction module (AIM) and a channel gated BSConv feed-forward network (CGBFN); S302: Calculate the spatial attention feature (AS-BSA) through an adaptive spatial BSConv, and achieve intra-block aggregation within the AS-BSA feature block through an adaptive interaction module (AIM) and a spatial gated BSConv feed-forward network (SGBFN); S303: Aggregate the inter-block features between the channel and spatial dimensions by alternately organizing AC-CBCA and AS-BSA; S304: Extract intermediate features in each dimension through multiple directly connected RDMABs, input the intermediate feature information into a convolutional layer, and smooth to obtain deep features.
[0019] Further, the specific steps of S301 include: The AC-CBCA module utilizes the interdependence between feature channels. Among them, the contrastive BSConv channel attention (CBCA) uses Contrast to evaluate the contrast of the feature map, adjusts the weight of each channel according to the sum of the standard deviation and mean of the feature map, applies BSConv and 1*1Conv to aggregate the feature map, generates a channel descriptor, and fully fuses the feature maps obtained from the two branches using AIM according to the type of self-attention mechanism, as Figure 2 (a) shows.
[0020] Redundant information in the channels hinders the feature expression ability. To overcome this limitation, this invention patent proposes a channel gating feed-forward network (CGFN), introducing channel gating BSConv (CGB) into the feed-forward network (FFN), as Figure 2 (a) shows, which is called the channel gating BSConv feed-forward network (CGBFN). The CGB in this CGBFN module is a gating mechanism composed of BSConv convolution and element-wise multiplication. Along the channel dimension, the feature map is divided into two parts: convolution and multiplicative bypass. For the input of a given CGBFN module , the calculation of CGBFN is as follows: , Among them, and represent linear projections, represents the GELU activation function, is the learnable parameter of the BSConv convolution, and are both in the space, where represents the hidden dimension in CGBFN, represents the CGBFN operation function, H and W are the height and width of the input image respectively, and C is the number of channels of the input image.
[0021] AIM can supplement the spatial window self-attention with channel information and enhance the channel self-attention from the spatial dimension. CGBFN can introduce additional non-linear spatial information into the FFN that only models channel relationships. Therefore, the AC-CBCA module can aggregate the features in each block, and the structure of the AIM module is as Figure 2 (c), (d) show.
[0022] According to the above description, using the following formula, extract the AC-CBCA features of the shallow features of the to-be-processed LR image: Among them, is the shallow feature of the LR image to be processed, represents the LayerNorm module, and represents element-wise multiplication, is the input feature of AC-CBCA, is the AC-CBCA feature output from the AC-CBCA sub-module, represents the CBCA module, represents the AIM module, represents the CGBFN module, represents the BSConv convolution, represents element-wise addition.
[0023] Furthermore, the specific steps of step S302 include: In the AS-BSA module, the BSConv spatial attention (BSA) in the AS-BSA module uses the relationship within the feature map space to generate a spatial attention feature map, applies BSConv and 1*1Conv to aggregate the feature map, generates a spatial descriptor, and fully fuses the feature maps obtained from the two branches using AIM according to the type of self-attention mechanism, as Figure 2 shown in (b).
[0024] To further restore the structural information, this invention patent adopts a spatial gated BSConv feed-forward network (SGBFN). This feed-forward network has two important components: a spatial gating mechanism and BSConv. Using BSConv to learn the local information between spatially adjacent pixels is very effective for learning the local similarity information of the image for restoration and reconstruction. The specific structure is as Figure 2 shown in (b). For the input X of a given SGBFN module, the calculation of SGBFN is as follows: Among them, represents the GELU function, , and are respectively Figure 2 the learnable parameters of the 3 BSConv convolutions in (b). SGBFN controls the information flow at each level in our pipeline, so that each level can focus on details complementary to other levels and can provide richer context information.
[0025] AIM can supplement spatial window self-attention with channel information to enhance channel self-attention in the spatial dimension. SGBFN can introduce additional non-linear spatial information into the FFN that only models channel relationships. Therefore, the AS-BSA module can aggregate features in each block, and the structure of the AIM module is as shown in Figure 2 (c) and (d).
[0026] According to the above description, using the following formula, extract the AS-BSA features of the shallow features of the LR image to be processed: where is the AC-CBCA feature output from the AC-CBCA sub-module and serves as the input to the AS-BSA sub-module, represents the BSA module, represents the AIM module, represents the CGBFN module, represents the BSConv convolution, represents element-wise addition, is the AS-BSA feature output from the AS-BSA sub-module.
[0027] Furthermore, the step S303 specifically includes: This invention patent alternately uses AC-CBCA and AS-BSA to capture features in two dimensions and utilizes their complementary advantages. AC-CBCA can better establish channel dependencies. At the same time, AS-BSA models the long-range spatial background, enhancing the spatial expressiveness of each feature map. AC-CBCA models the global channel context, which in turn helps AS-BSA capture spatial features and expand the receptive field. Therefore, channel and spatial information flow between consecutive dual-channel spatial attention module (DCSAM) blocks, enabling aggregation, as shown in Figure 3 .
[0028] Furthermore, the step S304 specifically includes: The intermediate feature extraction module consists of multiple directly connected RDMABs. Each RDMAB includes k dual-channel spatial Transformer blocks (DCSTBs), followed by a 3*3 convolutional layer for connection, as shown in Figure 4 . DCSTB solves the problem of insufficient feature information in the channel and spatial dimensions by combining parallel cascaded DCSAM with window shift-based multi-head attention (SW-MSA) in the standard Transformer block. The specific calculation process for DCSTB to extract features can be defined as: Among them, represents the SW-MSA module, is the shallow feature of the LR image to be processed, is the input feature of AC-CBCA and SW-MSA, is the AS-BSA feature output from the AS-BSA sub-module, represents the LayerNorm (LN) (layer normalization) module, and represents element-wise multiplication, represents the multi-layer perceptron (MLP) module, represents element-wise addition. When DCSTB initially receives the features , they are input into the first layer normalization (LN) operation. Subsequently, DCSAM parallel to the SW-MSA module is introduced.
[0029] For the specific calculation of the self-attention module, given the input feature size of , the specific process is as follows: First, the input feature is reconstructed into of size, which is divided into non-overlapping local windows of size M and the number of . Then, the self-attention in each window is calculated. For the local window feature , the vectors of query (Q), key (K), and value (V) are calculated through linear mapping: Among them, , , is a linear projection without bias. The window-based self-attention can be expressed as: Among them, d is the dimension of the query and the key, and E is the relative position encoding. In addition, the window shift division method is used to establish the connection between adjacent non-overlapping windows, and the shift size is set to half of the window size.
[0030] It can be known from Figure 4 that given the input feature of the i-th RDMAB, k DCSTBs are used to extract the intermediate feature ; the specific formula for this extraction is: Among them, i is the number of residual blocks in RDMAB for intermediate feature extraction, j is the number of layers of DCSTB in RDMAB, Denote the operation of the j-th DCSTB in the i-th RDMAB. Next, the intermediate feature extraction module consists of multiple directly connected RDMABs, and this process can be expressed as: where n is the number of intermediate feature extraction blocks, denote the input feature and output feature of the i-th RDMAB. We use the BSConv layer to smooth the intermediate features to obtain the deep features .
[0031] Furthermore, the specific step S4 is as follows: The image reconstruction module consists of pixel-shuffled standard convolutions, aiming to upsample the fused features and restore them to the HR size. Finally, the shallow and deep features are fused by adding global residual connections. This process can be expressed as: where, denote the high-resolution image finally reconstructed through the above steps.
[0032] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of the present invention.
Claims
1. An image super-resolution method, characterized in that: It includes the following steps: Step 1, obtain a low-resolution LR image; Step 2, extract the shallow features of the low-resolution LR image through a blueprint-separable convolutional layer; Step 3, the shallow features extracted in Step 2 are simultaneously input into a residual dual multi-attention module RDMAB that includes channel attention, spatial attention, and self-attention mechanisms to extract intermediate features in each dimension. The intermediate feature information is input into a convolutional layer and smoothed to obtain deep features; Step 4, fuse the shallow and deep feature information through convolutional operations to reconstruct the low-resolution LR image into a high-resolution HR image and output the high-resolution image HR.
2. The image super-resolution method according to claim 1, wherein: The specific steps of extracting the shallow features of the low-resolution LR image through the blueprint-separable convolutional layer in Step 2 are as follows: Through a convolutional layer, according to the following formula, extract the shallow features of the to-be-processed low-resolution LR image: Among them, represents the shallow features of the low-resolution image to be processed, represents the low-resolution image to be processed, is an operation that extracts the shallow features in using blueprint separable convolution.
3. The image super-resolution method according to claim 1 or 2, characterized in that: The specific steps of obtaining deep features in Step 3 are as follows: Step 3-1, calculate the channel attention feature AC-CBCA through an adaptive channel contrast blueprint-separable convolution BSConv, and achieve intra-block aggregation of the AC-CBCA feature block through an adaptive interaction module AIM and a channel gating BSConv feed-forward network CGBFN; Step 3-2, calculate the spatial attention feature AS-BSA through an adaptive spatial BSConv, and achieve intra-block aggregation of the AS-BSA feature block through an adaptive interaction module AIM and a spatial gating BSConv feed-forward network SGBFN; Step 3-3, alternately organize AC-CBCA and AS-BSA to aggregate the inter-block features between the channel and spatial dimensions; Step 3-4, through multiple directly connected RDMABs, extract intermediate features in each dimension. The intermediate feature information is input into a convolutional layer and smoothed to obtain deep features.
4. The image super-resolution method according to claim 3, wherein: The specific steps of calculating the channel attention feature through an adaptive channel contrast blueprint-separable convolution and achieving intra-block aggregation of the AC-CBCA feature block through an adaptive interaction module and a channel gating BSConv feed-forward network in Step 3-1 are as follows: Along the channel dimension, the feature map is divided into two parts: convolution and multiplication bypass. For the input of a given CGBFN module , the calculation of CGBFN is as follows: , Among them, and represent linear projections, represents the GELU activation function, are the learnable parameters of the BSConv convolution, and are both in space, where represents the hidden dimension in CGBFN, represents the CGBFN operation function, H and W are the height and width of the input image respectively, and C is the number of channels of the input image; After that, use the following formula to extract the AC-CBCA feature of the shallow features of the to-be-processed LR image: Among them, is the shallow feature of the LR image to be processed, represents the LayerNorm module and denotes element-wise multiplication, is the input feature of AC-CBCA, is the AC-CBCA feature output from the AC-CBCA sub-module, represents the CBCA module, represents the AIM module, represents the CGBFN module, ) represents the BSConv convolution, denotes element-wise addition.
5. The image super-resolution method according to claim 4, wherein: The specific method of calculating the spatial attention feature AS-BSA through an adaptive spatial BSConv and achieving intra-block aggregation of the AS-BSA feature block through an adaptive interaction module AIM and a spatial gating BSConv feed-forward network SGBFN in Step 3-2 is as follows: The BSConv in the AS-BSA module uses the spatial attention BSA to utilize the relationships within the feature map space to generate a spatial attention feature map. Apply BSConv and Conv to aggregate the feature maps, generate a spatial descriptor, and fully fuse the feature maps obtained from the two branches using AIM according to the type of self-attention mechanism. Among them, the calculation method of BSConv is as follows: For the input X of the given SGBFN module, the calculation of SGBFN is as follows: Among them, represents the GELU function, , and are the learnable parameters of three BSConv convolutions respectively; After that, use the following formula to extract the AS-BSA feature of the shallow features of the to-be-processed LR image: Among them, is the AC-CBCA feature output from the AC-CBCA sub-module and serves as the input to the AS-BSA sub-module, represents the BSA module, represents the AIM module, represents the CGBFN module, represents the BSConv convolution, represents element-wise addition, is the AS-BSA feature output from the AS-BSA sub-module.
6. The image super-resolution method according to claim 4, characterized in that: The specific method of extracting intermediate features in each dimension through multiple directly connected RDMABs and inputting the intermediate feature information into a convolutional layer and smoothing to obtain deep features in Step 3-4 is: each RDMAB includes k dual-channel spatial Transformer blocks DCSTB, followed by a 3*3 convolutional layer for connection. The specific calculation method for the DCSTB to extract features is as follows: Among them, represents the SW-MSA module, is the shallow feature of the LR image to be processed, is the input feature of AC-CBCA and SW-MSA, is the AS-BSA feature output from the AS-BSA sub-module, represents the LayerNorm module, representing element-wise multiplication, represents the multi-layer perceptron MLP module, represents element-wise addition. When DCSTB initially receives the features they are input into the first layer normalization LN operation. Subsequently, DCSAM parallel to the SW-MSA module is introduced; For the specific calculation of the self-attention module later, given the input feature size of , the details are as follows: First, reconstruct the input feature into in size, divide it into non-overlapping local windows of size M and the number of , then calculate the self-attention in each window. For the local window feature , calculate the vectors of query (Q), key (K), and value (V) through linear mapping: Among them, , , is a linear projection without deviation. Window-based self-attention can be expressed as: Among them, d is the dimension of the query and the keyword, and E is the relative position encoding. In addition, a window shift partitioning method is used to establish the connection between adjacent non-overlapping windows, and the shift size is set to half of the window size. After that, given the input feature of the i-th RDMAB , k DCSTBs are used to extract intermediate features ; The specific formula for this extraction is: where i is the number of residual blocks in the RDMAB for intermediate feature extraction, and j is the number of layers of the DCSTB in the RDMAB. represents the operation of the j-th DCSTB in the i-th RDMAB. Next, the intermediate feature extraction module consists of multiple directly connected RDMABs, and this process can be expressed as: where n is the number of intermediate feature extraction blocks, represent the input and output features of the i-th RDMAB. We use the BSConv layer to smooth the intermediate features to obtain the deep features .
7. The image super-resolution method according to claim 4, wherein: In the fourth step, the shallow and deep feature information is fused through convolution operation to reconstruct the low-resolution LR image into a high-resolution HR image, and the specific steps for outputting the high-resolution image HR are as follows: The image is reconstructed by the image reconstruction module, and the image reconstruction module consists of 3*3 standard convolutions with pixel shuffling. During the reconstruction process, global residual connections are added to fuse shallow and deep features, as follows: Among them, represents the high-resolution image finally reconstructed through the above steps.