Transform image super-division method based on cross-dimension

By using a transdimensional Transformer network in the image super-segment method, connecting spatial self-attention and channel self-attention units in series, the problem of poor super-segment effect in the prior art is solved, and a higher quality image super-segment effect is achieved.

CN120163712AActive Publication Date: 2025-06-17ANHUI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510361048.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-17
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

The super-segment effect of the image super-segment method in the prior art is poor, especially due to the high correlation between space and channels, the performance of a single-dimensional Transformer in super-resolution is limited.

Method used

A transdimensional image super-segment method is adopted to construct a transdimensional Transformer super-segment network through a series of spatial self-attention units and channel self-attention units to realize the aggregation of spatial information and channel information.

Benefits of technology

By connecting spatial self-attention and channel self-attention units in series, the model can more effectively capture the features of the spatial and channel dimensions, improve feature representation capabilities, and achieve excellent performance in image super-segment tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163712A_ABST
    Figure CN120163712A_ABST
Patent Text Reader

Abstract

The invention provides a trans-dimension-based Transform image super-division method. The method comprises the following steps: acquiring a low-resolution image to be processed; a low-resolution image to be processed is input into the trained cross-dimension Transform super-division network, and a high-resolution image after super-division is obtained; wherein the cross-dimension Transform super-division network realizes aggregation of space information and channel information through a space self-attention unit and a channel self-attention unit which are connected in series. According to the method, a space self-attention unit and a channel self-attention unit are connected in series, correlation between a channel and a space is modeled, and self-attention calculation of the space unit and the channel unit is guided, so that aggregation of space information and channel information is realized to further improve feature representation capability; and the excellent performance of the image under the super-resolution task is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a Transformer image super-resolution method based on cross-dimension. Background Art

[0002] In the big data era, the wide application of high-quality images in various fields has become increasingly prominent. In the field of social media, high-quality images help to improve the visual experience of users; in the medical field, high-quality images can assist professionals in making more accurate diagnoses and judgments; in the field of security, high-quality images help to maintain social security and surveillance. However, due to various factors such as equipment costs and natural environments, the images collected usually have the characteristics of low resolution, poor quality, and blurriness. Therefore, image enhancement technology has important research value and practical application significance in various fields. Among them, super-resolution technology (SR) is of great significance in image enhancement.

[0003] SR is a technology for generating high-resolution (HR) images from low-resolution (LR) images. In recent years, Convolutional Neural Network (CNN) has become a research hotspot due to its powerful ability to fit complex mappings. The super-resolution SR method based on the convolutional neural network CNN is widely popular due to its powerful ability to extract high-frequency details from images. However, it is difficult to establish global dependencies using CNN-based methods. As an alternative, Transformer-based methods utilize the powerful self-attention mechanism SA to model the global dependencies of the input data and show impressive performance.

[0004] Recently, Transformer has gained considerable popularity in low-level vision tasks including image super-resolution (SR). These networks utilize self-attention along the spatial or channel dimensions and have made remarkable progress. However, the high correlation between space and channels has severely limited the performance of single-dimension Transformer in super-resolution, resulting in poor super-resolution effects. Summary of the Invention

[0005] In view of the above defects of the prior art, the present invention provides a Transformer image super-resolution method based on cross-dimension to solve the technical problem of poor super-resolution effects of various image super-resolution methods in the prior art.

[0006] To achieve the above object and other related objects, the present invention provides a cross-dimensional Transformer image super-resolution method, including: obtaining a low-resolution image to be processed; inputting the to-be-processed low-resolution image into a trained cross-dimensional Transformer super-resolution network to obtain a super-resolved high-resolution image; wherein, the cross-dimensional Transformer super-resolution network aggregates spatial information and channel information through a series-connected spatial self-attention unit and channel self-attention unit.

[0007] In an embodiment of the present invention, the cross-dimensional Transformer super-resolution network is expressed by the formula: X HR = H RC (H SF (X LR ) + H DF (H SF (X LR ))) ; wherein, X LR is the low-resolution image, X HR is the high-resolution image, H SF (), H DF (), H RC () are a shallow feature extraction unit, a deep feature extraction unit, and a high-resolution reconstruction unit respectively.

[0008] In an embodiment of the present invention, the deep feature extraction unit is expressed by the formula: X d = H DF (X s ) = M D (M D-1 (… M2(M1(X s )))) ; wherein, X s and X d are respectively the input and output of the deep feature extraction unit, and M i is the i-th cross-dimensional Transformer module, i = {1, 2, …, D}.

[0009] In an embodiment of the present invention, the cross-dimensional Transformer module is expressed by the formula: X1 = CDSSA(LN(X0)) + X0; X2 = MLP(LN(X1)) + X1; X3 = CDCSA(LN(X2)) + X2; X4 = MLP(LN(X3)) + X3; wherein, X0 and X4 are respectively the input and output of the cross-dimensional Transformer module, X1~X3 are intermediate calculation features, LN() is layer normalization, CDSSA() is a spatial self-attention unit, MLP() is a multi-layer perceptron, and CDCSA() is a channel self-attention unit.

[0010] In an embodiment of the present invention, the spatial self-attention unit processes the first input feature as follows: According to the first input feature, matrices of queries, keys, and values are generated through linear projection and reconstructed to obtain matrix Q S , K S and V S ; According to the matrix Q S , K S , V S and the following formula, the intermediate feature Y1 is calculated: , where softmax() is normalization, is the transpose of matrix K S , d k is the feature dimension of matrix K S ; According to the intermediate feature Y1 and the following formula, the output feature Y of the spatial self-attention unit is calculated CDSSA : Y CDSSA = Y1·SEU(X in1 ), where X in1 is the first input feature, and SEU() is the spatial embedding unit.

[0011] In an embodiment of the present invention, the spatial embedding unit is expressed by the formula: SEU(X in1 ) = Sig(Conv(GELU(Conv(X in1 )))), where Conv() is the convolution operation, GELU() is the activation function, and Sig() is the Sigmoid() function.

[0012] In an embodiment of the present invention, the channel self-attention unit processes the second input feature as follows: According to the second input feature, matrices of queries, keys, and values are generated through linear projection along the channel dimension and reconstructed to obtain matrix Q C , K C and V C ; According to the matrix Q C , K C , V C and the following formula, the intermediate feature Y2 is calculated: , where softmax() is normalization, is the transpose of matrix Q C , and α is a learnable parameter; According to the intermediate feature Y2 and the following formula, the output feature Y of the channel self-attention unit is calculated CDCSA : Y CDCSA = Y2·CEU(X in2) + Y2, where X in2 is the second input feature, and CEU() is the channel embedding unit.

[0013] In an embodiment of the present invention, the channel embedding unit is represented by the formula: CEU(X in2 ) = Sig(Conv(GELU(Conv(Gap(X in2 ))))) , where Gap() is the global average pooling operation, Conv() is the convolution operation, GELU() is the activation function, and Sig() is the Sigmoid() function.

[0014] To achieve the above object and other related objects, the present invention also provides an electronic device, including a processor, a memory, and a communication bus; the communication bus is used to connect the processor and the memory; the processor is used to execute the computer program stored in the memory to implement the method provided in any one of the above embodiments.

[0015] To achieve the above object and other related objects, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to cause a computer to execute the method provided in any one of the above embodiments.

[0016] Advantages of the present invention: A Transformer image super-resolution method based on cross-dimension proposed by the present invention constructs a cross-dimension Transformer super-resolution network through a series of spatial self-attention units and channel self-attention units, and after training the super-resolution network, it can process the low-resolution image to be processed and obtain the super-resolved high-resolution image; the spatial self-attention unit can capture the features in the spatial dimension and input valuable information in the channel dimension. Correspondingly, the channel self-attention unit captures the features in the channel dimension and inputs valuable information in the spatial dimension. By connecting these two units in series, the correlation between the channel and the space is modeled, guiding the self-attention calculation of the spatial unit and the channel unit, aggregating the spatial information and the channel information to further improve the feature representation ability, and achieving excellent performance in the super-resolution task of the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1Flowchart of the image super-resolution method provided by an embodiment of the present invention; Figure 2 Architecture diagram of the cross-dimensional Transformer module provided by an embodiment of the present invention; Figure 3 Architecture diagram of the spatial self-attention unit provided by an embodiment of the present invention; Figure 4 Architecture diagram of the spatial embedding unit provided by an embodiment of the present invention; Figure 5 Architecture diagram of the channel self-attention unit provided by an embodiment of the present invention; Figure 6 Architecture diagram of the channel embedding unit provided by an embodiment of the present invention; Figure 7 Schematic structural diagram of an electronic device provided by an embodiment of the present invention.

[0019] Explanation of reference numerals: 101, processor; 102, memory. Detailed implementation manners

[0020] The following uses specific specific embodiments to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. In addition to the specific methods, devices, and materials used in the embodiments, according to the knowledge of those skilled in the art in the technical field and the description of the present invention, any methods, devices, and materials similar to or equivalent to those described in the embodiments of the present invention can also be used to implement the present invention.

[0021] It should be understood that the terms used in the embodiments of the present invention are for the purpose of describing specific specific implementation manners, rather than for limiting the protection scope of the present invention. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the technical field of the present invention.

[0022] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In some of these embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0023] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of methods and computer program products that can be implemented according to various embodiments disclosed in the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0024] Please refer to Figure 1 , Figure 1 A cross-dimensional Transformer image super-resolution method provided by an embodiment of the present invention includes: obtaining a low-resolution image to be processed; inputting the low-resolution image to be processed into a trained cross-dimensional Transformer super-resolution network to obtain a super-resolved high-resolution image; wherein, the cross-dimensional Transformer super-resolution network aggregates spatial information and channel information through a series-connected spatial self-attention unit and channel self-attention unit.

[0025] In this embodiment, the training process of the cross-dimensional Transformer super-resolution network is the same as that of a common network model. Specifically, it is as follows: (1) First, obtain an image pair as a training sample. The image pair consists of a low-resolution image I LR and a high-resolution image I HR . Among them, the low-resolution image I LR can be obtained by downsampling the I HR image; (2) Build a cross-dimensional Transformer super-resolution network; (3) Use the training sample to train the cross-dimensional Transformer super-resolution network to obtain a trained cross-dimensional Transformer super-resolution network.

[0026] In a specific embodiment of the present invention, the cross-dimensional Transformer super-resolution network is represented by the formula: X HR =H RC (H SF (X LR ) + H DF (H SF (X LR ))); wherein, X LRis a low-resolution image, X HR is a high-resolution image, H SF (), H DF (), H RC () are a shallow feature extraction unit, a deep feature extraction unit, and a high-resolution reconstruction unit respectively.

[0027] In this embodiment, the shallow feature extraction unit is used to extract low-level features of the input image, such as basic information like edges and textures, providing a preliminary feature representation for subsequent deep feature extraction. The shallow feature extraction unit can be, for example, a convolutional layer.

[0028] The deep feature extraction unit is used to capture high-level features of the image. By modeling the global dependencies in the spatial and channel dimensions, it enhances the expressive power of the features. It strengthens the acquisition of local information in the spatial-channel information interaction and generates feature pairs with short-range and long-range contexts.

[0029] The high-resolution reconstruction unit is used to upsample and fuse the information at these two levels to generate the final high-resolution image, completing the super-resolution task from low resolution to high resolution.

[0030] Among the above three units, the deep feature extraction unit realizes the aggregation of spatial information and channel information through a series-connected spatial self-attention unit and channel self-attention unit.

[0031] In a specific embodiment of the present invention, the deep feature extraction unit is represented by the formula: X d =H DF (X s )=M D (M D-1 (…M2(M1(X s )))); Wherein, X s and X d are respectively the input and output of the deep feature extraction unit, and M i is the i-th cross-dimension Transformer module, i = {1, 2, …, D}. In this embodiment, the deep feature extraction unit is composed of D cross-dimension Transformer modules (Cross-Dimension Transformer Block, CDTB). After D CDTBs, the output of the deep feature extraction unit is obtained.

[0032] Please refer to Figure 2 , in a specific embodiment of the present invention, the cross-dimension Transformer module is represented by the formula: X1 = CDSSA(LN(X0)) + X0; X2 = MLP(LN(X1)) + X1; X3 = CDCSA(LN(X2)) + X2; X4 = MLP(LN(X3)) + X3; Wherein, X0 and X4 are respectively the input and output of the cross - dimensional Transformer module, X1~X3 are intermediate calculation features, LN() is layer normalization, CDSSA() is a spatial self - attention unit, MLP() is a multi - layer perceptron, and CDCSA() is a channel self - attention unit.

[0033] In this embodiment, CDTB is composed of a spatial self - attention unit (Cross - Dimension Spatial SelfAttention, CDSSA) and a channel self - attention unit (Cross - Dimension Channel Self Attention, CDCSA) in series. When CDSSA and CDCSA generate outputs, a spatial embedding unit (Spatial Embedding Unit, SEU) and a channel embedding unit (Channel Embedding Unit, CEU) will respectively input valuable information in the spatial dimension and the channel dimension. CDSSA and CDCSA are respectively based on cross - dimensional self - attention in the spatial direction and cross - dimensional self - attention in the channel direction. By connecting CDSSA and CDCSA in series, CDTB can achieve feature aggregation between the spatial dimension and the channel dimension.

[0034] In this embodiment, LN() is Layer Normalization, which is a normalization operation used to standardize the output of a certain layer in a neural network so that its mean is 0 and its variance is 1. MLP() is Multi - LayerPerceptron, which is a fully - connected neural network used for non - linear transformation and feature extraction of input features, usually composed of multiple fully - connected layers (Linear Layer) and activation functions (such as ReLU).

[0035] Please refer to Figure 3 , in a specific embodiment of the present invention, the spatial self - attention unit processes the first input feature according to the following steps.

[0036] (1) According to the first input feature, generate matrices of query, key, and value through linear projection, and reconstruct them to obtain matrices Q S 、K S and V S . Denote the first input feature as X in1 ∈R C×H×W, the matrices of queries, keys, and values generated by linear projection can be represented as Q, K, V, where all matrices are in R C×H×W space, and the process is represented as: Q = X in1 W Q , K = X in1 W K , V = X in1 W V ; in this formula, W Q , W K , W V ∈R C×C , is a linear projection that omits the bias. Reconstruct Q, K, V into a two-dimensional matrix form in R C×HW (combining the height and width dimensions) to obtain matrices Q S , K S and V S .

[0037] (2) Calculate the intermediate feature Y1 according to the matrices Q S , K S , V S and the following formula: , where softmax() is normalization, is the transpose of matrix K S , and d k is the feature dimension of matrix K S . In the self-attention mechanism, the dot product of matrices Q S and K S is used to calculate the similarity (i.e., attention weights) between input features, and the transpose operation is to make the dimensions of matrix multiplication match. d k is used to scale the dot product result to prevent the value of the dot product from being too large, resulting in unstable gradients.

[0038] (3) Calculate the output feature Y CDSSA of the spatial self-attention unit according to the intermediate feature Y1 and the following formula: Y CDSSA = Y1·SEU(X in1 ) + Y1, where X in1 is the first input feature and SEU() is the spatial embedding unit. The design of the spatial embedding unit utilizes valuable channel information, enabling guidance in the channel dimension during self-attention calculation in the spatial dimension.

[0039] Please refer to Figure 4 , in a specific embodiment of the present invention, the spatial embedding unit is represented by the formula: SEU(X in1 ) = Sig(Conv(GELU(Conv(X in1)))) where Conv() is the convolution operation, GELU() is the activation function, and Sig() is the Sigmoid() function. SEU is used to extract valuable spatial information from the spatial dimension and embed it into the self-attention mechanism to enhance the expression ability of spatial features. Specifically, SEU generates spatial information weights through convolution and activation functions, and these weights are used to guide the calculation of spatial self-attention (CDSSA).

[0040] Please refer to Figure 5 , in a specific embodiment of the present invention, the channel self-attention unit processes the second input feature as follows: (1) According to the second input feature, matrices of query, key, and value are generated by linear projection along the channel dimension and reconstructed to obtain matrix Q C , K C and V C . The size of the reconstructed matrix is R HW×C .

[0041] (2) According to matrix Q C , K C , V C and the following formula, the intermediate feature Y2 is calculated: , where softmax() is normalization, is the transpose of matrix Q C , and α is a learnable parameter. The self-attention mechanism on the channel dimension is different from that on the spatial dimension. The channel dimension is usually much smaller than the spatial dimension. Using the learnable parameter α can more flexibly adapt to the feature distribution on the channel dimension, enabling the model to dynamically adjust the scaling factor according to the data, thereby better capturing the dependencies between channels.

[0042] (3) According to the intermediate feature Y2 and the following formula, the output feature Y of the channel self-attention unit is calculated CDCSA : Y CDCSA = Y2·CEU(X in2 ) + Y2, where X in2 is the second input feature and CEU() is the channel embedding unit. The design of the channel embedding unit utilizes valuable spatial information, enabling the guidance of the spatial dimension in the self-attention calculation on the channel dimension.

[0043] Please refer to Figure 6 , in a specific embodiment of the present invention, the channel embedding unit is expressed by the formula: CEU(X in2 ) = Sig(Conv(GELU(Conv(Gap(X in2))))), where Gap() is the global average pooling operation, Conv() is the convolution operation, GELU() is the activation function, and Sig() is the Sigmoid() function. The channel embedding unit is used to extract valuable channel information from the channel dimension and embed it into the self-attention mechanism to enhance the expression ability of channel features. Specifically, the CEU generates channel information weights through global average pooling (GAP), convolution, and activation functions, and these weights are used to guide the calculation of channel self-attention (CDCSA).

[0044] Please refer to Figure 7 , Figure 7 An electronic device provided by an embodiment of the present invention includes a processor 101, a memory 102, and a communication bus; the communication bus is used to connect the processor 101 and the memory 102; the processor 101 is used to execute a computer program stored in the memory 102 to implement the above-mentioned cross-dimensional Transformer image super-resolution method.

[0045] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to make a computer execute the above-mentioned cross-dimensional Transformer image super-resolution method.

[0046] In a specific embodiment of the present invention, the high-resolution reconstruction unit is expressed by the formula:[[]]END]] X h =H RC (X s +X d )=Reshape(Conv(Upsampling(X s ,X d ))); Among them, Upsampling() is the upsampling operation, which is an operation to enlarge the spatial dimensions (height and width) of the input feature map, usually implemented by interpolation or transposed convolution. In this embodiment, the transposed convolutional layer is used to upsample these feature maps to the desired scale by a scale factor s. Then, after passing through the Conv() convolution operation and Reshape(), the output size becomes B×sH×sW, that is, the high-resolution image X HR .

[0047] In the image enhancement task, the L1 loss function is used to calculate the error between the super-resolution result I SR and the real data I HR , and the backpropagation method is used to optimize the network parameters, so as to complete the training of the layer-by-layer context information aggregation network. Specifically, the above task of training the network using training samples includes the following steps: the image I in the LR-HR image pairLR Substitute it into the cross-dimensional Transformer super-resolution network to obtain image I SR ; According to image I SR in the LR-HR image pair HR and the following loss function calculation formula to calculate the loss loss: ; The above steps can conveniently help us achieve the training of network parameters. The trained cross-dimensional Transformer super-resolution network can directly be used to perform super-resolution on the input X LR image to obtain its corresponding X HR image.

[0048] Next, by comparing with other super-resolution algorithms, the effects of the present invention are further illustrated. When comparing, the training dataset and validation dataset used are the training set and validation set of the DF2K dataset, and five test sets are used to comprehensively evaluate its performance, namely: Set5 test set, Set14 test set, BSD100 test set, Urban100 test set, and Manga109 test set. During the evaluation, two evaluation metrics, Peak Signal to Noise Ratio (PSNR) and Structural Similarity (SSIM), are used to evaluate the super-resolution results. To verify the reliability of the results, the present invention compares the performance of the present invention and other super-resolution methods in image enhancement tasks with super-resolution multiples of ×2, ×3, and ×4.

[0049] Table 1: Comparison of multiple super-resolution algorithms (the evaluation parameters corresponding to 2-4 after the first Scale are PSNR, and the evaluation parameters corresponding to 2-4 after the first Scale are SSIM)

[0050] In this table, the first column Method is the abbreviation of each algorithm, Params refers to the number of parameters to be learned in the model, FLOPs refers to the number of floating-point operations performed by the model during the inference process, and Scale represents the super-resolution multiple. According to the comparison in Table 1, the present invention is superior to other state-of-the-art methods while maintaining lightweight parameters. The present invention aggregates the spatial and channel dimensions using CDTB and is superior to many CNN-based methods. The present invention combines spatial and channel features, obtains powerful representation capabilities, and achieves excellent results on various test sets, demonstrating the potential of the present invention in image enhancement tasks. The image super-resolution method used in the present invention constructs a network using the joint spatial-channel information of the image and can achieve higher-quality image enhancement tasks.

[0051] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A Transformer image super-resolution method based on cross-dimensionality, characterized in that: include: Obtain a low-resolution image to be processed; Inputting the low-resolution image to be processed into a trained cross-dimensional Transformer super-resolution network to obtain a super-resolution high-resolution image; Among them, the cross-dimensional Transformer super-resolution network realizes the aggregation of spatial information and channel information by connecting spatial self-attention units and channel self-attention units in series.

2. The method for super-resolution of an image based on a cross-dimensional Transformer according to claim 1, characterized in that: The cross-dimensional Transformer super-resolution network is expressed as follows: X HR =H RC (H SF (X LR )+H DF (H SF (X LR ))); Among them, X LR is the low-resolution image, X HR is the high-resolution image, H SF (), H DF (), H RC () are shallow feature extraction unit, deep feature extraction unit and high-resolution reconstruction unit respectively.

3. The method for super-resolution of an image based on a cross-dimensional Transformer according to claim 2, characterized in that: The deep feature extraction unit is expressed by the formula: X d =H DF (X s )=M D (M D-1 (…M2(M1(X s )))); Among them, X s and X d are the input and output of the deep feature extraction unit, respectively, i is the i-th cross-dimensional Transformer module, i={1,2,…,D}.

4. The method for super-resolution of an image based on a cross-dimensional Transformer according to claim 3, characterized in that: The cross-dimensional Transformer module is expressed by the formula: X1=CDSSA(LN(X0))+X0; X2=MLP(LN(X1))+X1; X3=CDCSA(LN(X2))+X2; X4=MLP(LN(X3))+X3; Among them, X0 and X4 are the input and output of the cross-dimensional Transformer module respectively, X1~X3 are intermediate calculation features, LN() is layer normalization, CDSSA() is the spatial self-attention unit, MLP() is the multi-layer perceptron, and CDCSA() is the channel self-attention unit.

5. The method for super-resolution of an image based on a cross-dimensional Transformer according to claim 4, characterized in that: The spatial self-attention unit processes the first input feature as follows: Based on the first input feature, a matrix of query, keyword and value is generated by linear projection, and then reconstructed to obtain the matrix Q S , K S and V S ; According to the matrix Q S , K S 、V S And the following formula is used to calculate the intermediate feature Y1: , where softmax() is normalized, is the matrix K S The transpose of k is the matrix K S The characteristic dimension of The output feature Y of the spatial self-attention unit is calculated according to the intermediate feature Y1 and the following formula CDSSA : Y CDSSA =Y1·SEU(X in1 )+Y1, where X in1 is the first input feature, and SEU() is the spatial embedding unit.

6. The method for super-resolution of an image based on a cross-dimensional Transformer according to claim 5, characterized in that: The spatial embedding unit is expressed by the formula: YOUR(X in1 )=Sig(Conv(GELU(Conv(X in1 )))), Among them, Conv() is the convolution operation, GELU() is the activation function, and Sig() is the Sigmoid() function.

7. The method for super-resolution of an image based on a cross-dimensional Transformer according to claim 4, characterized in that: The channel self-attention unit processes the second input feature as follows: According to the second input feature, a matrix of query, keyword and value is generated by linear projection along the channel dimension, and it is reconstructed to obtain the matrix Q C , K C and V C ; According to the matrix Q C , K C 、V C And the following formula is used to calculate the intermediate feature Y2: , where softmax() is normalized, The matrix Q C The transpose of , α is a learnable parameter; The output feature Y of the channel self-attention unit is calculated according to the intermediate feature Y2 and the following formula CDCSA : Y CDCSA =Y2·CEU(X in2 )+Y2, where X in2 is the second input feature, and CEU() is the channel embedding unit.

8. The method for super-resolution of an image based on a cross-dimensional Transformer according to claim 7, characterized in that: The channel embedding unit is expressed by the formula: CEU(X in2 )=Sig(Conv(GELU(Conv(Gap(X in2 ))))), Among them, Gap() is the global average pooling operation, Conv() is the convolution operation, GELU() is the activation function, and Sig() is the Sigmoid() function.

9. An electronic device, characterized in that: It comprises a processor, a memory and a communication bus; the communication bus is used to connect the processor and the memory; the processor is used to execute the computer program stored in the memory to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program is used to enable a computer to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Deep learning face recognition system and method based on self-attention mechanism

    CN110610129A

  • Ultrasonic image quantification method based on interactive fusion Transform

    CN114863111A

  • Single image super-resolution reconstruction method based on channel fusion self-attention mechanism

    CN116205789A

  • Single-frame image super-resolution method and system based on full-distance feature aggregation

    CN117218005A

  • Single-frame image super-resolution method and device based on mixed feature interaction Transformer

    CN117422614A