Single-frame image super-resolution method and apparatus based on hybrid feature interaction transformer

By introducing hybrid feature interaction Transformer into the image super-resolution method, the problem of existing methods ignoring feature correlation is solved, and more efficient image super-resolution performance is achieved.

WO2025129752A1PCT designated stage expired Publication Date: 2025-06-26HUAQIAO UNIVERSITY +1

Patent Information

Application Number
PCT/CN2023/143017
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2023-12-29
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing Transformer-based image super-resolution methods ignore the potential correlation between features of different dimensions, limiting the performance of image super-resolution methods.

Method used

A single-frame image super-resolution method based on hybrid feature interaction Transformer is proposed. By introducing mixed feature interaction self-attention units and mixed-scale feedforward neural networks, cross-dimensional feature interaction is encouraged, and the global feature expression ability and detail reconstruction ability of image super-resolution method are improved.

Benefits of technology

The performance of the image super-resolution method is significantly improved, and high-quality high-resolution image reconstruction can be achieved at lower parameter amounts and Flops values.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023143017_26062025_PF_FP_ABST
    Figure CN2023143017_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing. Disclosed are a single-frame image super-resolution (SR) method and apparatus based on a hybrid feature interaction transformer. The method comprises: acquiring a low-resolution image to be reconstructed; constructing a single-frame image SR model based on a hybrid feature interaction transformer and training same, so as to obtain a trained single-frame image SR model, wherein the single-frame image SR model comprises a shallow feature extraction unit, a deep feature extraction unit and an up-sampling reconstruction unit which are sequentially connected, and the deep feature extraction unit comprises P hybrid feature interaction transformer modules which are sequentially connected; and inputting the low-resolution image into the trained single-frame image SR model, extracting shallow features by means of the shallow feature extraction unit, inputting the shallow features into the deep feature extraction unit for extraction, so as to obtain deep features, and inputting the deep features into the up-sampling reconstruction unit for reconstruction to obtain a reconstructed high-resolution image. The problem of the reconstruction performance being affected due to an SR method for a transformer ignoring latent correlations between features from different dimensions is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Single-frame image super-resolution method and device based on hybrid feature interactive Transformer Technical Field

[0001] The present invention relates to the field of image processing, and in particular to a single-frame image super-resolution method and device based on a hybrid feature interactive Transformer. Background Art

[0002] Image super-resolution (SR) is a key task in computer vision and image processing. It aims to reconstruct high-quality high-resolution (HR) images from existing low-resolution (LR) images. Recently, SR methods based on convolutional neural networks (CNNs) have dominated the field of image SR due to their powerful feature representation, end-to-end trainable paradigm, and excellent performance. However, because the convolution operation extracts local features within a small neighborhood using a fixed sliding window, CNN-based SR methods have limited informative pixels. Currently, Transformer, as a novel alternative to CNNs, has achieved good performance on various low-level vision tasks.

[0003] For image SR, Liang et al. proposed an SR model based on Swin Transformer, namely SwinIR. SwinIR adopts a hierarchical design, limits the similarity calculation to a local window, and uses a moving window mechanism to enhance cross-window information interaction. However, SwinIR abandons global information reasoning due to the use of window-based self-attention, and the performance of the Transformer is limited. In order to activate more informative pixels that contribute to image SR, Chen et al. proposed HAT, in which channel attention is introduced to better aggregate cross-window information. Wang et al. proposed Omni-SR, which can simultaneously model pixel-level information interactions between spatial and window dimensions. However, existing Transformer-based SR methods generally capture the relationship between space and channels through serial or parallel operations, but ignore the potential correlation between features of different dimensions, thereby limiting the performance of Transformer-based SR methods.

[0004] Summary of the Invention

[0005] In response to the technical problems mentioned above, the purpose of the embodiments of this application is to propose a single-frame image super-resolution method and device based on a hybrid feature interaction Transformer, which overcomes the problem that the existing Transformer method ignores the potential correlation between features of different dimensions, and significantly improves the global feature expression ability and detail reconstruction ability of the image super-resolution method by encouraging cross-dimensional feature interaction.

[0006] In a first aspect, the present invention provides a single-frame image super-resolution method based on a hybrid feature interactive Transformer, comprising the following steps:

[0007] Acquire a low-resolution image to be reconstructed;

[0008] A single-frame image super-resolution model based on a hybrid feature interactive Transformer is constructed and trained to obtain a trained single-frame image super-resolution model, wherein the single-frame image super-resolution model includes a shallow feature extraction unit, a deep feature extraction unit, and an upsampling reconstruction unit connected in sequence, and the deep feature extraction unit includes P hybrid feature interactive Transformer modules connected in sequence;

[0009] The low-resolution image to be reconstructed is input into the trained single-frame image super-resolution model, the shallow features are extracted by the shallow feature extraction unit, the shallow features are input into the deep feature extraction unit to extract the deep features, and the deep features are input into the upsampling reconstruction unit to reconstruct a high-resolution reconstructed image.

[0010] Preferably, the hybrid feature interaction Transformer module includes an efficient local feature extraction unit, a first normalization layer, a hybrid feature interaction self-attention unit, a second normalization layer and a hybrid scale feedforward neural network. The input of the hybrid feature interaction Transformer module passes through the efficient local feature extraction unit and the first normalization layer in sequence. The output of the first normalization layer is added to the input of the hybrid feature interaction Transformer module to obtain a first hybrid feature. The first hybrid feature passes through the hybrid feature interaction self-attention unit and the second normalization layer in sequence. The output of the second normalization layer is added to the first hybrid feature to obtain a second hybrid feature. The second hybrid feature is input into the hybrid scale feedforward neural network to obtain the output of the hybrid feature interaction Transformer module.

[0011] Preferably, the efficient local feature extraction unit includes a first shifted convolution layer, a first GeLU activation function layer, a second shifted convolution layer, an SE module and a third shifted convolution layer connected in sequence. The calculation process of the efficient local feature extraction unit is as follows: H ELF (·)=F shift-conv (F SE (Fshift-conv (GeLU(F shift-conv (·))))) ;

[0012] Among them, H ELF (·) represents the function of efficient local feature extraction unit, F shift-conv (·) represents the shifted convolution operation of the first shifted convolution layer, the second shifted convolution layer, or the third shifted convolution layer, F SE (·) represents the function of the SE module, and GeLU(·) represents the GeLU activation function.

[0013] Preferably, the hybrid feature interaction self-attention unit includes a local window self-attention branch, a deep convolution branch and a bidirectional feature interaction unit, the bidirectional feature interaction unit includes a spatial interaction unit and a channel interaction unit, the channel interaction unit includes a global average pooling layer, a first convolution layer, a first batch of normalization layers, a second GeLU activation function layer, a second convolution layer and a first Sigmoid activation function layer connected in sequence, the spatial interaction unit includes a third convolution layer, a second batch of normalization layers, a third GeLU activation function layer, a fourth convolution layer and a second Sigmoid activation function connected in sequence, the local window self-attention branch includes a query linear layer, a key linear layer, a value linear layer and a local window self-attention module, the deep convolution branch includes a first deep convolution layer with a convolution kernel size of 3×3, the local features output by the first deep convolution layer are input into the channel interaction unit to obtain channel-level dynamic weights, and the channel-level dynamic weights are input into the local window self-attention branch to adaptively correct the value feature map output by the value linear layer; the global features output by the local window self-attention module are input into the spatial interaction unit to obtain spatial-level dynamic weights, and the spatial-level dynamic weights are input into the deep convolution branch to adaptively correct the local features.

[0014] As an example, the calculation process of the hybrid feature interactive self-attention unit is as follows: the first feature map of the input hybrid feature interactive self-attention unit is Input the first depth convolution layer to obtain local features in, Represents real multidimensional space, C, H, and W represent the number of channels, length, and width of the first feature map, respectively. Represents three-dimensional data with a shape of C×H×W and a window size of S. Its expression is as follows: local =DwConv 3×3 (X);

[0015] Among them, DwConv 3×3 (·) represents the function of the first depthwise convolutional layer;

[0016] The local feature F localInput channel interaction unit to obtain channel-level dynamic weights Its expression is as follows: ca =CI(F local );

[0017] Where CI(·) represents the function of the channel interaction unit;

[0018] Split the first feature map X into N non-overlapping windows of size S×S Where N = H × W / S 2 , Indicates shape is NS 2 The two-dimensional data of ×C is transformed into non-overlapping windows X by query linear layer, key linear layer and value linear layer respectively. win Converted into query feature graphs respectively Key feature map Sum value feature graph The expression is as follows: Q, K, V = L Q (X win ), L K (X win ), L V (X win );

[0019] Among them, L Q , L K , L V Represent the functions of query linear layer, key linear layer, and value linear layer respectively;

[0020] The data format of the value feature map V is changed from NS 2 ×C is converted to C×H×W and combined with the channel-level dynamic weight W ca Multiply them to adaptively correct the value feature map V, and then restore the data format to NS2×C. The corrected result is recorded as V′;

[0021] Perform calculations on the local window self-attention module to obtain global features The expression is as follows:

[0022] Where T represents the transposed matrix, and Softmax represents the Softmax function;

[0023] The global feature F global The data format is determined by NS 2 ×C is converted to C×H×W and input into the spatial interaction unit to obtain the spatial level dynamic weight Its expression is as follows: sa =SI(Fglobal);

[0024] Where Si(·) represents the function of the spatial interaction unit;

[0025] By adding the spatial level dynamic weight W sa With the global feature F local Multiply to the global feature F local Perform adaptive correction, and the result after correction is recorded as F′ local ;

[0026] Finally, the global feature F local and F′ local Add to obtain mixed features

[0027] Preferably, the mixed-scale feedforward neural network includes a first branch, a second branch and a fifth convolutional layer, the first branch includes a second depth convolutional layer, a first ReLU activation function layer, a third depth convolutional layer and a second ReLU activation function layer connected in sequence, and the second branch includes a fourth depth convolutional layer, a third ReLU activation function layer, a fifth depth convolutional layer and a fourth ReLU activation function layer connected in sequence, wherein the convolution kernel size of the second depth convolutional layer and the fifth depth convolutional layer is 7×7, and the convolution kernel size of the third depth convolutional layer and the fourth depth convolutional layer is 5×5. The specific calculation process is as follows:

[0028] The second feature map of the input mixed-scale feedforward neural network is fed along the channel dimension Divide X' into two equal parts and get the features after division and Indicates the shape The three-dimensional data will and The first and second branches are input respectively for mixed cross feature extraction, and the first cross feature and the second cross feature are output respectively. The first cross feature and the second cross feature are spliced ​​and input into the fifth convolution layer. The output of the fifth convolution layer is added to the second feature map to obtain the mixed scale feature. Its expression is as follows:

[0029] Among them, ReLU(·) represents the ReLU activation function, DwConv 5×5 (·) and DwConv 7×7 (·) denotes the functions of the depthwise convolutional layers with convolution kernels of 5×5 and 7×7, respectively. Conv 1×1 (·) represents the function of the fifth convolution layer with a convolution kernel size of 1×1, [·] represents the splicing operation, represent the first and second characteristics respectively, Represent the first cross feature and the second cross feature respectively.

[0030] As a preference, the specific structure and calculation process of the single-frame image super-resolution model are as follows:

[0031] The shallow feature extraction unit uses the sixth convolutional layer. The calculation process of the shallow feature extraction unit is as follows: F0 = Conv 3×3 (I LR );

[0032] Among them, F0 represents the shallow feature Conv 3×3 (·) represents the function of the sixth convolution layer with a convolution kernel of 3×3, I LR Represents a low-resolution image;

[0033] P hybrid feature interaction Transformer modules are used to extract features, and F0 is transferred to the end of the network using long skip connections. It is added to the output of the Pth hybrid feature interaction Transformer module for residual learning. Its expression is as follows: F i =MF i (F i-1 ), i∈[1,P]; F P0 =MF P (…(MF 2 (MF 1 (F0))))+F0;

[0034] Among them, F i-1 represents the output of the i-1th hybrid feature interaction Transformer module, MF P Represents the function of the Pth hybrid feature interaction Transformer module, MF 1 Represents the function of the first hybrid feature interaction Transformer module, MF 2 Represents the function of the second hybrid feature interaction Transformer module, MF i represents the function of the i-th mixed feature interaction Transformer module, F i represents the output of the i-th hybrid feature interaction Transformer module, F P0 Represents deep features,

[0035] The upsampling reconstruction unit includes a sub-pixel convolution layer with a scale factor of scale and a seventh convolution layer with a convolution kernel of 3×3, and its expression is as follows: SR =Conv 3×3 (fup (F P0 ));

[0036] Among them, f up (·) represents the function of the sub-pixel convolution layer, Conv 3×3 (·) represents the function of the seventh convolutional layer, I SR represents the high-resolution reconstructed image, Represents three-dimensional data with a shape of 3×(H×scale)×(W×scale).

[0037] In a second aspect, the present invention provides a single-frame image super-resolution device based on a hybrid feature interactive Transformer, comprising:

[0038] An image acquisition module is configured to acquire a low-resolution image to be reconstructed;

[0039] A model construction module is configured to construct and train a single-frame image super-resolution model based on a hybrid feature interactive Transformer to obtain a trained single-frame image super-resolution model, wherein the single-frame image super-resolution model includes a shallow feature extraction unit, a deep feature extraction unit, and an upsampling reconstruction unit connected in sequence, and the deep feature extraction unit includes P hybrid feature interactive Transformer modules connected in sequence;

[0040] The reconstruction module is configured to input the low-resolution image to be reconstructed into a trained single-frame image super-resolution model, extract shallow features through a shallow feature extraction unit, input the shallow features into a deep feature extraction unit to extract deep features, and input the deep features into an upsampling reconstruction unit to reconstruct a high-resolution reconstructed image.

[0041] In a third aspect, the present invention provides an electronic device comprising one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] (1) The hybrid feature interaction self-attention unit in the single-frame image super-resolution method based on the hybrid feature interaction Transformer proposed in the present invention adopts a dual-branch structure combined with a bidirectional feature interaction unit. The dual-branch structure introduces an additional deep convolution branch parallel to the local window self-attention unit on the basis of the standard local window self-attention unit, which can enhance the cross-window feature interaction capability of the Transformer. The bidirectional feature interaction unit can provide complementary clues for the dual-branch structure, fully consider the complementarity between different types of features, and can significantly improve information utilization and image super-resolution performance.

[0044] (2) The single-frame image super-resolution method based on hybrid feature interaction Transformer proposed in this invention can overcome the problem that the existing Transformer method ignores the potential correlation between features of different dimensions. By encouraging cross-dimensional feature interaction, the global feature expression ability and detail reconstruction ability of the image super-resolution method are significantly improved.

[0045] (3) Compared with the existing single-frame image super-resolution methods, the single-frame image super-resolution method based on hybrid feature interactive Transformer proposed in this invention has lower parameter count and Flops value, the best comprehensive performance, and can achieve high-performance image super-resolution reconstruction with less computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0047] FIG1 is a diagram of an exemplary device architecture in which an embodiment of the present application may be applied;

[0048] FIG2 is a flow chart of a single-frame image super-resolution method based on a hybrid feature interactive Transformer according to an embodiment of the present application;

[0049] FIG3 is a schematic structural diagram of an efficient local feature extraction unit of a single-frame image super-resolution method based on a hybrid feature interactive Transformer according to an embodiment of the present application;

[0050] FIG4 is a schematic diagram of the structure of a hybrid feature interactive self-attention unit of a single-frame image super-resolution method based on a hybrid feature interactive Transformer according to an embodiment of the present application;

[0051] FIG5 is a schematic diagram of the structure of a hybrid-scale feedforward neural network of a single-frame image super-resolution method based on a hybrid feature interactive Transformer according to an embodiment of the present application;

[0052] FIG6 is a schematic structural diagram of a hybrid feature interactive Transformer module of a single-frame image super-resolution method based on a hybrid feature interactive Transformer according to an embodiment of the present application;

[0053] FIG7 is a structural diagram of a single-frame image super-resolution model based on a hybrid feature interactive Transformer of a single-frame image super-resolution method based on a hybrid feature interactive Transformer according to an embodiment of the present application;

[0054] FIG8 is a schematic diagram of a single-frame image super-resolution device based on a hybrid feature interactive Transformer according to an embodiment of the present application;

[0055] FIG9 is a schematic diagram of the structure of a computer device suitable for implementing an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0056] To make the objectives, technical solutions, and advantages of the present invention more apparent, the present invention will be further described in detail below with reference to the accompanying drawings. It is apparent that the embodiments described are only some, not all, of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are intended to fall within the scope of protection of the present invention.

[0057] FIG1 shows an exemplary device architecture 100 to which a single-frame image super-resolution method based on a hybrid feature interactive Transformer or a single-frame image super-resolution device based on a hybrid feature interactive Transformer according to an embodiment of the present application can be applied.

[0058] As shown in FIG1 , device architecture 100 may include terminal device 1 101, terminal device 2 102, terminal device 3 103, network 104, and server 105. Network 104 is used as a medium for providing communication links between terminal device 1 101, terminal device 2 102, terminal device 3 103, and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0059] It should be noted that the single-frame image super-resolution method based on the hybrid feature interactive Transformer provided in the embodiment of the present application can be executed by the server 105, or by the terminal device 1 101, the terminal device 2 102, and the terminal device 3 103. Accordingly, the single-frame image super-resolution device based on the hybrid feature interactive Transformer can be set in the server 105, or in the terminal device 1 101, the terminal device 2 102, and the terminal device 3 103.

[0060] It should be understood that the number of terminal devices, networks, and servers in FIG1 is merely illustrative. Any number of terminal devices, networks, and servers may be provided as needed. If the processed data does not need to be acquired remotely, the above-described apparatus architecture may not include a network, but only servers or terminal devices.

[0061] FIG2 shows a single-frame image super-resolution method based on a hybrid feature interactive Transformer provided in an embodiment of the present application, comprising the following steps:

[0062] S1, obtaining a low-resolution image to be reconstructed.

[0063] Specifically, a low-resolution image to be reconstructed is collected, where the low-resolution image is a single-frame image.

[0064] S2, construct and train a single-frame image super-resolution model based on a hybrid feature interactive Transformer to obtain a trained single-frame image super-resolution model. The single-frame image super-resolution model includes a shallow feature extraction unit, a deep feature extraction unit, and an upsampling reconstruction unit connected in sequence. The deep feature extraction unit includes P hybrid feature interactive Transformer modules connected in sequence.

[0065] In a specific embodiment, the hybrid feature interaction Transformer module includes an efficient local feature extraction unit, a first normalization layer, a hybrid feature interaction self-attention unit, a second normalization layer and a hybrid scale feedforward neural network. The input of the hybrid feature interaction Transformer module passes through the efficient local feature extraction unit and the first normalization layer in sequence. The output of the first normalization layer is added to the input of the hybrid feature interaction Transformer module to obtain a first hybrid feature. The first hybrid feature passes through the hybrid feature interaction self-attention unit and the second normalization layer in sequence. The output of the second normalization layer is added to the first hybrid feature to obtain a second hybrid feature. The second hybrid feature is input into the hybrid scale feedforward neural network to obtain the output of the hybrid feature interaction Transformer module.

[0066] In a specific embodiment, the efficient local feature extraction unit includes a first shifted convolution layer, a first GeLU activation function layer, a second shifted convolution layer, an SE module, and a third shifted convolution layer connected in sequence. The calculation process of the efficient local feature extraction unit is as follows: H ELF (·)=F shift-conv (F SE (F shift-conv (GeLU(F shift-conv (·))))) ;

[0067] Among them, H ELF (·) represents the function of efficient local feature extraction unit, F shift-conv (·) represents the shifted convolution operation of the first shifted convolution layer, the second shifted convolution layer, or the third shifted convolution layer, F SE (·) represents the function of the SE module, and GeLU(·) represents the GeLU activation function.

[0068] In a specific embodiment, the hybrid feature interaction self-attention unit includes a local window self-attention branch, a deep convolution branch and a bidirectional feature interaction unit, the bidirectional feature interaction unit includes a spatial interaction unit and a channel interaction unit, the channel interaction unit includes a global average pooling layer, a first convolution layer, a first batch normalization layer, a second GeLU activation function layer, a second convolution layer and a first Sigmoid activation function layer connected in sequence, the spatial interaction unit includes a third convolution layer, a second batch normalization layer, a third GeLU activation function layer, a fourth convolution layer and a second Sigmoid activation function connected in sequence, the local window self-attention branch includes a query linear layer, a key linear layer, a value linear layer and a local window self-attention module, the deep convolution branch includes a first deep convolution layer with a convolution kernel size of 3×3, the local features output by the first deep convolution layer are input into the channel interaction unit to obtain channel-level dynamic weights, and the channel-level dynamic weights are input into the local window self-attention branch to adaptively correct the value feature map output by the value linear layer; the global features output by the local window self-attention module are input into the spatial interaction unit to obtain spatial-level dynamic weights, and the spatial-level dynamic weights are input into the deep convolution branch to adaptively correct the local features.

[0069] In a specific embodiment, the calculation process of the mixed feature interactive self-attention unit is as follows: the first feature map of the input mixed feature interactive self-attention unit is Input the first depth convolution layer to obtain local features in, Represents real multidimensional space, C, H, W represent the number of channels, length and width of the first feature map respectively, and the window size is S. Represents three-dimensional data with a shape of C×H×W, and its expression is as follows: local=DwConv 3×3 (X);

[0070] Among them, DwConv 3×3 (·) represents the function of the first depthwise convolutional layer;

[0071] The local feature F local Input channel interaction unit to obtain channel-level dynamic weights Its expression is as follows: ca =CI(F local );

[0072] Where CI(·) represents the function of the channel interaction unit;

[0073] Split the first feature map X into N non-overlapping windows of size S×S Where N = H × W / S 2 , Indicates shape is NS 2 The two-dimensional data of ×C is transformed into non-overlapping windows X by query linear layer, key linear layer and value linear layer respectively. win Converted into query feature graphs respectively Key feature map Sum value feature graph The expression is as follows: Q, K, V = L Q (X win ), LK(X win ), L V (X win );

[0074] Among them, L Q , L K , L V Represent the functions of query linear layer, key linear layer, and value linear layer respectively;

[0075] The data format of the value feature map V is changed from NS 2 ×C is converted to C×H×W and combined with the channel-level dynamic weight W ca Multiply to adaptively correct the value feature map V, and then restore the data format to NS 2 ×C, the corrected result is recorded as V′;

[0076] Perform calculations on the local window self-attention module to obtain global features The expression is as follows:

[0077] Where T represents the transposed matrix, and Softmax represents the Softmax function;

[0078] The global feature Fglobal The data format is determined by NS 2 ×C is converted to C×H×W and input into the spatial interaction unit to obtain the spatial level dynamic weight Its expression is as follows: sa =SI(F global );

[0079] Where SI(·) represents the function of spatial interaction unit;

[0080] By adding the spatial level dynamic weight W sa With the global feature F local Multiply to the global feature F local Perform adaptive correction, and the result after correction is recorded as F′ local ;

[0081] Finally, the global feature F local and F′ local Add to obtain mixed features

[0082] In a specific embodiment, the mixed-scale feedforward neural network includes a first branch, a second branch and a fifth convolutional layer, the first branch includes a second depth convolutional layer, a first ReLU activation function layer, a third depth convolutional layer and a second ReLU activation function layer connected in sequence, and the second branch includes a fourth depth convolutional layer, a third ReLU activation function layer, a fifth depth convolutional layer and a fourth ReLU activation function layer connected in sequence, wherein the convolution kernel size of the second depth convolutional layer and the fifth depth convolutional layer is 7×7, and the convolution kernel size of the third depth convolutional layer and the fourth depth convolutional layer is 5×5. The specific calculation process is as follows:

[0083] The second feature map of the input mixed-scale feedforward neural network is fed along the channel dimension Divide X' into two equal parts and get the features after division and Indicates the shape The three-dimensional data will and The first and second branches are input respectively for mixed cross feature extraction, and the first cross feature and the second cross feature are output respectively. The first cross feature and the second cross feature are spliced ​​and input into the fifth convolution layer. The output of the fifth convolution layer is added to the second feature map to obtain the mixed scale feature. Its expression is as follows:

[0084] Among them, ReLU(·) represents the ReLU activation function, DwConv 5×5 (·) and DwConv 7×7 (·) denotes the functions of the depthwise convolutional layers with convolution kernels of 5×5 and 7×7, respectively. Conv 1×1 (·) represents the function of the fifth convolution layer with a convolution kernel size of 1×1, [·] represents the splicing operation, represent the first and second characteristics respectively, Represent the first cross feature and the second cross feature respectively.

[0085] Specifically, referring to FIG3, an efficient local feature extraction unit can be constructed first. The efficient local feature extraction unit is sequentially composed of a first shifted convolution layer, a first GeLU activation function layer, a second shifted convolution layer, an SE module, and a third shifted convolution layer. The SE module is a Squeeze-Excitation Module. Referring to FIG4, a hybrid feature interaction self-attention unit is constructed. The hybrid feature interaction self-attention unit is constructed on the basis of the standard local window self-attention unit by adding two key designs: (1) a dual-branch structure, including a local window self-attention branch and a deep convolution branch; (2) a bidirectional feature interaction unit. Specifically, by designing a simple dual-branch structure, a deep convolution layer parallel to the standard local window self-attention unit is introduced to enhance cross-window feature interaction. The bidirectional feature interaction unit includes a spatial interaction unit and a channel interaction unit. The information of the deep convolution branch first flows into the local window self-attention branch through the spatial interaction unit; then, the information of the local window self-attention branch flows into the deep convolution branch through the spatial interaction unit. Therefore, the bidirectional feature interaction unit proposed in the embodiment of the present application can provide complementary clues for the dual-branch structure to enhance information utilization. Specifically, the channel interaction unit is composed of a global average pooling layer, a first convolution layer with a convolution kernel size of 3×3, a first batch of normalization layers, a second GeLU activation function layer, a second convolution layer with a convolution kernel size of 3×3, and a first Sigmoid activation function layer cascaded. The spatial interaction unit is composed of a third convolution layer with a convolution kernel size of 3×3, a second batch of normalization layers, a third GeLU activation function layer, a fourth convolution layer with a convolution kernel size of 3×3, and a second Sigmoid activation function cascaded. Then, referring to Figure 5, a mixed-scale feedforward neural network is constructed, which includes two multi-scale deep convolution branches. The two multi-scale deep convolution branches realize mixed feature extraction by alternately using a deep convolution layer with a convolution kernel size of 5×5 and a deep convolution layer with a convolution kernel size of 7×7. Each deep convolution layer is connected to a ReLU activation function layer, and finally the outputs of the two branches are fused using the fifth convolution layer with a convolution kernel size of 1×1 to obtain a mixed-scale feature.

[0086] Furthermore, referring to Figure 6, an efficient local feature extraction unit, a hybrid feature interactive self-attention unit and a mixed-scale feedforward neural network are integrated to construct a hybrid feature interactive Transformer module. The hybrid feature interactive Transformer module is composed of an efficient local feature extraction unit, a first-layer normalization layer, a hybrid feature interactive self-attention unit, a second-layer normalization layer, and a mixed-scale feedforward neural network in cascade.

[0087] Finally, referring to FIG7 , a single-frame image super-resolution model based on a hybrid feature interaction Transformer is constructed and trained to obtain a trained single-frame image super-resolution model.

[0088] S3: Input the low-resolution image to be reconstructed into the trained single-frame image super-resolution model, extract shallow features through the shallow feature extraction unit, input the shallow features into the deep feature extraction unit to extract deep features, and input the deep features into the upsampling reconstruction unit to reconstruct a high-resolution reconstructed image.

[0089] In a specific embodiment, the specific structure and calculation process of the single-frame image super-resolution model are as follows:

[0090] The shallow feature extraction unit uses the sixth convolutional layer. The calculation process of the shallow feature extraction unit is as follows: F0 = Conv 3×3 (I LR );

[0091] Among them, F0 represents the shallow feature Conv 3×3 (·) represents the function of the sixth convolution layer with a convolution kernel of 3×3, I LR Represents a low-resolution image;

[0092] P hybrid feature interaction Transformer modules are used to extract features, and F0 is transferred to the end of the network using long skip connections. It is added to the output of the Pth hybrid feature interaction Transformer module for residual learning. Its expression is as follows: F i =MF i (F i-1 ), i∈[1,P]; F P0 =MF P (…(MF 2 (MF 1 (F0))))+F0;

[0093] Among them, F i-1 represents the output of the i-1th hybrid feature interaction Transformer module, MF PRepresents the function of the Pth hybrid feature interaction Transformer module, MF 1 Represents the function of the first hybrid feature interaction Transformer module, MF 2 Represents the function of the second hybrid feature interaction Transformer module, MF i represents the function of the i-th mixed feature interaction Transformer module, F i represents the output of the i-th hybrid feature interaction Transformer module, F P0 Represents deep features,

[0094] The upsampling reconstruction unit includes a sub-pixel convolution layer with a scale factor of scale and a seventh convolution layer with a convolution kernel of 3×3, and its expression is as follows: SR =Conv 3×3 (f up (F P0 ));

[0095] Among them, f up (·) represents the function of the sub-pixel convolution layer, Conv 3×3 (·) represents the function of the seventh convolutional layer, I SR represents the high-resolution reconstructed image, Represents three-dimensional data with a shape of 3×(H×scale)×(W×scale).

[0096] Specifically, the trained single-frame image super-resolution module is used to reconstruct the low-resolution image to be reconstructed to obtain a reconstruction result. The trained single-frame image super-resolution module consists of three parts: a shallow feature extraction unit, a deep feature extraction unit, and an upsampling reconstruction unit. For a given low-resolution image to be reconstructed, The scaling factor scale is used as input, where the value of scale is the required magnification, for example, scale is 2, 3, 4 or 8.

[0097] A single-frame image super-resolution method based on a hybrid feature interaction Transformer proposed in an embodiment of the present application is compared with the most advanced single-frame image super-resolution method. In this comparative experiment, DIV2K is used as the training set, Set5, Se14, BSD100 and Urban100 are used as test sets, and the target scaling factor is 2. The quantitative indicators PSNR and SSIM are used to evaluate the quality of the reconstructed image. Higher PSNR and SSIM values ​​correspond to higher SR performance. The quantitative indicators parameter quantity (Params) and Flops are used to measure the model scale and execution speed. The lower the parameter quantity, the smaller the model scale, and the lower the Flops value, the faster the model execution speed. In order to meet the needs of real application scenarios, designing an image super-resolution method with low parameter quantity and low Flops value that can generate reconstructed images with high PSNR and SSIM is an important goal in the field of image super-resolution. As shown in Table 1, compared with other methods, the method proposed in the embodiment of the present application achieves the highest PSNR and SSIM in the four test sets with the lowest parameter quantity and lowest Flops value. Therefore, Table 1 fully illustrates that the single-frame image super-resolution method based on hybrid feature interaction Transformer proposed in the embodiment of the present application demonstrates the best overall performance compared with other methods.

[0098] Table 1

[0099] The above steps S1-S3 do not merely represent the order of the steps, but are symbolic representations of the steps.

[0100] Further referring to Figure 8, as an implementation of the methods shown in the above figures, the present application provides an embodiment of a single-frame image super-resolution device based on a hybrid feature interactive Transformer. The device embodiment corresponds to the method embodiment shown in Figure 2, and the device can be specifically applied to various electronic devices.

[0101] The present invention provides a single-frame image super-resolution device based on a hybrid feature interactive Transformer, comprising:

[0102] An image acquisition module 1 is configured to acquire a low-resolution image to be reconstructed;

[0103] Model construction module 2 is configured to construct and train a single-frame image super-resolution model based on a hybrid feature interactive Transformer to obtain a trained single-frame image super-resolution model, wherein the single-frame image super-resolution model includes a shallow feature extraction unit, a deep feature extraction unit, and an upsampling reconstruction unit connected in sequence, and the deep feature extraction unit includes P hybrid feature interactive Transformer modules connected in sequence;

[0104] The reconstruction module 3 is configured to input the low-resolution image to be reconstructed into the trained single-frame image super-resolution model, extract shallow features through the shallow feature extraction unit, input the shallow features into the deep feature extraction unit to extract deep features, and input the deep features into the upsampling reconstruction unit to reconstruct a high-resolution reconstructed image.

[0105] Reference is now made to FIG9 , which illustrates a schematic diagram of the structure of a computer device 900 suitable for implementing an electronic device (e.g., the server or terminal device shown in FIG1 ) according to an embodiment of the present application. The electronic device shown in FIG9 is merely an example and should not limit the functionality or scope of use of the embodiments of the present application.

[0106] As shown in FIG9 , a computer device 900 includes a central processing unit (CPU) 901 and a graphics processing unit (GPU) 902, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 903 or a program loaded from a storage unit 909 into a random access memory (RAM) 904. Various programs and data required for the operation of the computer device 900 are also stored in the RAM 904. The CPU 901, GPU 902, ROM 903, and RAM 904 are connected to each other via a bus 905. An input / output (I / O) interface 906 is also connected to the bus 905.

[0107] The following components are connected to the I / O interface 906: an input section 907 including a keyboard, a mouse, etc.; an output section 908 including a display such as a liquid crystal display (LCD), a speaker, etc.; a storage section 909 including a hard disk, etc.; and a communication section 910 including a network interface card such as a LAN card, a modem, etc. The communication section 910 performs communication processing via a network such as the Internet. A drive 911 may also be connected to the I / O interface 906 as needed. A removable medium 912, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed in the drive 911 as needed, so that a computer program read therefrom can be installed into the storage section 909 as needed.

[0108] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 910, and / or installed from the removable medium 912. When the computer program is executed by the central processing unit (CPU) 901 and the graphics processing unit (GPU) 902, the above-mentioned functions defined in the method of the present application are executed.

[0109] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable medium, or any combination thereof. Computer-readable media may be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, apparatuses, or components, or any combination thereof. More specific examples of computer-readable media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution device, apparatus, or component. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution apparatus, device, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical cable, RF, or any suitable combination thereof.

[0110] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based device that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0112] The modules involved in the embodiments described in this application may be implemented in software or hardware, and may also be set in a processor.

[0113] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application. Industrial Applicability

[0114] The present invention discloses a single-frame image super-resolution method and device based on a hybrid feature interactive Transformer. The method and device include an extraction unit comprising P hybrid feature interactive Transformer modules connected in sequence. A low-resolution image is input into a trained single-frame image super-resolution model, shallow features are extracted by a shallow feature extraction unit, the shallow features are input into a deep feature extraction unit to extract deep features, and the deep features are input into an upsampling reconstruction unit to reconstruct a high-resolution reconstructed image. This solves the problem that the Transformer SR method ignores the potential correlation between features of different dimensions, thereby affecting the reconstruction performance.

Claims

1. A single-frame image super-resolution method based on a hybrid feature interaction Transformer, characterized in that The method includes the following steps: Obtain a low-resolution image to be reconstructed; Construct and train a single-frame image super-resolution model based on a hybrid feature interaction Transformer to obtain a trained single-frame image super-resolution model. The single-frame image super-resolution model includes a shallow feature extraction unit, a deep feature extraction unit, and an upsampling and reconstruction unit connected in sequence. The deep feature extraction unit includes P hybrid feature interaction Transformer modules connected in sequence; Input the low-resolution image to be reconstructed into the trained single-frame image super-resolution model, extract shallow features through the shallow feature extraction unit, input the shallow features into the deep feature extraction unit to extract deep features, and input the deep features into the upsampling and reconstruction unit to reconstruct a high-resolution reconstructed image.

2. The single-frame image super-resolution method based on a hybrid feature interaction Transformer according to claim 1, wherein The hybrid feature interaction Transformer module includes an efficient local feature extraction unit, a first normalization layer, a hybrid feature interaction self-attention unit, a second normalization layer, and a hybrid-scale feed-forward neural network. The input of the hybrid feature interaction Transformer module passes through the efficient local feature extraction unit and the first normalization layer in sequence. The output of the first normalization layer is added to the input of the hybrid feature interaction Transformer module to obtain a first hybrid feature. The first hybrid feature passes through the hybrid feature interaction self-attention unit and the second normalization layer in sequence. The output of the second normalization layer is added to the first hybrid feature to obtain a second hybrid feature. The second hybrid feature is input into the hybrid-scale feed-forward neural network to obtain the output of the hybrid feature interaction Transformer module.

3. The single-frame image super-resolution method based on the hybrid feature interaction Transformer according to claim 2, wherein The efficient local feature extraction unit includes a first displacement convolutional layer, a first GeLU activation function layer, a second displacement convolutional layer, an SE module, and a third displacement convolutional layer connected in sequence. The calculation process of the efficient local feature extraction unit is as follows: H ELF (·) = F shift-conv (F SE (F shift-conv (GeLU(F shift-conv (·))))) ; Among them, H ELF (·) represents the function of the efficient local feature extraction unit, F shift-conv (·) represents the The displacement convolution operation of the first displacement convolution layer, the second displacement convolution layer, or the third displacement convolution layer, F SE (·) represents the function of the SE module, and GeLU(·) represents the GeLU activation function.

4. The single-frame image super-resolution method based on the hybrid feature interaction Transformer according to claim 2, wherein The hybrid feature interaction self-attention unit includes a local window self-attention branch, a depth convolution branch, and a bidirectional feature interaction unit. The bidirectional feature interaction unit includes a spatial interaction unit and a channel interaction unit. The channel interaction unit includes a global average pooling layer, a first convolutional layer, a first batch normalization layer, a second GeLU activation function layer, a second convolutional layer, and a first Sigmoid activation function layer connected in sequence. The spatial interaction unit includes a third convolutional layer, a second batch normalization layer, a third GeLU activation function layer, a fourth convolutional layer, and a second Sigmoid activation function connected in sequence. The local window self-attention branch includes a query linear layer, a key linear layer, a value linear layer, and a local window self-attention module. The depth convolution branch includes a first depth convolutional layer with a convolution kernel size of 3×3. The local features output by the first depth convolutional layer are input into the channel interaction unit to obtain channel-level dynamic weights. The channel-level dynamic weights are input into the local window self-attention branch to adaptively correct the value feature map output by the value linear layer; The global features output by the local window self-attention module are input into the spatial interaction unit to obtain spatial-level dynamic weights. The spatial-level dynamic weights are input into the depth convolution branch to adaptively correct the local features.

5. The single-frame image super-resolution method based on the hybrid feature interaction Transformer according to claim 4, wherein The calculation process of the hybrid feature interaction self-attention unit is as follows: The first feature map input to the hybrid feature interaction self-attention unit Input the first depth convolution layer to obtain the local features Among them, Represents a multi-dimensional real space, where C, H, and W respectively represent the number of channels, length, and width of the first feature map, represents three-dimensional data with a shape of C×H×W, and the window size is S. Its expression is as follows: F local = DwConv 3×3 (X); Among them, DwConv 3×3 (·) represents the function of the first depth convolution layer; Input the local feature F local into the channel interaction unit to obtain channel-level dynamic weights Its expression is as follows: W ca = CI(F local ); Among them, CI(·) represents the function of the channel interaction unit; Partition the first feature map X into N non-overlapping windows with a window size of S×S where N = H × W / S 2 , Indicating a two-dimensional data of shape NS 2 ×C, and respectively converting the non-overlapping window X through the query linear layer, the key linear layer, and the value linear layer win into query feature maps respectively Key feature diagram Sum feature map Its expression is as follows: Q, K, V = L Q (X win ), L K (X win ), L V (X win ) Among them, L Q , L K , L V respectively represent the functions of the query linear layer, the key linear layer, and the value linear layer; Convert the data format of the value feature map V from NS 2 ×C to C×H×W and multiply it with the channel-level dynamic weight W ca to adaptively correct the value feature map V, and then restore the data format to NS 2 ×C. Denote the corrected result as V'. Perform the calculation of the local window self-attention module to obtain global features The expression is as follows: Among them, T represents the transpose matrix, and Softmax represents the Softmax function; Convert the data format of the global feature F global from NS 2 ×C to C×H×W, and input it into the spatial interaction unit to obtain the spatial-level dynamic weight Its expression is as follows: W sa = SI(F global ) Among them, SI(·) represents the function of the spatial interaction unit; By multiplying the spatial-level dynamic weight W sa with the global feature F local to adaptively correct the global feature F local , and the corrected result is denoted as F′ local ; Finally, add the global feature F local and F′ local to obtain a hybrid feature 6. The single-frame image super-resolution method based on a hybrid feature interaction Transformer according to claim 2, wherein The hybrid scale feed-forward neural network includes a first branch, a second branch, and a fifth convolutional layer. The first branch includes a second depth convolutional layer, a first ReLU activation function layer, a third depth convolutional layer, and a second ReLU activation function layer connected in sequence. The second branch includes a fourth depth convolutional layer, a third ReLU activation function layer, a fifth depth convolutional layer, and a fourth ReLU activation function layer connected in sequence. Among them, the convolutional kernels of the second depth convolutional layer and the fifth depth convolutional layer are 7×7, and the convolutional kernels of the third depth convolutional layer and the fourth depth convolutional layer are 5×5. The specific calculation process is as follows: Along the channel dimension, the second feature map input to the hybrid-scale feedforward neural network Divide X' into two equal parts to obtain the divided features and Indicates a shape of 3D data of, will and The first branch and the second branch are respectively input for hybrid cross - feature extraction, and the first cross - feature and the second cross - feature are respectively output. After splicing the first cross - feature and the second cross - feature, they are input into the fifth convolutional layer. The output of the fifth convolutional layer is added to the second feature map to obtain the hybrid scale feature Its expression is as follows: where ReLU(·) represents the ReLU activation function, DwConv 5×5 (·) and DwConv 7×7 (·) represent the functions of depth convolution layers with convolution kernels of 5×5 and 7×7 respectively, Conv 1×1 (·) represents the function of the fifth convolution layer with a convolution kernel size of 1×1, and [·] represents the concatenation operation work respectively represent the first feature and the second feature, respectively represent the first cross feature and the second cross feature.

7. The single-frame image super-resolution method based on a hybrid feature interaction Transformer according to claim 1, wherein The specific structure and calculation process of the single-frame image super-resolution model are as follows: The shallow feature extraction unit uses a sixth convolutional layer. The calculation process of the shallow feature extraction unit is as follows: F0 = Conv 3×3 (I LR ); Among them, F0 represents the shallow features Conv 3×3 (·) represents the function of the sixth convolutional layer with a 3×3 convolutional kernel, and I LR represents the low-resolution image; Extract features using P of the hybrid feature interaction Transformer modules, and use a long skip connection to pass F0 to the end of the network, and add it to the output of the Pth hybrid feature interaction Transformer module for residual learning. Its expression is as follows: F i = MF i (F i-1 ), i ∈ [1, P]; F P0 = MF P (…(MF 2 (MF 1 (F0)))) + F0; Among them, F i-1 represents the output of the (i - 1)-th hybrid feature interaction Transformer module, MF P represents the function of the P-th hybrid feature interaction Transformer module, MF 1 represents the function of the 1st hybrid feature interaction Transformer module, MF 2 represents the function of the 2nd hybrid feature interaction Transformer module, MF i represents the function of the i-th hybrid feature interaction Transformer module, F i represents the output of the i-th hybrid feature interaction Transformer module, F P0 represents the deep feature, The upsampling and reconstruction unit includes a sub-pixel convolutional layer with a scale factor of scale and a seventh convolutional layer with a convolution kernel of 3×3 Its expression is as follows: I SR = Conv 3×3 (f up (F P0 )); Among them, f up (·) represents the function of the sub-pixel convolutional layer, Conv 3×3 (·) represents the function of the seventh convolutional layer, I SR represents the high-resolution reconstructed image, represents three-dimensional data with a shape of 3×(H×scale)×(W×scale).

8. A single-frame image super-resolution device based on a hybrid feature interaction Transformer, which applies the method described in any one of claims 1-7, characterized in that includes: An image acquisition module configured to acquire a low-resolution image to be reconstructed; A model construction module configured to construct and train a single-frame image super-resolution model based on the hybrid feature interaction Transformer to obtain a trained single-frame image super-resolution model. The single-frame image super-resolution model includes a shallow feature extraction unit, a deep feature extraction unit, and an upsampling and reconstruction unit connected in sequence. The deep feature extraction unit includes P hybrid feature interaction Transformer modules connected in sequence; A reconstruction module configured to input the low-resolution image to be reconstructed into the trained single-frame image super-resolution model, extract shallow features through the shallow feature extraction unit, input the shallow features into the deep feature extraction unit to extract deep features, input the deep features into the upsampling and reconstruction unit, and reconstruct a high-resolution reconstructed image.

9. The single-frame image super-resolution device based on the hybrid feature interaction Transformer according to claim 8, wherein The hybrid feature interaction Transformer module includes an efficient local feature extraction unit, a first layer of normalization layer, a hybrid feature interaction self-attention unit, a second layer of normalization layer, and a hybrid scale feed-forward neural network.

10. An electronic device, including: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Video super-resolution based on enhanced deep feature extraction and residual up-down sampling blocks

    CN114387161A

  • Single image super-resolution reconstruction method and system based on CNN and Transform hybrid network

    CN114926337A

  • Single-frame image super-resolution method and system based on cross-layer mixed attention Transform

    CN117173025A

  • Image super-resolution method and device based on cross attention mechanism and Swin-Transform

    CN117237197A

Cited By

  • Ball mill granularity soft measurement method based on large time sequence model

    CN120449128A

  • Fetal echocardiography information processing method and system for congenital heart disease

    CN120471915A

  • Multi-level attention visible light guided infrared image super-resolution method and system

    CN120495087A

  • OCT image super-resolution reconstruction method and system based on generative adversarial network

    CN120525723A

  • OCT image super-resolution reconstruction method and system based on generative adversarial network

    CN120525723B