Super-resolution image generation method

By introducing multi-layer feature fusion and detail control mechanisms into the super-resolution image generation model, the problems of insufficient details and artifacts of super-resolution image in the prior art are solved, and high-quality image generation is achieved.

CN120147134APending Publication Date: 2025-06-13CHENGDU UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510308126.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing deep learning methods based on diffusion models are insufficient in detail and artifacts when generating super-resolution images.

Method used

A super-resolution image generation model consisting of pre-trained encoder, noise generation module, fusion noise module, detail control network, rotation feature extraction module, Unet network and decoder is adopted. The model enhances image detail and reduces artifacts through multi-layer stitching and feature fusion.

Benefits of technology

The generated super-resolution images are rich in detail and have no artifacts, which significantly improves image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147134A_ABST
    Figure CN120147134A_ABST
Patent Text Reader

Abstract

The invention discloses a super-resolution image generation method, which comprises the following steps of: inputting an image of which the resolution is to be improved into a trained super-resolution image generation model to obtain an image of which the resolution is improved; wherein the trained super-resolution image generation model comprises a pre-trained encoder, a noise generation module, a fusion noise module, a first channel splicing module, a trained detail control network, a trained rotation feature extraction module, a pre-trained Unet network and a pre-trained decoder. The super-resolution image generated by the method is rich in details and does not have artifacts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of super-resolution image generation, and more specifically, to a super-resolution image generation method. Background Art

[0002] Super-Resolution (SR) refers to the process of using optics and related optical knowledge to restore image details and other data information based on known image information. Simply put, it is to increase the resolution of an image to prevent its image quality from deteriorating and even obtain higher image quality. The methods of super-resolution include traditional methods and deep learning methods. Compared with traditional methods, deep learning methods are far ahead in performance and have better performance in image super-resolution.

[0003] In recent years, deep learning methods based on diffusion models have gradually outperformed other deep learning methods and achieved better performance. However, there are still the following problems: 1) The generated super-resolution images lack details; 2) There are artifacts in the generated super-resolution images.

[0004] Therefore, how to provide a super-resolution image generation method whose generated super-resolution images are rich in details and free of artifacts is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a super-resolution image generation method.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] A super-resolution image generation method includes the following steps:

[0008] Input the image to be upscaled into a trained super-resolution image generation model to obtain an image with increased resolution;

[0009] Among them, the trained super-resolution image generation model includes a pre-trained encoder, a noise generation module, a fusion noise module, a first channel splicing module, a trained detail control network, a trained rotation feature extraction module, a pre-trained Unet network, and a pre-trained decoder;

[0010] The output end of the pre-trained encoder and the output end of the noise generation module are both connected to the input end of the fusion noise module;

[0011] The output end of the pre-trained encoder and the output end of the noise generation module are both connected to the input end of the first channel splicing module;

[0012] The output end of the first channel splicing module is successively connected to the input end of the pre-trained Unet network through the trained detail control network and the trained rotation feature extraction module;

[0013] The output end of the fusion noise module is connected to the input end of the pre-trained Unet network;

[0014] The output end of the pre-trained Unet network is connected to the input end of the pre-trained decoder.

[0015] Preferably, the fusion noise module performs the following operations:

[0016]

[0017]

[0018] α i = 1 - β i ;

[0019] where Z n represents the noise latent variable output by the fusion noise module; β i represents the variance of the Gaussian noise generated by the noise generation module; β i ∈(0, 1); N represents the number of selected β i ; ∈ represents the Gaussian noise generated by the noise generation module; Z represents the latent variable output by the pre-trained encoder.

[0020] Preferably, the trained detail control network includes a detail retention module and a detail control module;

[0021] The input end of the detail retention module is connected to the output end of the first channel splicing module;

[0022] The output end of the detail retention module is connected to the input end of the trained rotation feature extraction module through the detail control module;

[0023] The detail retention module includes a first 1*1 convolutional layer, a first 7*7 convolutional layer, a second 1*1 convolutional layer, a convolutional feedforward neural network, an average pooling processing unit, and a max pooling processing unit;

[0024] The input end of the first 1*1 convolutional layer is connected to the output end of the first channel splicing module;

[0025] The output end of the first 1*1 convolutional layer is successively connected to the input ends of the first 7*7 convolutional layer, the second 1*1 convolutional layer, the convolutional feedforward neural network, the average pooling processing unit, and the max pooling processing unit;

[0026] The output end of the maximum pooling processing unit is connected to the input end of the detail control module;

[0027] The detail control module has the same downsampling structure as the Unet network.

[0028] Preferably, the average pooling processing unit includes an average pooling layer, a bilinear interpolation layer, a feature subtraction module, a self-attention module, and a feature fusion module;

[0029] The input end of the average pooling layer is connected to the output end of the convolutional feedforward neural network;

[0030] The output end of the average pooling layer is connected to the input end of the bilinear interpolation layer;

[0031] The input end of the average pooling layer and the output end of the bilinear interpolation layer are both connected to the input end of the feature subtraction module;

[0032] The output end of the feature subtraction module is connected to the input end of the feature fusion module through the self-attention module;

[0033] The output end of the average pooling layer is connected to the input end of the feature fusion module.

[0034] Preferably, the maximum pooling processing unit replaces the average pooling layer in the average pooling processing unit with a maximum pooling layer, and the remaining structure is the same as that of the average pooling processing unit.

[0035] Preferably, the self-attention module includes a convolutional module, a first multiplication operation unit, a softmax layer, and a second multiplication operation unit;

[0036] The convolutional module is used to perform three 1*1 convolutional operations on the output of the feature subtraction module, and outputs Q value, K value, and V value;

[0037] The first multiplication operation unit is used to perform a multiplication operation on the Q value and the K value, and outputs the product Q*K;

[0038] The softmax layer is used to perform a softmax operation on the product Q*K, and outputs softmax(Q*K);

[0039] The second multiplication operation unit performs a multiplication operation on softmax(Q*K) and the V value, and outputs the product softmax(Q*K)*V, and the product softmax(Q*K)*V is the output of the self-attention module.

[0040] Preferably, the feature fusion module includes a first MBB module, a second MBB module, a third MBB module, an upsampling module, a fourth MBB module, a second channel splicing module, a first 3×3 convolutional layer, a channel attention layer, and a third 1×1 convolutional layer;

[0041] The input end of the first MBB module is connected to the input end of the bilinear interpolation layer;

[0042] The output end of the first MBB module is sequentially connected to the input end of the second channel splicing module through the second MBB module, the third MBB module, and the upsampling module;

[0043] The input end of the fourth MBB module is connected to the output end of the self-attention module;

[0044] The output end of the fourth MBB module is sequentially connected to the input end of the third 1×1 convolutional layer through the first 3×3 convolutional layer and the channel attention layer;

[0045] The output end of the third 1×1 convolutional layer is the output end of the feature fusion module.

[0046] Preferably, the first MBB module, the second MBB module, the third MBB module, and the fourth MBB module have the same structure, and each includes a second 3×3 convolutional layer, a third 3×3 convolutional layer, a fourth 3×3 convolutional layer, a fifth 3×3 convolutional layer, and an addition operation unit;

[0047] The input end of the second 3×3 convolutional layer is connected to the input end of the third 3×3 convolutional layer as the input end of the MBB module;

[0048] The output end of the third 3×3 convolutional layer is sequentially connected to the input end of the addition operation unit through the fourth 3×3 convolutional layer and the fifth 3×3 convolutional layer;

[0049] The output end of the second 3×3 convolutional layer is connected to the input end of the addition operation unit;

[0050] The output end of the addition operation unit is the output end of the MBB module.

[0051] Preferably, the rotation feature extraction module includes a first branch, a second branch, a third branch, and an arithmetic mean operation unit;

[0052] The first branch includes a first Z-pool layer, a second 7×7 convolution, a first batch norm layer, a first sigmoid layer, and a third multiplication operation unit that are sequentially connected end to end from input to output;

[0053] The input end of the first Z-pool layer and the output end of the first sigmoid layer are both connected to the input end of the third multiplication operation unit;

[0054] The second branch includes an H-axis counterclockwise rotation unit, a second Z-pool layer, a third 7*7 convolution, a second batch norm layer, a second sigmoid layer, a fourth multiplication operation unit, and an H-axis clockwise rotation unit that are connected end to end in sequence from input to output; among them, the H-axis counterclockwise rotation unit is used to rotate its input 90 degrees counterclockwise along the H axis; the H-axis clockwise rotation unit is used to rotate its input 90 degrees clockwise along the H axis;

[0055] The output end of the H-axis counterclockwise rotation unit and the output end of the second sigmoid layer are both connected to the input end of the fourth multiplication operation unit;

[0056] The third branch includes a W-axis counterclockwise rotation unit, a third Z-pool layer, a fourth 7*7 convolution, a third batch norm layer, a third sigmoid layer, a fifth multiplication operation unit, and a W-axis clockwise rotation unit that are connected end to end in sequence from input to output; among them, the W-axis counterclockwise rotation unit is used to rotate its input 90 degrees counterclockwise along the W axis; the W-axis clockwise rotation unit is used to rotate its input 90 degrees clockwise along the W axis;

[0057] The output end of the W-axis counterclockwise rotation unit and the output end of the third sigmoid layer are both connected to the input end of the fifth multiplication operation unit;

[0058] The output ends of the third multiplication operation unit, the H-axis clockwise rotation unit, and the W-axis clockwise rotation unit are all connected to the input end of the arithmetic mean operation unit;

[0059] The output end of the arithmetic mean operation unit is the output end of the rotation feature extraction module;

[0060] The input ends of the first Z-pool layer, the H-axis counterclockwise rotation unit, and the W-axis counterclockwise rotation unit are the input ends of the rotation feature extraction module.

[0061] Preferably, the trained super-resolution image generation model is obtained based on the following steps:

[0062] Obtain a training data set; the training data set contains a number of low-resolution images;

[0063] Input the dataset into the super-resolution image generation model for training, and minimize the loss function to update the network parameters of the detail control network and the rotation feature extraction module, so as to obtain the trained detail control network and the trained rotation feature extraction module, and then the trained super-resolution image generation model can be obtained; where the expression of the loss function is ∈ represents the Gaussian noise generated by the noise generation module; ∈ θ (x) represents the predicted noise output by the pre-trained Unet network; x represents the noise latent feature output by the fusion noise module after inputting the training image; N(0,1) represents the Gaussian noise that satisfies the standard normal distribution; E represents taking the average value.

[0064] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a super-resolution image generation method, and the generated super-resolution image is rich in details and has no artifacts. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0066] Figure 1 It is a schematic structural diagram of the trained super-resolution image generation model provided by the present invention;

[0067] Figure 2 It is a schematic structural diagram of the trained detail control network provided by the present invention;

[0068] Figure 3 It is a schematic structural diagram of the detail retention network provided by the present invention;

[0069] Figure 4 It is a schematic structural diagram of the self-attention module provided by the present invention;

[0070] Figure 5 It is a schematic structural diagram of the feature fusion module provided by the present invention;

[0071] Figure 6 It is a schematic structural diagram of the MBB module provided by the present invention;

[0072] Figure 7 It is a schematic structural diagram of the trained rotation feature extraction module provided by the present invention;

[0073] Figure 8 It is a schematic structural diagram of the pre-trained Unet network provided by the present invention. Specific Embodiments

[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0075] As Figures 1 - 8 shown, an embodiment of the present invention discloses a super-resolution image generation method, including the following steps:

[0076] Input the image to be upscaled into a trained super-resolution image generation model (as Figure 1 shown) to obtain an image with increased resolution;

[0077] Among them, the trained super-resolution image generation model includes a pre-trained encoder, a noise generation module, a fusion noise module, a first channel splicing module, a trained detail control network, a trained rotation feature extraction module, a pre-trained Unet network, and a pre-trained decoder;

[0078] The output end of the pre-trained encoder and the output end of the noise generation module are both connected to the input end of the fusion noise module;

[0079] The output end of the pre-trained encoder and the output end of the noise generation module are both connected to the input end of the first channel splicing module;

[0080] The output end of the first channel splicing module is sequentially connected to the input end of the pre-trained Unet network through the trained detail control network and the trained rotation feature extraction module;

[0081] The output end of the fusion noise module is connected to the input end of the pre-trained Unet network;

[0082] The output end of the pre-trained Unet network is connected to the input end of the pre-trained decoder.

[0083] Specifically:

[0084] Input the image to be upscaled into the pre-trained encoder to obtain a latent variable Z;

[0085] Input the latent variable Z and the Gaussian noise generated by the noise generation module into the fusion noise module for corresponding fusion operations to obtain a noise latent variable Z n ;

[0086] The fusion noise module specifically performs the following fusion operations to obtain the noise latent variable Z n :

[0087]

[0088]

[0089] α i = 1 - β i ;

[0090] where Z n represents the noise latent variable output by the fusion noise module; β i represents the variance of the Gaussian noise generated by the noise generation module; β i ∈(0, 1); N represents the number of selected β i ; ∈ represents the Gaussian noise generated by the noise generation module; Z represents the latent variable output by the pre-trained encoder.

[0091] Furthermore, the Gaussian noise generated by the noise generation module is Gaussian noise that satisfies the standard normal distribution.

[0092] The noise latent variable Z n and the latent variable Z are input into the first-channel splicing module for splicing in the channel dimension to obtain the spliced feature Z p ;

[0093] The spliced feature Z p is input into the trained detail control network to obtain the additional detail feature Z c ;

[0094] The additional detail feature Z c is input into the trained rotation feature extraction module to obtain the rotation feature Z z ;

[0095] The rotation feature Z z and the noise latent variable Z n are input into the pre-trained Unet network to obtain the predicted noise and the denoised feature; the denoised feature is equal to the noise latent variable Z n minus the predicted noise;

[0096] The denoised feature is input into the pre-trained decoder to obtain an image with improved resolution.

[0097] Furthermore, the trained detail control network includes a detail retention module and a detail control module;

[0098] The input end of the detail retention module is connected to the output end of the first-channel splicing module;

[0099] The output end of the detail retention module is connected to the input end of the trained rotation feature extraction module through the detail control module;

[0100] The detail retention module includes a first 1×1 convolutional layer, a first 7×7 convolutional layer, a second 1×1 convolutional layer, a convolutional feedforward neural network, an average pooling processing unit, and a max pooling processing unit;

[0101] The input end of the first 1×1 convolutional layer is connected to the output end of the first channel splicing module;

[0102] The output end of the first 1×1 convolutional layer is sequentially connected to the input ends of the first 7×7 convolutional layer, the second 1×1 convolutional layer, the convolutional feedforward neural network, the average pooling processing unit, and the max pooling processing unit;

[0103] The output end of the max pooling processing unit is connected to the input end of the detail control module;

[0104] The detail control module has the same downsampling structure as the Unet network.

[0105] Further, the average pooling processing unit includes an average pooling layer, a bilinear interpolation layer, a feature subtraction module, a self-attention module, and a feature fusion module;

[0106] The input end of the average pooling layer is connected to the output end of the convolutional feedforward neural network;

[0107] The output end of the average pooling layer is connected to the input end of the bilinear interpolation layer;

[0108] The input ends of the average pooling layer and the output end of the bilinear interpolation layer are both connected to the input end of the feature subtraction module;

[0109] The output end of the feature subtraction module is connected to the input end of the feature fusion module through the self-attention module;

[0110] The output end of the average pooling layer is connected to the input end of the feature fusion module.

[0111] Further, the max pooling processing unit replaces the average pooling layer in the average pooling processing unit with a max pooling layer, and the remaining structure is the same as that of the average pooling processing unit.

[0112] Specifically: the concatenated feature Z pIt is successively processed by a first 1*1 convolutional layer, a first 7*7 convolutional layer, a second 1*1 convolutional layer, and a convolutional feedforward neural network to obtain an input feature F (with a size of H*W*C); the input feature F is input into an average pooling layer for average pooling operation to obtain a low-frequency feature FL = Avgpool(F) (with a size of H / 2*W / 2*C); where Avgpool represents the average pooling operation;

[0113] The low-frequency feature FL is input into a bilinear interpolation layer for bilinear interpolation operation to upsample the low-frequency feature FL to H*W*C, obtaining an upsampled feature Upsample(FL); where Upsample represents the bilinear interpolation operation;

[0114] The upsampled feature Upsample(FL) and the input feature F are input into a feature subtraction module to obtain a high-frequency feature FH = F - Upsample(FL);

[0115] The high-frequency feature FH is input into a self-attention module to obtain a high-frequency feature FH1 (with a size of H*W*C);

[0116] The low-frequency feature FL and the high-frequency feature FH1 are input into a feature fusion module for feature fusion to obtain a first fused feature M (with a size of H*W*C);

[0117] The first fused feature M is input into a max pooling layer for max pooling operation to obtain a low-frequency feature ML = Maxpool(M) (with a size of H / 2*W / 2*C); where Maxpool represents the max pooling operation;

[0118] The low-frequency feature ML is input into a bilinear interpolation layer for bilinear interpolation operation to upsample the low-frequency feature ML to H*W*C, obtaining an upsampled feature Upsample(ML); where Upsample represents the bilinear interpolation operation;

[0119] The upsampled feature Upsample(ML) and the input feature M are input into a feature subtraction module to obtain a high-frequency feature MH = M - Upsample(ML));

[0120] The high-frequency feature MH is input into a self-attention module to obtain a high-frequency feature MH1 (with a size of H*W*C);

[0121] The low-frequency feature ML and the high-frequency feature MH1 are input into a feature fusion module for feature fusion to obtain a second fused feature K (with a size of H*W*C);

[0122] The second fused feature K is input into a detail control module to obtain an additional detail feature Z c 。

[0123] Furthermore, the self-attention module includes a convolution module, a first multiplication operation unit, a softmax layer, and a second multiplication operation unit;

[0124] The convolution module is used to perform three 1×1 convolution operations on the output of the feature subtraction module, and output Q value, K value, and V value;

[0125] The first multiplication operation unit is used to perform a multiplication operation on the Q value and the K value, and output the product Q*K;

[0126] The softmax layer is used to perform a softmax operation on the product Q*K, and output softmax(Q*K);

[0127] The second multiplication operation unit performs a multiplication operation on softmax(Q*K) and the V value, and outputs the product softmax(Q*K)*V, and the product softmax(Q*K)*V is the output of the self-attention module.

[0128] Specifically:

[0129] 1) For the self-attention module in the average pooling processing unit:

[0130] Input the high-frequency feature FH into the convolution module to perform three 1×1 convolution operations to obtain the Q value, the K value, and the V value;

[0131] Input the Q value and the K value into the first multiplication operation unit to perform a multiplication operation to obtain the product Q*K;

[0132] Input the product Q*K into the softmax layer to perform a softmax operation to obtain softmax(Q*K);

[0133] Input softmax(Q*K) and the V value into the second multiplication operation unit to perform a multiplication operation to obtain the product softmax(Q*K)*V;

[0134] The product softmax(Q*K)*V is the high-frequency feature FH1;

[0135] 2) For the self-attention module in the max pooling processing unit:

[0136] Input the high-frequency feature MH into the convolution module to perform one convolution to obtain the Q value, the K value, and the V value;

[0137] Input the Q value and the K value into the first multiplication operation unit to perform a multiplication operation to obtain the product Q*K;

[0138] Input the product Q*K into the softmax layer to perform a softmax operation to obtain softmax(Q*K);

[0139] Input the softmax(Q*K) and V values into the second multiplication operation unit for multiplication operation to obtain the product softmax(Q*K)*V;

[0140] The product softmax(Q*K)*V is the high-frequency feature MH1.

[0141] Furthermore, the feature fusion module includes a first MBB module, a second MBB module, a third MBB module, an upsampling module, a fourth MBB module, a second channel splicing module, a first 3*3 convolutional layer, a channel attention layer, and a third 1*1 convolutional layer;

[0142] The input end of the first MBB module is connected to the input end of the bilinear interpolation layer;

[0143] The output end of the first MBB module is sequentially connected to the input end of the second channel splicing module through the second MBB module, the third MBB module, and the upsampling module;

[0144] The input end of the fourth MBB module is connected to the output end of the self-attention module;

[0145] The output end of the fourth MBB module is sequentially connected to the input end of the third 1*1 convolutional layer through the first 3*3 convolutional layer and the channel attention layer;

[0146] The output end of the third 1*1 convolutional layer is the output end of the feature fusion module.

[0147] Specifically:

[0148] 1) For the feature fusion module in the average pooling processing unit:

[0149] Process the low-frequency feature FL through the first MBB module, the second MBB module, the third MBB module, and the upsampling module to obtain the feature TF (with a size of H*W*C);

[0150] Process the high-frequency feature FH1 through the fourth MBB module to obtain the feature RF (with a size of H*W*C);

[0151] Input the feature TF and the feature RF into the second channel splicing module for channel dimension splicing and then process them through the first 3*3 convolutional layer, the channel attention layer, and the third 1*1 convolutional layer to obtain the first fusion feature M;

[0152] 1) For the feature fusion module in the max pooling processing unit:

[0153] The low-frequency feature ML is processed through the first MBB module, the second MBB module, the third MBB module, and the upsampling module to obtain the feature TM (with dimensions H*W*C).

[0154] The high-frequency feature MH1 is processed through the fourth MBB module to obtain the feature RM (with dimensions H*W*C);

[0155] The feature TM and the feature RM are input into the second-channel splicing module for splicing in the channel dimension, and then processed through the first 3*3 convolutional layer, the channel attention layer, and the third 1*1 convolutional layer to obtain the second fusion feature K.

[0156] Furthermore, the structures of the first MBB module, the second MBB module, the third MBB module, and the fourth MBB module are the same, and each includes a second 3*3 convolutional layer, a third 3*3 convolutional layer, a fourth 3*3 convolutional layer, a fifth 3*3 convolutional layer, and an addition operation unit;

[0157] The input end of the second 3*3 convolutional layer is connected to the input end of the third 3*3 convolutional layer as the input end of the MBB module;

[0158] The output end of the third 3*3 convolutional layer is sequentially connected to the input end of the addition operation unit through the fourth 3*3 convolutional layer and the fifth 3*3 convolutional layer;

[0159] The output end of the second 3*3 convolutional layer is connected to the input end of the addition operation unit;

[0160] The output end of the addition operation unit is the output end of the MBB module.

[0161] Specifically:

[0162] 1) Taking the first MBB module in the average pooling processing unit as an example:

[0163] The low-frequency feature FL is processed through the second 3*3 convolutional layer to obtain the feature FL(1);

[0164] The low-frequency feature FL is processed through the third 3*3 convolutional layer, the fourth 3*3 convolutional layer, and the fifth 3*3 convolutional layer to obtain the feature FL(2);

[0165] The feature FL(2) and the feature FL(3) are input into the addition operation unit for addition operation to obtain the feature FL(3);

[0166] 2) Taking the first MBB module in the max pooling processing unit as an example:

[0167] The low-frequency feature ML is processed through the second 3*3 convolutional layer to obtain the feature ML(1);

[0168] The low-frequency feature ML is processed through the third 3×3 convolutional layer, the fourth 3×3 convolutional layer, and the fifth 3×3 convolutional layer to obtain the feature ML(2).

[0169] The feature ML(2) and the feature ML(3) are input into an addition operation unit for addition operation to obtain the feature ML(3).

[0170] The processing procedures of all other MBB modules are the same as that of the first MBB module, except for the input and output.

[0171] Further, the rotation feature extraction module includes a first branch, a second branch, a third branch, and an arithmetic mean operation unit.

[0172] The first branch includes a first Z-pool layer, a second 7×7 convolution, a first batch norm layer, a first sigmoid layer, and a third multiplication operation unit that are connected end to end in sequence from the input to the output.

[0173] The input end of the first Z-pool layer and the output end of the first sigmoid layer are both connected to the input end of the third multiplication operation unit.

[0174] The second branch includes an H-axis counterclockwise rotation unit, a second Z-pool layer, a third 7×7 convolution, a second batch norm layer, a second sigmoid layer, a fourth multiplication operation unit, and an H-axis clockwise rotation unit that are connected end to end in sequence from the input to the output. Among them, the H-axis counterclockwise rotation unit is used to rotate its input 90 degrees counterclockwise along the H axis; the H-axis clockwise rotation unit is used to rotate its input 90 degrees clockwise along the H axis.

[0175] The output end of the H-axis counterclockwise rotation unit and the output end of the second sigmoid layer are both connected to the input end of the fourth multiplication operation unit.

[0176] The third branch includes a W-axis counterclockwise rotation unit, a third Z-pool layer, a fourth 7×7 convolution, a third batch norm layer, a third sigmoid layer, a fifth multiplication operation unit, and a W-axis clockwise rotation unit that are connected end to end in sequence from the input to the output. Among them, the W-axis counterclockwise rotation unit is used to rotate its input 90 degrees counterclockwise along the W axis; the W-axis clockwise rotation unit is used to rotate its input 90 degrees clockwise along the W axis.

[0177] The output end of the W-axis counterclockwise rotation unit and the output end of the third sigmoid layer are both connected to the input end of the fifth multiplication operation unit.

[0178] The output terminals of the third multiplication operation unit, the H-axis clockwise rotation unit, and the W-axis clockwise rotation unit are all connected to the input terminal of the arithmetic mean operation unit;

[0179] The output terminal of the arithmetic mean operation unit is the output terminal of the rotation feature extraction module;

[0180] The input terminals of the first Z-pool layer, the H-axis counterclockwise rotation unit, and the W-axis counterclockwise rotation unit are the input terminals of the rotation feature extraction module.

[0181] Specifically:

[0182] 1) Input the additional detailed feature Z c (with a shape of C’*H’*W’) into the first Z-pool layer to obtain a reduced feature X1 (with a shape of 2*H’*W’); wherein, the first Z-pool layer is used for reducing the channel dimension and reducing the number of channels to 2;

[0183] Process the reduced feature X1 through a second 7*7 convolution, a first batch norm layer, and a first sigmoid layer to obtain a first-branch attention weight W1; wherein, the shape of the tensor output by the first batch norm layer is 1*H’*W’;

[0184] Input the first-branch attention weight W1 and the additional detailed feature Z c into the third multiplication operation unit for multiplication operation to obtain a first-branch additional detailed feature W1*Z c ;

[0185] 2) Input the additional detailed feature Z c (with a shape of C’*H’*W’) into the H-axis counterclockwise rotation unit, and rotate the additional detailed feature Z c counterclockwise by 90 degrees along the H axis to obtain a rotated tensor Z 1 c (with a shape of W’*H’*C’)

[0186] Input the rotated tensor Z 1 c into the second Z-pool layer to obtain a reduced feature X2 (with a shape of 2*H’*C’); wherein, the second Z-pool layer is used for reducing the width dimension and reducing the width to 2;

[0187] Process the reduced feature X2 through a third 7*7 convolution, a second batch norm layer, and a second sigmoid layer to obtain a second-branch attention weight W2; wherein, the shape of the tensor output by the second batch norm layer is 1*H’*C’;

[0188] Input the attention weight W2 of the second branch and the rotation tensor into the fourth multiplication operation unit for multiplication operation to obtain the initial additional detailed features of the second branch

[0189] Input the initial additional detailed features of the second branch W2* into the H-axis clockwise rotation unit to rotate the initial additional detailed features of the second branch clockwise by 90 degrees along the H axis to obtain the final additional detailed features of the second branch;

[0190] 3) Input the additional detailed features Z c (with the shape of C’*H’*W’) into the W-axis counterclockwise rotation unit to rotate the additional detailed features Z c counterclockwise by 90 degrees along the W axis to obtain the rotation tensor (with the shape of H’*C’*W’)

[0191] Input the rotation tensor into the third Z-pool layer to obtain the reduced feature X3 (with the size of 2*C’*W’); wherein, the third Z-pool layer is used for reducing the length dimension and reducing the length to 2;

[0192] Process the reduced feature X3 through the fourth 7*7 convolution, the third batch norm layer, and the third sigmoid layer to obtain the attention weight W3 of the third branch; wherein, the shape of the tensor output by the third batch norm layer is 1*C’*W’;

[0193] Input the attention weight W3 of the third branch and the rotation tensor into the fifth multiplication operation unit for multiplication operation to obtain the initial additional detailed features of the second branch

[0194] Input the initial additional detailed features of the third branch into the W-axis clockwise rotation unit to rotate the initial additional detailed features of the third branch clockwise by 90 degrees along the W axis to obtain the final additional detailed features of the third branch;

[0195] Input the additional detailed features W1*Z of the first branch c , the final additional detailed features of the third branch, and the final additional detailed features of the third branch into the arithmetic mean operation unit; calculate the arithmetic mean of the additional detailed features W1*Z c of the first branch, the final additional detailed features of the third branch, and the final additional detailed features of the third branch to obtain the rotation feature Z z .

[0196] Further, the trained super-resolution image generation model is obtained based on the following steps:

[0197] Obtain a training data set; the training data set includes a number of low-resolution images;

[0198] Input the data set into the super-resolution image generation model for training, and minimize the loss function to update the network parameters of the detail control network and the rotation feature extraction module, so as to obtain the trained detail control network and the trained rotation feature extraction module, that is, the trained super-resolution image generation model can be obtained; where the expression of the loss function is ∈ represents the Gaussian noise generated by the noise generation module; ∈ θ (x) represents the predicted noise output by the pre-trained Unet network; x represents the noise latent feature output by the fusion noise module after inputting the training image; N(0,1) represents the Gaussian noise that satisfies the standard normal distribution; E represents taking the average value.

[0199] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, refer to the description of the method part.

[0200] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A super-resolution image generation method, characterized in that: The following steps are involved: Inputting the image to be improved in resolution into the trained super-resolution image generation model to obtain an image with improved resolution; The trained super-resolution image generation model includes a pre-trained encoder, a noise generation module, a fusion noise module, a first channel splicing module, a trained detail control network, a trained rotation feature extraction module, a pre-trained Unet network and a pre-trained decoder; The output end of the pre-trained encoder and the output end of the noise generation module are both connected to the input end of the fusion noise module; The output end of the pre-trained encoder and the output end of the noise generation module are both connected to the input end of the first channel splicing module; The output end of the first channel splicing module is connected to the input end of the pre-trained Unet network through the trained detail control network and the trained rotation feature extraction module in sequence; The output end of the fusion noise module is connected to the input end of the pre-trained Unet network; The output end of the pre-trained Unet network is connected to the input end of the pre-trained decoder.

2. A super-resolution image generation method according to claim 1, characterized in that: The fusion noise module performs the following operations: α i =1-β i ; Among them, Z n represents the noise latent variable output by the fusion noise module; β i represents the variance of the Gaussian noise generated by the noise generation module; β i ∈(0,1); N represents the selected β i The number of obtained; ∈ represents the Gaussian noise generated by the noise generation module; Z represents the potential variable output by the pre-trained encoder.

3. The super-resolution image generation method according to claim 1, characterized in that: The trained detail control network includes a detail preservation module and a detail control module; An input end of the detail preservation module is connected to an output end of the first channel splicing module; The output end of the detail preservation module is connected to the input end of the trained rotation feature extraction module through the detail control module; The detail preservation module includes a first 1*1 convolutional layer, a first 7*7 convolutional layer, a second 1*1 convolutional layer, a convolutional feedforward neural network, an average pooling processing unit and a maximum pooling processing unit; The input end of the first 1*1 convolutional layer is connected to the output end of the first channel splicing module; The output end of the first 1*1 convolutional layer is connected to the input end of the maximum pooling processing unit through the first 7*7 convolutional layer, the second 1*1 convolutional layer, the convolutional feedforward neural network, and the average pooling processing unit in sequence; The output end of the maximum pooling processing unit is connected to the input end of the detail control module; The detail control module has the same downsampling structure as the Unet network.

4. The super-resolution image generation method according to claim 3, characterized in that: The average pooling processing unit includes an average pooling layer, a bilinear difference layer, a feature subtraction module, a self-attention module and a feature fusion module; The input end of the average pooling layer is connected to the output end of the convolutional feedforward neural network; The output end of the average pooling layer is connected to the input end of the bilinear difference layer; The input end of the average pooling layer and the output end of the bilinear difference layer are both connected to the input end of the feature subtraction module; The output end of the feature subtraction module is connected to the input end of the feature fusion module through the self-attention module; The output end of the average pooling layer is connected to the input end of the feature fusion module.

5. The super-resolution image generation method according to claim 4, characterized in that: The maximum pooling processing unit replaces the average pooling layer in the average pooling processing unit with a maximum pooling layer, and the rest of the structure is the same as that of the average pooling processing unit.

6. The super-resolution image generation method according to claim 5, characterized in that: The self-attention module includes a convolution module, a first multiplication unit, a softmax layer, and a second multiplication unit; The convolution module is used to perform three 1*1 convolution operations on the output of the feature subtraction module, and output a Q value, a K value, and a V value; The first multiplication unit is used to perform a multiplication operation on the Q value and the K value, and output a product Q*K; The softmax layer is used to perform a softmax operation on the product Q*K and output softmax(Q*K); The second multiplication unit performs a multiplication operation on softmax(Q*K) and the V value, and outputs the product softmax(Q*K)*V, where the product softmax(Q*K)*V is the output of the self-attention module.

7. The super-resolution image generation method according to claim 6, characterized in that: The feature fusion module includes a first MBB module, a second MBB module, a third MBB module, an upsampling module, a fourth MBB module, a second channel splicing module, a first 3*3 convolutional layer, a channel attention layer, and a third 1*1 convolutional layer; An input end of the first MBB module is connected to an input end of the bilinear difference layer; The output end of the first MBB module is connected to the input end of the second channel splicing module through the second MBB module, the third MBB module, and the upsampling module in sequence; An input end of the fourth MBB module is connected to an output end of the self-attention module; The output end of the fourth MBB module is connected to the input end of the third 1*1 convolutional layer through the first 3*3 convolutional layer and the channel attention layer in sequence; The output end of the third 1*1 convolutional layer is the output end of the feature fusion module.

8. The method for generating a super-resolution image according to claim 7, characterized in that: The first MBB module, the second MBB module, the third MBB module and the fourth MBB module have the same structure, and all include a second 3*3 convolutional layer, a third 3*3 convolutional layer, a fourth 3*3 convolutional layer, a fifth 3*3 convolutional layer and an addition operation unit; An input end of the second 3*3 convolutional layer is connected to an input end of the third 3*3 convolutional layer as an input end of the MBB module; The output end of the third 3*3 convolutional layer is connected to the input end of the addition operation unit through the fourth 3*3 convolutional layer and the fifth 3*3 convolutional layer in sequence; The output end of the second 3*3 convolutional layer is connected to the input end of the addition operation unit; The output end of the adding unit is the output end of the MBB module.

9. The super-resolution image generation method according to claim 8, characterized in that: The rotation feature extraction module includes a first branch, a second branch, a third branch and an arithmetic mean calculation unit; The first branch includes a first Z-pool layer, a second 7*7 convolution, a first batch norm layer, a first sigmoid layer and a third multiplication unit which are sequentially connected from input to output; The input end of the first Z-pool layer and the output end of the first sigmoid layer are both connected to the input end of the third multiplication operation unit; The second branch includes an H-axis counterclockwise rotation unit, a second Z-pool layer, a third 7*7 convolution, a second batch norm layer, a second sigmoid layer, a fourth multiplication unit and an H-axis clockwise rotation unit, which are connected in sequence from input to output; wherein the H-axis counterclockwise rotation unit is used to rotate its input 90 degrees counterclockwise along the H axis; and the H-axis clockwise rotation unit is used to rotate its input 90 degrees clockwise along the H axis; The output end of the H-axis counterclockwise rotation unit and the output end of the second sigmoid layer are both connected to the input end of the fourth multiplication unit; The third branch includes a W-axis counterclockwise rotation unit, a third Z-pool layer, a fourth 7*7 convolution, a third batch norm layer, a third sigmoid layer, a fifth multiplication unit and a W-axis clockwise rotation unit, which are connected in sequence from input to output; wherein the W-axis counterclockwise rotation unit is used to rotate its input 90 degrees counterclockwise along the W axis; and the W-axis clockwise rotation unit is used to rotate its input 90 degrees clockwise along the W axis; The output end of the W-axis counterclockwise rotation unit and the output end of the third sigmoid layer are both connected to the input end of the fifth multiplication unit; The output ends of the third multiplication unit, the H-axis clockwise rotation unit and the W-axis clockwise rotation unit are all connected to the input end of the arithmetic mean value calculation unit; The output end of the arithmetic mean value calculation unit is the output end of the rotation feature extraction module; The input ends of the first Z-pool layer, the H-axis counterclockwise rotation unit, and the W-axis counterclockwise rotation unit are input ends of the rotation feature extraction module.

10. The super-resolution image generation method according to claim 9, characterized in that: The trained super-resolution image generation model is obtained based on the following steps: Obtain a training data set; the training data set includes a plurality of low-resolution images; The data set is input into the super-resolution image generation model for training, and the network parameters of the detail control network and the rotation feature extraction module are updated by minimizing the loss function, so as to obtain the trained detail control network and the trained rotation feature extraction module, and thus the trained super-resolution image generation model can be obtained; wherein the expression of the loss function is: ∈ represents the Gaussian noise generated by the noise generation module; ∈ θ (x) represents the predicted noise output by the pre-trained Unet network; x represents the noise potential characteristics output by the fusion noise module after the training image is input; N(0,1) represents Gaussian noise that satisfies the standard normal distribution; E represents averaging.