A Blind Restoration Method for Real Degraded Images Based on Cross-Attention Mechanism

By introducing a cross-attention mechanism in image blind repair, optimizing the correlation between potential coding and feature maps, the problem of insufficient faithfulness and details in blind repair of real degraded images is solved, and higher quality image repair is achieved.

CN115829876BActive Publication Date: 2025-05-30NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211616971.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-15
Publication Date
2025-05-30
Estimated Expiration
2042-12-15

AI Technical Summary

Technical Problem

When the prior art blind repair of real degraded images without supervision, the reconstructed images are not faithful, the texture details are not rich enough, and the model generalization ability is weak.

Method used

The blind repair method of real degraded images based on the cross attention mechanism is adopted to optimize the latent coding through multi-head self-attention and multi-head cross attention, enhance the semantic feature correlation between the spatial features of the feature map and the potential coding, and improve the semantic feature expression ability.

Benefits of technology

It significantly improves the faithfulness and texture detail richness of the reconstructed images, and enhances the generalization ability of the model in various practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115829876B_ABST
    Figure CN115829876B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image processing. Specifically, it is a blind restoration method for real degraded images based on a cross-attention mechanism. By introducing an attention mechanism, multi-head self-attention optimization is performed on the latent encoding, realizing the semantic feature weight assignment of the optimal latent encoding. Multi-head cross-attention optimization is used for both the latent encoding and the multi-resolution scale feature maps, realizing the introduction of the spatial features of the multi-scale feature maps into the latent encoding, enhancing the correlation between the spatial features of the feature maps and the semantic features of the latent encoding, significantly improving the expression ability of the latent encoding, and solving the key problems of low fidelity of the reconstructed images and insufficient richness of texture details.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and specifically, it is a blind restoration method for real degraded images based on a cross-attention mechanism. Background Art

[0002] With the progress of the times and technology, image processing technology has been widely applied in various fields of modern society. As one of the major fields, image restoration has a wide range of applications. In each process of image generation, transmission, and storage, due to the limitations of the imaging system and digital imaging devices themselves and the vulnerability of the imaging process to various external environmental interferences, information in the image is lost, resulting in a degraded image. For example, relative motion between the camera and the scene causes motion blur; out-of-focus leads to defocus blur; Gaussian blur caused by solar radiation and atmospheric turbulence; continuous noise interference in the imaging system; various compression distortions and other image degradation methods. Therefore, how to blindly restore real degraded images without supervision has always been a hot research point in image processing.

[0003] Blind restoration of an image refers to an image restoration method that estimates the point spread function and the high-definition original image only by using the original degraded and blurred image. Traditional linear image restoration algorithms need to specifically design a corresponding inverse degradation function under the condition of clear image degradation methods to restore the degraded image. In the face of complex degradation and unknown types, the efficiency and practicality of traditional algorithms are poor. Currently, the main methods for blind restoration of degraded images are: the scheme based on encoder optimization, the scheme based on latent coding optimization, and the scheme based on latent space embedding. In the scheme based on encoder optimization, the generative adversarial network (GAN) and the encoder are jointly trained to let the encoder learn how to map the image to the latent space of the GAN. However, there is an overfitting problem with the encoder, resulting in a large structural difference between the reconstructed image and the input image. Especially for real-world images, the generalization ability of the model is very weak, and joint training leads to a huge number of network parameters. In the scheme based on latent coding optimization, the optimal latent coding corresponding to the real image in the latent space is iteratively optimized by the gradient descent method to minimize the per-pixel loss between the input and the reconstructed image. However, multiple iterations of optimization are required for each input image, consuming a huge amount of resources and having extremely low efficiency. The scheme based on latent space embedding is the current optimal solution. It can quickly realize the latent coding mapping by using the encoder and can also iteratively obtain a better latent coding. By embedding the optimized latent coding during the GAN generation process, the quality and efficiency of the reconstructed image are greatly improved. However, the texture of the reconstructed image is prone to excessive smoothing, lacking high-frequency details and having local artificial artifacts, resulting in insufficient faithfulness of the reconstructed image.

[0004] In addition, for the latent encoding generated by iterative optimization using an encoder or gradient descent method, the semantic features in the latent encoding are still highly coupled, and the expressive ability of the semantic feature information is insufficient. As a result, the overall structure of the generated reconstructed image is unnatural, artificial artifacts are likely to occur in local areas, and the texture is prone to excessive smoothing, lacking high-frequency detailed feature information, resulting in low fidelity of the reconstructed image and insufficient richness of texture details. It is usually trained and used under supervised or semi-supervised conditions, with the training set being high-quality clear images. In actual application scenarios, the blind restoration effect on real degraded blurred images is very poor, and blind restoration cannot be performed without supervision. Summary of the Invention

[0005] To solve the above problems, the present invention discloses a blind restoration method for real degraded images based on a cross-attention mechanism. By introducing an attention mechanism, multi-head self-attention optimization is performed on the latent encoding, realizing the semantic feature weight allocation of the optimal latent encoding. Multi-head cross-attention optimization is used for both the latent encoding and the multi-resolution scale feature maps, realizing the introduction of the spatial features of the multi-scale feature maps into the latent encoding, enhancing the correlation between the spatial features of the feature maps and the semantic features of the latent encoding, significantly improving the semantic feature expression ability of the latent encoding, and solving the key problems of low fidelity of the reconstructed image and insufficient richness of texture details.

[0006] The specific technical solution adopted by the present invention is as follows:

[0007] A blind restoration method for real degraded images based on a cross-attention mechanism, comprising the following steps:

[0008] Step 1: Obtain a highly degraded image dataset for training;

[0009] Step 2: Preprocess the training dataset in Step 1, perform scale scaling, and generate labels for the images;

[0010] Step 3: Use the encoder in U-Net to perform latent encoding mapping on the input image to obtain a preliminary latent encoding, which is consistent with the dimension of the W+ latent encoding;

[0011] Step 4: Use the decoder in U-Net to generate multi-resolution scale feature maps;

[0012] Step 5: Use the attention mechanism to optimize the latent codes and multi-scale feature maps generated in Steps 3 and 4. Use the multi-head self-attention mechanism to optimize the latent codes, optimizing the encoder's selection of semantic features in the latent codes. Take the feature map as the information source for query matching and the latent code as the query flag, and use the multi-head cross-attention mechanism to introduce the spatial features in the feature map into the latent code, enhancing the local details and global context consistency of the feature map, and completing the optimization of the latent code to improve its semantic expression ability;

[0013] Step 6: Use the latent codes optimized in Step 5 as the input and feed them into the pre-trained StyleGAN2 generator. Embed the multi-scale feature maps in Step 4 into the corresponding generation layers in the StyleGAN2 generation process to achieve the embedding and expansion of the latent space of the pre-trained generator, and then obtain the reconstructed image;

[0014] Step 7: Use multiple loss functions such as perceptual loss, pixel-level loss, adversarial loss, and frequency-domain loss to calculate the loss values between the GT of the input image and the reconstructed image, perform backpropagation processing on the network, and perform iterative optimization of the network hyperparameters to finally obtain the trained model;

[0015] Step 8: Based on the trained model, perform blind restoration and reconstruction on the real degraded blurred image. Feed the blurred image into the model trained in Step 7 for blind restoration to obtain a high-quality and highly faithful reconstructed image.

[0016] For a further improvement of the present invention, the blurred dataset for training in Step 1 is generated by mixing and combining different degradation methods such as using different types of blur kernels, downsampling blur, JPEG compression distortion, and adding noise. The degradation formula is as follows:

[0017]

[0018] where is the highly degraded blurred image generated, is the high-quality image, is the convolution operation, is the blur kernel (Gaussian blur kernel or anisotropic blur kernel), r is the downsampling ratio factor, is the additive Gaussian noise, and JPEG q is the JPEG compression for determining the quality factor q.

[0019] For a further improvement of the present invention, in steps 3 and 4, each encoding and decoding block layer in the encoder and decoder of U-Net (where the encoding block layer is a downsampling operation and the decoding block layer is an upsampling operation) has a residual connection structure. The backbone is composed of a combination of convolutional layers with a kernel size of 3*3 and 1*1, and the branch is a convolutional layer with a kernel size of 3*3. The finally generated latent encoding dimension is 16*512.

[0020] For a further improvement of the present invention, the multi-resolution scale feature maps generated in step 4 are all subjected to scale and translation processing. Among them, the convolutional layer for scale processing has a kernel size of 3*3, and the convolutional layer for translation processing has a kernel size of 1*1.

[0021] For a further improvement of the present invention, in step 5, the preliminary 16*512-dimensional latent encoding and the 8*8*256-512*512*16 multi-scale feature maps generated in steps 3 and 4 are optimized using the attention mechanism. Among them, the latent encoding is optimized using the multi-head self-attention mechanism, taking the feature map as the information source for query matching and the latent encoding as the query flag. The multi-head cross-attention mechanism is used to optimize between the latent encoding and the multi-scale feature map, introducing the spatial features in the multi-scale feature map into the latent encoding to enhance the local details and global context consistency of the feature map, and completing the optimization of the latent encoding to improve its semantic expression ability.

[0022] The multi-head cross-attention formula is similar to the multi-head self-attention formula. The difference is that in the multi-head self-attention, the latent encoding is used to generate Q, K, and V, while in the multi-head cross-attention, the multi-scale feature map is used to generate K and V, and the latent encoding is used to generate Q. The multi-head self-attention formula is as follows:

[0023]

[0024]

[0025] MHA(Q, K, V) = [Attention(Q, K, V)] h=1:H W O

[0026] The above is the formula of the multi-head self-attention mechanism, where Q is the query matrix, K is the keys matrix, V is the values matrix, q is the 512-dimensional query tokens, is the set of query tokens, and are both and are both learnable mapping matrices in the self-attention module. H is the number of attention heads, d is the feature dimension and is equal to 512 / H, is also a learnable mapping matrix for performing the fusion operation of the final result.

[0027] For a further improvement of the present invention, the input for pre-training StyleGAN2 in step 6 is the optimized 16 * 512-dimensional latent code in step 5, and the 8 * 8 * 256 - 512 * 512 * 16 multi-scale feature maps embedded during the generation process of StyleGAN2 are the multi-scale feature maps that have undergone scale and translation processing in step 4.

[0028] For a further improvement of the present invention, in step 7, the GT of the input image and the reconstructed image are combined to calculate the loss using a loss function, where the combination consists of a perceptual loss based on VGG-19, a per-pixel loss of MSE, an adversarial loss, and a frequency-domain loss of FFT. The loss function is defined as follows:

[0029]

[0030] The above is the perceptual loss function, where is the reconstructed image, I ∈ R H*W*C is the reference GT image, H represents the height of the image, W represents the width of the image, C represents the three RGB channels. In the present invention, I, Φ is the pre-trained VGG-19 network. In the experiment, the outputs of a total of 7 layers, namely conv1_2, conv2_2, conv3_2 to conv7_2, which have not passed through the LeakyReLU activation function, are selected. is the L1 norm operation on the output of the VGG-19 network, where L mse The root mean square loss function is defined as follows:

[0031]

[0032] The above is the root mean square loss function, where G represents the pre-trained StyleGAN2 generator, W represents the 16 * 512-dimensional latent code, N is the scalar in the image, that is, N = H * W * C. Where L adv The adversarial loss function is defined as follows:

[0033]

[0034] The above is the adversarial loss function, represents the formula abbreviation for encoding and mapping the reconstructed image. D is the discriminator of StyleGAN2, and softplus is the smooth approximation of the ReLU activation function, which is used to ensure that the output is always positive. Where L fft The frequency-domain loss function is defined as follows:

[0035]

[0036] The above is the frequency-domain loss function, where, is the feature map generated in U-Net, i is the i-th layer in the multi-resolution scale feature map, and t i is the total number of layers of the generated feature maps accumulated, is the fast Fourier transform operation. The total loss function combination and the proportion of each loss weight are as follows:

[0037] L total = λ per L per + λ mse L mse + λ adv L adv + λ fft L fft

[0038] The above is the total loss function. The λ before each item above * is the corresponding loss function ratio coefficient, which are 10:2:2:1 respectively. Among them, λ per L per is the perceptual loss function based on the VGG-19 network, and λ mse L mse is the root mean square loss function, and λ adv L adv is the adversarial loss function, and λ fft L fft is the high-frequency loss function of FFT.

[0039] Advantages of the present invention: By randomly combining multiple degradation methods to generate a highly degraded blurred image training set, the present invention realistically simulates the complex degradation situation of real-world images, improves the generalization ability of the model in various practical applications, and realizes the blind restoration task of real degraded images under unsupervised conditions; the present invention introduces the FFT loss function in the frequency domain to enhance the model's attention to high-frequency feature information, making the texture and local details of the reconstructed image richer. Traditional loss functions usually choose the MSE loss function, perceptual loss function, and regularization loss function, resulting in the model paying more attention to low-frequency feature information, thus causing the texture of the result to be overly smooth. Brief Description of the Drawings

[0040] Figure 1 is a schematic diagram of the overall framework of the model of the present invention.

[0041] Figure 2 is a schematic diagram of the Transformer block in the present invention.

[0042] Figure 3 is a schematic diagram of embedding the multi-scale feature map into the intermediate layer of the StyleGAN2 generation process in the present invention.

[0043] Figure 4It is a comparison chart of the experimental results of the present invention. Detailed implementation manners

[0044] To deepen the understanding of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. The embodiments are only used to explain the present invention and do not limit the protection scope of the present invention.

[0045] A blind restoration method for real degraded images based on cross-attention mechanism includes the following steps:

[0046] Step 1: Obtain a highly degraded image dataset for training:

[0047] It is generated by mixing and combining different degradation methods such as using different types of blur kernels, downsampling blur, JPEG compression distortion, and adding noise. The degradation formula is as follows:

[0048]

[0049] Where is the generated highly degraded blurred image, is the high-quality image, is the convolution operation, is the blur kernel (Gaussian blur kernel or anisotropic blur kernel), r is the downsampling ratio factor, is the additive Gaussian noise, JPEG q is the JPEG compression for determining the quality factor q.

[0050] Step 2: Preprocess the training dataset in Step 1, perform scale scaling, and generate labels for the images.

[0051] Step 3: Use the encoder in U-Net to perform potential coding mapping on the input image to obtain a preliminary potential coding, which has the same dimension as the W+ potential coding;

[0052] Step 4: Use the decoder in U-Net to generate feature maps of multi-resolution scales;

[0053] In the encoder and decoder of U-Net in the above Steps 3 and 4, each encoding and decoding block layer has a residual connection structure. The backbone is composed of a combination of convolutional layers with a convolutional kernel size of 3*3 and 1*1, and the branch is a convolutional layer with a convolutional kernel size of 3*3. The finally generated potential coding dimension is 16*512; the feature maps of multi-resolution scales generated in Step 4 have all undergone scale and translation processing. Among them, the convolutional layer for scale processing has a convolutional kernel size of 3*3, and the convolutional layer for translation processing has a convolutional kernel size of 1*1.

[0054] Step 5: Optimize the preliminary 16 * 512-dimensional latent code and the 8 * 8 * 256 - 512 * 512 * 16 multi-scale feature maps generated in Step 3 and Step 4 using the attention mechanism. Among them, use the multi-head self-attention mechanism to optimize the latent code, take the feature map as the information source for query matching and the latent code as the query flag, and use the multi-head cross-attention mechanism to optimize between the latent code and the multi-scale feature map, introduce the spatial features in the multi-scale feature map into the latent code, enhance the local details and global context consistency of the feature map, and complete the optimization of the latent code to improve its semantic expression ability;

[0055] Among them, in the above multi-head self-attention, use the latent code to generate Q, K, and V, while in the multi-head cross-attention, use the multi-scale feature map to generate K and V, and use the latent code to generate Q. The formula for multi-head self-attention is as follows:

[0056]

[0057]

[0058] MHA(Q, K, V) = [Attention(Q, K, V)] h=1:H W O

[0059] In the above formula, Q is the query matrix, K is the keys matrix, V is the values matrix, q is the 512-dimensional query tokens, is the set of query tokens, and both and are learnable mapping matrices in the self-attention module. H is the number of attention heads, d is the feature dimension and is equal to 512 / H, is also a learnable mapping matrix for performing the fusion operation of the final result.

[0060] Step 6: Use the latent code optimized in Step 5 as the input and send it into the pre-trained StyleGAN2 generator, and embed the multi-scale feature map in Step 4 into the corresponding generation layer in the StyleGAN2 generation process to achieve the embedding expansion of the latent space of the pre-trained generator, and then obtain the reconstructed image: The input of the pre-trained StyleGAN2 is the 16 * 512-dimensional latent code optimized in Step 5, and the 8 * 8 * 256 - 512 * 512 * 16 multi-scale feature map embedded in the StyleGAN2 generation process is the multi-scale feature map processed by scale and translation in Step 4.

[0061] Step 7: Use multiple loss functions such as perceptual loss, pixel-level loss, adversarial loss, and frequency-domain loss to calculate the loss values of the GT of the input image and the reconstructed image, perform backpropagation on the network, and perform iterative optimization of the network hyperparameters to finally obtain a trained model; calculate the combined loss of the loss functions for the GT of the input image and the reconstructed image, where the combination consists of perceptual loss based on VGG-19, per-pixel loss of MSE, adversarial loss, and frequency-domain loss of FFT. The definitions of each part of the loss function are as follows:

[0062]

[0063] In the above perceptual loss function, is the reconstructed image, I ∈ R H*W*C is the reference GT image, H represents the height of the image, W represents the width of the image, C represents the three RGB channels. In the present invention, I, Φ is the pre-trained VGG-19 network. In the experiment, the outputs of 7 layers, namely conv1_2, conv2_2, conv3_2 to conv7_2, which have not passed through the LeakyReLU activation function, are selected. Perform L1 norm operation on the output of the VGG-19 network.

[0064]

[0065] In the above root mean square loss function, G represents the pre-trained StyleGAN2 generator, W represents the 16*512-dimensional latent code, and N is the scalar in the image, that is, N = H*W*C.

[0066]

[0067] In the above adversarial loss function, represents the formula abbreviation for encoding and mapping the reconstructed image. D is the discriminator of StyleGAN2, and softplus is the smooth approximation method of the ReLU activation function, which is used to ensure that the output is always positive.

[0068]

[0069] In the above frequency-domain loss function, is the feature map generated in U-Net. i is the i-th layer in the multi-resolution scale feature map, and t i is the total number of layers of the generated feature maps accumulated, is the fast Fourier transform operation.

[0070] L total = λ per L per + λmse L mse + λ adv L adv + λ fft L fft

[0071] In the above total loss function, λ before each term * is the corresponding loss function proportionality coefficient, which are 10:2:2:1 respectively. Among them, λ per L per is the perceptual loss function based on the VGG-19 network, and λ mse L mse is the root mean square loss function, and λ adv L adv is the adversarial loss function, and λ fft L fft is the high-frequency loss function of the FFT.

[0072] Step 8: Perform blind restoration and reconstruction on the real degraded blurred image based on the trained model, and send the blurred image into the model trained in Step 7 for blind restoration to obtain a high-quality and highly faithful reconstructed image.

[0073] As Figure 4 shown, send the blurred image highly degraded in the real world into the model trained in Step 7. The generated reconstructed and restored image has a more natural face structure and richer local texture details, and the faithfulness is very high. As Figure 4 shown, the experimental results will be compared with the best GFPGAN model in the current blind restoration field. Among them, in the first column, the ear area of the baby, in the second column, the eye pupil area of the woman, in the third column, the mole on the boy's arm and face, and in the fourth column, the eye corner wrinkles and mouth shape of the man. The reconstructed images of the present invention are of better generation quality than GFPGAN in the above high-frequency detail areas, and in areas such as the double eyelids on the face and the texture on the lips in each column of images, the present invention has richer details than GFPGAN. It is proved that the blind restoration and reconstruction images of the present invention have richer texture details and the overall structure is natural, and the input image and the reconstructed image have higher faithfulness.

[0074] In the above embodiment, Figure 1 similar residual connection operations are used inside the encoder-decoder or the same task is completed using the Transformer encoding block; Figure 2 in the attention mechanism, the replacement types are selected, for example, the multi-head cross-attention is replaced with cross-attention, the multi-head self-attention is replaced with self-attention or channel attention, etc., but the operation purposes are the same; Figure 3In the multi-scale feature map embedding, a channel attention mechanism or other operations are added to the channel segmentation operation to achieve the best channel segmentation purpose, which is also to perform proportional segmentation on the channels.

[0075] The above are exemplary embodiments of the present invention, and thus do not limit the protection scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the protection scope of the present invention.

Claims

1. A blind restoration method for real degraded images based on cross-attention mechanism, characterized in that, it includes the following steps: Step 1: Obtain a highly degraded image dataset for training; Step 2: Preprocess the training dataset in Step 1, perform scale scaling, and generate labels for the images; Step 3: Use the encoder in U-Net to perform potential encoding mapping on the input image to obtain preliminary potential encoding, which is consistent with the dimension of W+ potential encoding; Step 4: Use the decoder in U-Net to generate feature maps of multi-resolution scales; Step 5: Use the attention mechanism to optimize the potential encoding and multi-scale feature maps generated in Steps 3 and 4. Use the multi-head self-attention mechanism to optimize the potential encoding, optimize the selection of semantic features in the potential encoding by the encoder, use the feature map as the information source for query matching and the potential encoding as the query flag, and use the multi-head cross-attention mechanism to introduce the spatial features in the feature map into the potential encoding, enhance the local details and global context consistency of the feature map, and complete the optimization of the potential encoding to improve its semantic expression ability; Step 6: Take the latent code optimized in Step 5 as the input and send it into the pre-trained StyleGAN2 generator, and embed the multi-scale feature maps in Step 4 into the corresponding generation layers in the StyleGAN2 generation process to realize the embedding expansion of the potential space of the pre-trained generator, and then obtain the reconstructed image; Step 7: Use the loss function to calculate the loss value between the GT of the input image and the reconstructed image, perform backpropagation processing on the network, perform iterative optimization of the network hyperparameters, and finally obtain the trained model; Step 8: Based on the trained model, perform blind restoration and reconstruction on the real degraded blurred image, and send the blurred image into the model trained in Step 7 for blind restoration to obtain a high-quality and highly faithful reconstructed image.

2. The blind restoration method for real degraded images based on cross-attention mechanism according to claim 1, characterized in that, the blurred dataset for training in Step 1 is generated by mixing and combining different types of degradation methods, and the degradation formula is as follows: where is the generated highly degraded blurred image, is the high-quality image, is the convolution operation, is the blur kernel, r is the downsampling ratio factor, n δ is the additive Gaussian noise, JPEG q is the JPEG compression for determining the quality factor q.

3. The blind restoration method for real degraded images based on cross-attention mechanism according to claim 1, characterized in that, in the encoder and decoder of U-Net in Steps 3 and 4, each encoding and decoding block layer has a residual connection structure, where the backbone is composed of convolutional layers with a convolution kernel size of 3*3 and 1*1, and the branch is a convolutional layer with a convolution kernel size of 3*3. The finally generated potential encoding dimension is 16*512.

4. The blind restoration method for real degraded images based on cross-attention mechanism according to claim 3, characterized in that, the multi-resolution scale feature maps generated in Step 4 have all undergone scale and translation processing. Among them, the convolution kernel size in the convolutional layer for scale processing is 3*3, and the convolution kernel size in the convolutional layer for translation processing is 1*1.

5. The blind restoration method for real degraded images based on cross-attention mechanism according to claim 4, characterized in that, In step 5, the preliminary 16×512-dimensional latent code and the 8×8×256 - 512×512×16 multi-scale feature maps generated in steps 3 and 4 are optimized using the attention mechanism. Among them, the multi-head self-attention mechanism is used to optimize the latent code, taking the feature map as the information source for query matching and the latent code as the query flag, and the multi-head cross-attention mechanism is used to optimize between the latent code and the multi-scale feature map, introducing the spatial features in the multi-scale feature map into the latent code to enhance the local details and global context consistency of the feature map, and completing the optimization of the latent code to improve its semantic expression ability.

6. The real degraded image blind restoration method based on the cross-attention mechanism according to claim 5, wherein, in step 5, Q, K, and V are generated using the latent code in the multi-head self-attention, and the formula is as follows: MHA(Q, K, V) = [Attention(Q, K, V)] h=1:H W O Among them, Q is the query matrix, K is the keys matrix, V is the values matrix, q is the 512-dimensional query tokens, is the set of query tokens, and both and are learnable mapping matrices in the self-attention module. H is the number of attention heads, d is the feature dimension and is equal to 512 / H, is also a learnable mapping matrix for performing the fusion operation of the final result.

7. The real degraded image blind restoration method based on the cross-attention mechanism according to claim 6, wherein, the input of the pre-trained StyleGAN2 in step 6 is the 16×512-dimensional latent code optimized in step 5, and the 8×8×256 - 512×512×16 multi-scale feature map embedded during the generation of StyleGAN2 is the multi-scale feature map processed by scale and translation in step 4.

8. The real degraded image blind restoration method based on the cross-attention mechanism according to claim 7, wherein, In step 7, the GT of the input image and the reconstructed image are combined by a loss function to calculate the loss. The combination consists of a perceptual loss based on VGG-19, a pixel-wise loss of MSE, an adversarial loss, and a frequency domain loss of FFT, where L per The perceptual loss function is defined as follows: Among them, is the reconstructed image, I ∈ R H*W*C is the reference GT image, H represents the height of the image, W represents the width of the image, C represents the three RGB channels, I, Φ is the pre-trained VGG-19 network. In the experiment, the outputs of 7 layers, namely conv1_2, conv2_2, conv3_2 to conv7_2, which have not passed through the LeakyReLU activation function, are selected. is to perform the L1 norm operation on the output of the VGG-19 network, where L mse The root mean square loss function is defined as follows: Among them, G represents the pre-trained StyleGAN2 generator, W represents the 16 * 512-dimensional latent code, N is the scalar in the image, that is, N = H * W * C, where L adv The adversarial loss function is defined as follows: Among them, represents the formula abbreviation for encoding and mapping the reconstructed image. D is the discriminator of StyleGAN2, and softplus is the smooth approximation method of the ReLU activation function, which is used to limit the output to always be positive. Among them, L fft The frequency domain loss function is defined as follows: Among them, is the feature map generated in U-Net, i is the i-th layer in the multi-resolution scale feature map, and t i is the total number of cumulative layers of the generated feature map, is the fast Fourier transform operation. The total loss function combination and the proportion of each loss weight are as follows: L total = λ per L per + λ mse L mse + λ adv L adv + λ fft L fft λ before each of the above items * is the corresponding loss function proportionality coefficient, λ per = 0.1, λ mse = 0.02, λ adv = 0.02, λ fft = 0.01, where λ per L per is the perceptual loss function based on the VGG-19 network, λ mse L mse is the root mean square loss function, λ adv L adv is the adversarial loss function, λ fft L fft is the high-frequency loss function of the FFT.

Citation Information

Patent Citations

  • Attention mechanism-based image blind deblurring method and system

    CN111709895A

  • Face image blind restoration method and system

    CN113763268A