Image Reconstruction Method Using Dynamic Convolution Kernel Correction of Pre-trained StyleGAN Model
The generator convolution kernel of the pre-trained StyleGAN model is corrected through multi-scale encoder and attention module, which solves the problem of the lack of effective correction of the dynamic convolution kernel, improves the quality of image reconstruction and simplifies the training process.
Patent Information
- Application Number
- CN202310060359.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-01-16
AI Technical Summary
In the existing image reconstruction method using pre-trained StyleGAN model, the dynamic convolution kernel in the generator lacks effective correction, resulting in difficulty in training and poor reconstruction image quality, and failure to effectively model long-distance image feature, and loss of global spatial information.
The multi-scale encoder and attention module are used to generate style codewords and residual codewords, correct the static convolution kernel in the generator of the pre-trained StyleGAN model to be a dynamic convolution kernel, and correct the residual codewords to train the encoding network to improve the quality of image reconstruction.
This significantly improves the quality of image reconstruction, reduces parameters and calculation amount, simplifies the training process, and achieves efficient image reconstruction effect.
Smart Images

Figure CN116071258B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and digital image processing, and in particular to an image reconstruction method using a pre-trained StyleGAN model to dynamically correct convolutional kernels. Background Art
[0002] At present, generative adversarial networks and transformers have achieved good results in the field of computer vision. As an image generation task in computer vision, image reconstruction has also seen much development. Image reconstruction requires encoding the input image into a codeword and then generating a reconstructed image through a generator. The training cost of generative adversarial networks is too high, increasing the difficulty of high-definition image reconstruction. Due to the powerful high-definition image generation ability of the StyleGAN model and the feature of introducing style codewords from side branches, many methods use the pre-trained StyleGAN model to achieve high-definition image reconstruction, improving image quality and efficiency.
[0003] In the image reconstruction method using a pre-trained StyleGAN model, an encoder needs to be trained to encode the input image into a style codeword, and then the style codeword is input into the generator of the pre-trained StyleGAN model from the side branch to obtain the reconstructed image. Mean square error loss, perceptual loss, and identity loss are established between the input image and the reconstructed image to complete the training. Reconstructing the image only using the style codeword may lose high-frequency details in the image. Some methods adjust the parameters in the generator by fine-tuning or training a convolutional neural network to supplement image details.
[0004] However, there are two problems in the existing image reconstruction work using a pre-trained StyleGAN model: 1) The dynamic convolutional kernels in the generator are not effectively corrected. Fine-tuning the generator or training a convolutional neural network has a large number of parameters, which may cause training difficulties; 2) Long-distance modeling of image features is not carried out, only local information is concerned, and global spatial information may be lost.
[0005] Therefore, it is very necessary to provide an efficient image reconstruction method using a pre-trained StyleGAN model. Summary of the Invention
[0006] The object of the present invention is to provide an image reconstruction method using a pre-trained StyleGAN model to correct dynamic convolution kernels, which adopts a method of correcting the dynamic convolution kernels in the generator using the correction codewords output by the attention module, so as to improve the quality of the reconstructed images. By combining a multi-scale encoder and an attention module, style codewords and residual codewords are obtained and fed into the generator of the pre-trained StyleGAN model. The static convolution kernels in the generator are modulated into dynamic convolution kernels using the style codewords, and the dynamic convolution kernels are corrected using the residual codewords. An encoding network is trained, the method is simple and efficient, the number of parameters and the amount of computation are very low, the problem due to the lack of effective correction of the dynamic convolution kernels in the generator is well solved, and the quality of the reconstructed images is significantly improved.
[0007] The object of the present invention is achieved as follows:
[0008] An image reconstruction method using a pre-trained StyleGAN model to correct dynamic convolution kernels, characterized in that a multi-scale encoder and an attention module are combined to obtain style codewords and residual codewords and fed into the generator of the pre-trained StyleGAN model. The static convolution kernels in the generator are modulated into dynamic convolution kernels using the style codewords, and the dynamic convolution kernels are corrected using the residual codewords. An encoding network is trained to achieve the improvement of the quality of the reconstructed images, specifically including the following steps:
[0009] Step 1: The input image is fed into the multi-scale encoder to obtain multi-scale features, an attention module is constructed, image information is extracted from the multi-scale features to obtain style codewords and residual codewords;
[0010] Step 2: The style codewords and residual codewords are fed into the generator of the pre-trained StyleGAN model. The static convolution kernels in the generator are modulated into dynamic convolution kernels using the style codewords, and the dynamic convolution kernels are corrected using the residual codewords;
[0011] Step 3: The multi-scale encoder and the attention module are combined to train an encoding network, and the corrected generator is used to achieve the reconstruction of the image.
[0012] The attention module is composed of a number of self-attention modules and cross-attention modules connected in series, and the attention module parameters for generating style codewords and residual codewords are not shared.
[0013] The specific content of Step 2 includes:
[0014] Step 2-1: The style codewords are mapped into modulation codewords s through a fully connected layer and multiplied pixel by pixel on the static convolution kernels to obtain dynamic convolution kernels, the dimension size of which is ××3×3, is the output channel number of the convolution kernel, and is the input channel number of the convolution kernel;
[0015] Step 2-2: Map the residual codeword through two independent fully-connected layers to obtain the column residual codeword P and the row residual codeword Q, whose dimensional sizes are × and × respectively, where is the feature dimension of the residual codeword. Multiply the column residual codeword P and the row residual codeword Q matrices to obtain a residual matrix, whose dimensional size is ×;
[0016] Step 2-3: Copy the residual matrix 3×3 times and add it to the dynamic convolution kernel to obtain the corrected dynamic convolution kernel whose dimensional size is ××3×3.
[0017] The combination method of the multi-scale encoder and the attention module is as follows: The cross-attention module in the attention module extracts information from the features in ascending order of scale, and each scale feature corresponds to an independent cross-attention module.
[0018] Compared with the prior art, the present invention has the characteristics of low parameter quantity, low calculation quantity, and significantly improved reconstructed image quality. The method is simple, efficient, and preferably solves the problem of poor reconstructed image quality caused by the lack of effective correction of the dynamic convolution kernel of the generator. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a schematic flow diagram of the present invention;
[0020] Figure 2 is a schematic diagram of the encoder output features;
[0021] Figure 3 is a schematic diagram of the self-attention module;
[0022] Figure 4 is a schematic diagram of the comparison of the reconstructed image quality between the present invention and the prior art. DETAILED DESCRIPTION OF THE INVENTION
[0023] In order to more clearly and clearly illustrate the technical means, technical improvements and beneficial effects of the present invention, the present invention will be described in detail below with reference to the accompanying drawings.
[0024] The image reconstruction method using the pre-trained StyleGAN model to correct the dynamic convolution kernel of the present invention includes the following specific steps:
[0025] Step 1: Send the input image into the multi-scale encoder to obtain multi-scale features, construct an attention module, extract image information from the multi-scale features, and obtain a style codeword and a residual codeword;
[0026] Step 2: Send the style codeword and the residual codeword into the generator of the pre-trained StyleGAN model, modulate the static convolution kernel in the generator with the style codeword to obtain a dynamic convolution kernel, and use the residual codeword to correct the dynamic convolution kernel;
[0027] Step 3: Combine the multi-scale encoder and the attention module to train an encoding network, and use the generator corrected in Step 2 to improve the quality of the reconstructed image.
[0028] The specific steps of Step 1 are as follows: Send the input image into the multi-scale encoder to obtain multi-scale features, construct two attention modules to extract image information from the multi-scale features, and obtain the style codeword and the residual codeword respectively. Both attention modules are composed of a number of self-attention modules and cross-attention modules connected in series, and the attention module parameters for generating the style codeword and the residual codeword are not shared.
[0029] Step 2 specifically includes the following steps:
[0030] Step 2-1: Map the style codeword through a fully connected layer to obtain the modulation codeword s, and multiply it pixel by pixel on the static convolution kernel to obtain the dynamic convolution kernel, whose dimension size is ××3×3, where is the output channel number of the convolution kernel and is the input channel number of the convolution kernel.
[0031] Step 2-2: Map the residual codeword through two independent fully connected layers to obtain the column residual codeword P and the row residual codeword Q, whose dimension sizes are × and × respectively, where is the feature dimension of the residual codeword. Multiply the column residual codeword P and the row residual codeword Q matrices to obtain the residual matrix, whose dimension size is ×;
[0032] Step 2-3: Copy the residual matrix 3×3 times and add it to the dynamic convolution kernel to obtain the corrected dynamic convolution kernel whose dimension size is ××3×3.
[0033] The specific content of Step 2-1 is as follows: Map the style codeword through a fully connected layer to obtain the modulation codeword s, whose dimension is 1××1×1. Take out the static convolution kernel, whose dimension is ××3×3, and obtain the dynamic convolution kernel from the following formula (a):
[0034] W d = s⊙W0 (a);
[0035] where: the dimension size of is ××3×3; ⊙ represents pixel-by-pixel multiplication.
[0036] The specific content of Step 2-2 is as follows: Map the residual codeword through two independent fully connected layers to obtain the column residual codeword P and the row residual codeword Q, whose dimension sizes are × and × respectively. Obtain the residual matrix from the following formula (b):
[0037]
[0038] where: the dimension of is ×; represents matrix multiplication.
[0039] The specific steps of step 2-3 are as follows: Copy the residual matrix obtained in step 2-2 three by three to obtain an extended residual matrix with dimensions ××3×3, add it to the dynamic convolution kernel obtained in step 2-1, and obtain the corrected dynamic convolution kernel from the following formula (c)
[0040]
[0041] The specific steps of step 3 are as follows: Use the corrected generator in step 2 to obtain a reconstructed image, establish a mean square error loss, a perceptual loss, and an identity loss between the input image and the reconstructed image to train an encoding network composed of a multi-scale encoder and an attention module, and realize the reconstruction of the image.
[0042] Embodiment 1
[0043] Refer to Figure 1 , the present invention combines a multi-scale encoder and an attention module, sends an input image into the multi-scale encoder, uses the attention module to obtain a style codeword and a residual codeword, and sends them into the generator of the pre-trained StyleGAN model. Modulate the static convolution kernel in the generator into a dynamic convolution kernel with the style codeword, and use the residual codeword to correct the dynamic convolution kernel, and train an encoding network composed of a multi-scale encoder and an attention module to improve the quality of the reconstructed image, specifically including the following steps:
[0044] S1: Send the input image into the multi-scale encoder to obtain multi-scale features, construct an attention module, extract image information from the multi-scale features, and obtain a style codeword and a residual codeword;
[0045] Refer to Figure 2 , the specific steps of this step are as follows:
[0046] Step 0: Use a convolutional neural network to construct a multi-scale encoder, send an RGB input image with a height of H and a width of W into the multi-scale encoder, and the multi-scale encoder encodes the input image into three feature maps with different scales,,.
[0047] Step 1: Use a neural network to construct two attention modules, which are respectively used to generate a style codeword and a correction codeword. The attention modules are both composed of a number of self-attention modules and cross-attention modules connected in series. The two attention modules extract image information from the multi-scale features and obtain a style codeword and a correction codeword respectively.
[0048] S2: Send the style codeword and the residual codeword into the generator of the pre-trained StyleGAN model, modulate the static convolution kernel in the generator into a dynamic convolution kernel with the style codeword, and use the residual codeword to correct the dynamic convolution kernel;
[0049] Refer toFigure 3 , this step is specifically as follows:
[0050] Step 0. Map the style codeword through a fully connected layer to obtain the modulation codeword s, multiply it pixel by pixel on the static convolution kernel to obtain the dynamic convolution kernel;
[0051] Step 1. Map the residual codeword through two independent fully connected layers to obtain the column residual codeword P and the row residual codeword Q, multiply the column residual codeword P and the row residual codeword Q matrices to obtain the residual matrix;
[0052] Step 2. Copy the residual matrix 3×3 times and add it to the dynamic convolution kernel to obtain the corrected dynamic convolution kernel Perform a convolution operation on the features in the generator using this convolution kernel.
[0053] S3: Combine the multi-scale encoder and the attention module, train an encoding network, and use the corrected generator to improve the quality of the reconstructed image.
[0054] The present invention uses the multi-scale encoder and the attention module to obtain the style codeword and the residual codeword and send them into the generator of the pre-trained StyleGAN model, and uses the residual codeword to correct the dynamic convolution kernel, thereby improving the quality of the reconstructed image.
[0055] Refer to Figure 4 , for the comparison schematic diagram of the reconstructed image quality between the present invention and Method A: Richardson E, Alaluf Y, Patashnik O, et al. Encoding in style: a stylegan encoder for image-to-image translation[C]. Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2021: 2287-2296 and Method B: Alaluf Y, Patashnik O, Cohen-Or D. Restyle: A residual-based stylegan encoder via iterative refinement[C]. Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021: 6711-6720. It can be seen from the figure that the present invention has improved the quality of the reconstructed image.
[0056] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An image reconstruction method using dynamic convolution kernel correction of a pre-trained StyleGAN model, characterized in that, The method includes the following specific steps: Step 1: Feed the input image into a multi-scale encoder to obtain multi-scale features, construct an attention module, extract image information from the multi-scale features to obtain a style codeword and a residual codeword; Step 2: Feed the style codeword and the residual codeword into the generator of the pre-trained StyleGAN model, modulate the static convolutional kernels in the generator into dynamic convolutional kernels with the style codeword, and correct the dynamic convolutional kernels with the residual codeword; Step 3: Combine the multi-scale encoder and the attention module, train an encoding network, and use the corrected generator to achieve image reconstruction; where: Step 2 specifically includes: Step 2-1: Map the style codewords through a fully connected layer to obtain modulation codewords s, and multiply them pixel by pixel on the static convolution kernel to obtain the dynamic convolution kernel W d , whose dimension size is C out ×C in ×3×3, C out is the output channel number of the convolution kernel, and C in is the input channel number of the convolution kernel; Step 2-2: Map the residual codeword through two independent fully-connected layers to obtain a column residual codeword P and a row residual codeword Q, with their dimension sizes being C out ×L and L×C in , where L is the feature dimension of the residual codeword. Multiply the column residual codeword P and the row residual codeword Q matrices to obtain a residual matrix res, with its dimension size being C out ×C in ; Step 2-3: Copy the residual matrix res three times in 3x3 and add it to the dynamic convolution kernel W d to obtain the corrected dynamic convolution kernel whose dimension size is C out × C in × 3 × 3.
2. The image reconstruction method according to claim 1, wherein The attention module is composed of a number of self-attention modules and cross-attention modules connected in series, and the attention module parameters for generating the style codeword and the residual codeword are not shared.
3. The image reconstruction method according to claim 1, characterized in that The combination of the multi-scale encoder and the attention module is specifically as follows: The cross-attention module in the attention module extracts information from the features in ascending order of scale, and each scale feature corresponds to an independent cross-attention module.