Face restoration method based on diffusion model and visual style prompt learning
By applying a face recovery method based on diffusion model and visual style prompt learning in the field of face reconstruction, the problem of accurate prediction of face prior representation and characteristics in low-quality pictures is solved, and efficient face reconstruction is achieved and reconstruction quality is improved.
Patent Information
- Application Number
- CN202411923775.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-27
AI Technical Summary
Given low-quality pictures, how to accurately predict the face prior representation and prior information in the generative model and effectively use the face prior features for efficient reconstruction is still a challenge in the current field of face reconstruction.
A face recovery method based on diffusion model and visual style prompt learning is adopted. Through a style encoder, denoising encoder and reconstruction network, combined with a style modulation aggregation layer, the encoding priors of low-quality images in hidden space are efficiently predicted, and the corresponding face feature information is predicted using compact encoding priors to achieve efficient face reconstruction.
Accurate prediction of high-quality representation of low-quality face images in the generative model is achieved, clear facial visual features are generated, the perceived quality of face reconstruction is improved, and face visual style prompts are used to predict face attributes.
Smart Images

Figure CN120047354A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image restoration and reconstruction, and in particular to a face restoration method based on a diffusion model and visual style prompt learning. Background Art
[0002] Blind face image reconstruction is an important branch of image restoration and reconstruction algorithms. Its corresponding goal is to reconstruct a low-quality face image (such as a blurred face image) without given the low-quality distribution, and obtain a face image as close as possible to the original image, making it look as real as possible and as close as possible to the original image.
[0003] Currently, due to the limited information available in low-quality pictures, methods based on face prior information dominate the existing methods. Current methods mainly include priors based on geometric information, such as using facial feature points and semantic maps of faces as guiding information. However, due to the limited guiding information provided by such methods, such as facial texture information, expression information, identity information, etc., there are still certain limitations in the reconstruction results of such algorithms. Another type of method mainly uses pre-trained generative models or face prior features in a feature vector quantization dictionary as guidance. However, when the degree of low quality is too severe, some key face information may be difficult to accurately estimate, resulting in using inaccurate face information for guidance, thus affecting the reconstruction effect.
[0004] Therefore, in the case of a given low-quality picture, how to accurately predict the face prior representation and prior information in the generative model, and effectively utilize the face prior features for efficient reconstruction, remains a challenge in the current face reconstruction field. Summary of the Invention
[0005] The technical problem to be solved by the embodiments of the present invention is to provide a face restoration method based on a diffusion model and visual style prompt learning, which can efficiently predict the encoding prior of a low-quality image in the latent space, use the compact encoding prior to predict the corresponding face feature information, and combine the face feature information to achieve efficient face reconstruction.
[0006] To solve the above technical problem, the embodiments of the present invention provide a face restoration method based on a diffusion model and visual style prompt learning, and the method includes the following steps:
[0007] Obtain a low-quality face image as input;
[0008] First, input the low-quality image into a style encoder to obtain an initial encoding, where is the initial encoding, I deis the corresponding low-quality image. And the initial encoding is used to guide the sampling and generation of the denoising encoding. In order to be able to sample the denoising encoding w 0 , the present invention realizes denoising by using T steps of denoising steps. For the t-th step, the corresponding formula is:
[0009]
[0010] where the noise is generated by the formula with variance and the random noise is
[0011] The denoising encoding is input into the face feature generator to obtain the corresponding face visual features. The corresponding expression formula is where is the corresponding set of face feature maps, represents the face feature map of the i-th layer; N represents the number of corresponding encoding vectors, and the total number is N.
[0012] Meanwhile, the low-quality face image, the denoising encoding, and the set of face feature maps are input into the reconstruction network to obtain the corresponding face reconstruction result. The corresponding formula is:
[0013]
[0014] where I out is the corresponding reconstruction output, with the corresponding image scale of height h, width w, and the number of channels 3. is the corresponding random style.
[0015] Among them, the reconstruction network introduces a novel style modulation aggregation layer to capture multi-scale feature information while using the denoising encoding and the random encoding to guide the reconstruction process. The encoder of the reconstruction network calculates the multi-scale encoded features where is the inference feature map of the i-th layer, and its calculation process is as follows:
[0016]
[0017] where SMART(·) is the style modulation aggregation layer, SC(·,·) is the style convolution; when i mod 2 = 1, the corresponding face feature map, denoising encoding, and random encoding are input into this layer for calculation. Otherwise, they are input into the style convolution and further input into the downsampling function (·)↓ 2 , for downsampling; [·,·] is the corresponding concatenation function; When i = N, the global encoding can be further obtained, and the corresponding formula is:
[0018] The corresponding decoder uses the above encoder features, denoising encoding, global encoding, and random encoding to calculate the formula for obtaining the decoder features as follows:
[0019]
[0020] Where Through the formula The final image is obtained Where, (·)↓ r And (·)↑ r Are the corresponding upsampling and downsampling functions, with r as the corresponding multiple, and the corresponding f(·) is the feature summation function.
[0021] Among them, the style modulation aggregation layer is used to capture multi-scale feature information while using denoising encoding to guide the reconstruction process. Given the input features of the image And the corresponding style hint information (In the encoder of the reconstruction network is In the decoder of the reconstruction network is ). The expression formula of this style modulation aggregation layer is:
[0022]
[0023] Where, Mod(·) is the corresponding modulation function, which uses the style vector obtained through transformation to scale the given convolution kernel parameters. And the corresponding Demod(·) is used to renormalize the scaled parameters for model training. Among them, Aggregation(·) is the corresponding convolution layer, which is used to aggregate the local and global face features extracted from different convolution kernel sizes.
[0024] Among them, the model of the style encoder is the e4e encoder; the encoding denoiser is a network constructed by combining channel and spatial self-attention mechanisms, which uses the time step t variable, the previous denoising encoding, and the initial encoding as inputs at the same time to obtain the denoising encoding result of the next step. The face feature generator comes from the StyleGAN model; the corresponding discriminator is used to train the reconstruction network to make the output result of the reconstruction network sharper.
[0025] Among them, the method needs to be trained in three stages. The first stage trains the style encoder to enable the encoder to extract the initial encoding from low-quality pictures. The second stage trains the denoiser to enable the denoiser to sample a high-quality denoising encoding based on the initial encoding. The third stage trains the reconstruction network to reconstruct the low-quality pictures using the predicted denoising encoding, random encoding, and face feature prior.
[0026] Among them, the denoiser performs backpropagation by calculating the diffusion loss, perceptual loss, and identity loss, and updates the model parameters using the stochastic gradient descent method. The formula for the diffusion loss is:
[0027]
[0028] where ∈ is the noise added during the diffusion process, and is the denoising result corresponding to the t-th step. The formula for calculating the perceptual loss is that when the model executes to the t = T step, the corresponding denoising code can obtain the corresponding face features by being input into the pre-trained face feature generator, and the corresponding expression formula is:
[0029]
[0030] where, is the corresponding set of face features, and I ve is the corresponding reconstructed image. S(·) is the corresponding pre-trained face feature generator. The formula for calculating the corresponding perceptual loss is:
[0031]
[0032] where VGG(·) is the corresponding pre-trained VGG network. The formula for calculating the identity loss is as follows:
[0033]
[0034] where cos(·) represents the cosine function and R(·) represents the pre-trained face image recognition model; the denoiser is jointly constrained using the diffusion loss, perceptual loss, and identity loss, and the total objective function is and the loss value is calculated according to the total objective function, and the network parameters θ of the denoiser are updated using the chain rule of differentiation p , and the network parameter update formula is η represents the learning rate in the hyperparameters, represents the parameter gradient corresponding to the denoiser module.
[0035] And the corresponding reconstruction network is jointly constrained by the perceptual loss, identity loss, and generative adversarial loss. The formula for calculating the perceptual loss is The formula for calculating the identity loss is The adversarial loss The formula for is where D(·) is the corresponding discriminator model and γ is the regularization weight; using the perceptual loss The identity loss And the adversarial loss To jointly constrain the reconstruction network and the discriminator network model, the total objective function is obtained as And calculate the loss value according to the total objective function, and use the chain rule of differentiation to update the reconstruction network parameter θ g And the network parameter θ of the discriminator d ; where The network parameter update formula is And η represents the learning rate in the hyperparameters, And represent the parameter gradients corresponding to the reconstruction network and the discriminator respectively.
[0036] Implementing the embodiments of the present invention has the following beneficial effects:
[0037] 1. The present invention combines a diffusion model to accurately predict the high-quality representation of a given low-quality face image in the corresponding generative model, so as to generate clear face visual features as prompt information.
[0038] 2. The present invention introduces a style modulation aggregation layer to capture multi-scale feature information and uses denoising coding to guide the generation of relevant styles and attributes of the face, thereby strengthening the face reconstruction process.
[0039] 3. The present invention introduces a visual style prompt framework for high-quality blind face image reconstruction, effectively uses face visual style prompts to predict face attributes, and uses a style modulation aggregation layer to improve the perceptual quality of the reconstructed face. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, obtaining other drawings based on these drawings still belongs to the scope of the present invention.
[0041] Figure 1 It is a flowchart of a face reconstruction model based on diffusion model and visual style prompt learning provided by an embodiment of the present invention;
[0042] Figure 2 It is an inference and training flowchart of a face restoration method based on diffusion model and visual style prompt learning provided by an embodiment of the present invention;
[0043] Figure 3 It is a diffusion and denoising flowchart of a denoiser in a face restoration method based on diffusion model and visual style prompt learning provided by an embodiment of the present invention;
[0044] Figure 4 This is the calculation process of the style modulation aggregation layer in a face restoration method based on diffusion model and visual style prompt learning provided by an embodiment of the present invention;
[0045] Figure 5 This is the face image reconstruction effect in a face restoration method based on diffusion model and visual style prompt learning provided by an embodiment of the present invention;
[0046] Figure 6 This is the specific application of a face restoration method based on diffusion model and visual style prompt learning provided by an embodiment of the present invention in face feature point detection;
[0047] Figure 7 This is the specific application of a face restoration method based on diffusion model and visual style prompt learning provided by an embodiment of the present invention in the face emotion recognition task. Detailed implementation manners
[0048] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.
[0049] As Figure 1 shown, in an embodiment of the present invention, a face restoration method based on diffusion model and visual style prompt learning is proposed, and the method includes the following steps:
[0050] Step S1, obtain a low-quality face image as the input; in the training stage, the low-quality image can be synthesized by using the following formula with a high-quality face image as the input. The synthesis formula is as follows:
[0051]
[0052] where, I de represents the corresponding low-quality image; I gt represents the high-quality face image; k σ is the corresponding Gaussian blur kernel, and σ is used to control the blur degree; n δ is the corresponding Gaussian random noise, and δ is used to control the degree; is the corresponding JPEG compression function, and is used to control the corresponding quality degree. Among them, (·)↓ r is the downsampling function, and r is the corresponding multiple. During the training process, the ranges of these parameters are σ ∈ [0.2, 10], r ∈ [1, 8], δ ∈ [0, 20],
[0053] Step S2: Input the obtained low-quality face image into the style encoder model to obtain an initial encoding, and use the encoding denoiser to optimize the initial encoding to obtain a denoised encoding. Input the denoised encoding into the face feature generator to generate rough face features. The generated face features, denoised encoding, and random noise will be input into the reconstruction network to obtain a reconstructed face image.
[0054] Specifically, before step S1, as Figure 2 shown, first construct a style hint module. The style hint module includes a style encoder and a denoiser module. Then construct a feature generator, a reconstruction network, and a corresponding discriminator network. The reconstruction network includes an autoencoder and a mapping network.
[0055] Both the style encoder and the face feature generator are pre-trained models. The role of the style encoder is to map the image into the latent space of the face feature generator to obtain an initial encoding. The role of the face feature generator is to sample the latent space and generate corresponding face features, which are used as the face feature library in the method design. The model of the style encoder can be the e4e encoder (Tov, Omer, Yuval Alaluf, Yotam Nitzan, Or Patashnik, and Daniel Cohen-Or. "Designing an encoder for stylegan image manipulation." ACM Transactions on Graphics (TOG) 40, no. 4 (2021): 1-14.).
[0056] The models of the mapping network, face feature generator, and discriminator come from the StyleGAN model (T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, ``Analyzing and improving the image quality of styleGAN,” in Proc. CVPR, 2020, pp. 8110-8119.).
[0057] The denoiser module is a network constructed using channel self-attention, spatial self-attention, and a gating mechanism, which is an extension of the network of the FFCLIP model (Zhu, Yiming, Hongyu Liu, Yibing Song, Ziyang Yuan, Xintong Han, Chun Yuan, Qifeng Chen, and Jue Wang. "One model to edit them all: Free-form text-driven image manipulation with semantic modulations." Advances in Neural Information Processing Systems 35 (2022): 25146-25159.). The corresponding difference is that the present invention improves the network by adding the time step t variable as an input and expanding the dimension to the calculation from multi-vector to multi-vector.
[0058] The reconstruction network introduces a novel style modulation aggregation layer to capture multi-scale feature information while using denoising coding and stochastic coding to guide the reconstruction process.
[0059] First, the low-quality image is input into the style encoder E(·) to obtain an initial encoding. And use the initial encoding to guide the sampling and generation of the denoising coding.
[0060] As Figure 2 shown, in order to be able to sample the denoising coding w 0 , the present invention realizes denoising by using a diffusion model to combine with the denoiser module. Before performing denoising, it is necessary to learn how to denoise. The diffusion model contains a diffusion and a denoising process. The diffusion process is mainly to continuously add noise to the given encoding, making it gradually degenerate into a vector or tensor obeying a Gaussian distribution. The diffusion process can be defined as:
[0061]
[0062] Among them, this diffusion process can be directly expressed as α t = 1 - β t , And the corresponding denoising process is mainly to learn how to gradually remove the noise of w t to obtain a cleaner version w t-1 . The corresponding expression can be written as:
[0063]
[0064] To obtain a high-quality style code w 0 , it is necessary to perform T steps of denoising steps to achieve denoising. For the t-th step, the corresponding formula is:
[0065]
[0066] where the noise is generated by the formula with variance and the random noise is P(·) is the corresponding denoiser module. The denoiser takes the given intermediate result the initial code and the time step t as inputs and performs four identical time-step-aware encoding mapping modules TACC for calculation. The corresponding process is: where The calculation details of are as follows:
[0067]
[0068] Then calculate the corresponding channel-based and space-based cross-attention, as well as the gated and bias results:
[0069]
[0070] where [·] is the concatenation operation; FC(·) is the fully connected layer; MLP(·) is a perceptron containing two layers of fully connected layers; Softmax(·) is the softmax activation function; σ(·) is the sigmoid activation function; φ(·) is the LeakyReLU activation function; LayerNorm(·) is the corresponding LayerNorm normalization layer.
[0071] Input the denoised code w 0 into the face feature generator to obtain the corresponding face visual features. The corresponding expression formula is where is the corresponding set of face feature maps, represents the i-th layer of face feature maps; N represents the number of corresponding encoding vectors, and the total number is N.
[0072] At the same time, input the low-quality face image, the denoised code, and the set of face feature maps into the reconstruction network to obtain the corresponding face reconstruction result. The corresponding formula is:
[0073]
[0074] where is the corresponding reconstruction output, with the corresponding image scale of height h, width w, and number of channels 3. is the corresponding random style, generated by the formula It is calculated. Among them, F(·) is the corresponding mapping network, which is optimized together with the parameters in the reconstruction network.
[0075] Among them, the reconstruction network introduces a novel style modulation aggregation layer to capture multi-scale feature information while using denoising coding and random coding to guide the reconstruction process. The encoder of the reconstruction network calculates multi-scale encoded features Among them is the inference feature map of the i-th layer, and its calculation process is as follows:
[0076]
[0077] Among them, SMART(·) is the style modulation aggregation layer, and SC(·,·) is the style convolution; when i mod 2 = 1, the corresponding face feature map, denoising coding, and random coding are input into this layer for calculation. Otherwise, it is input into the style convolution and further input into the downsampling function (·)↓ 2 , for downsampling; [·,·] is the corresponding concatenation function; When i = N, the global encoding can be further obtained, and the corresponding formula is:
[0078] The corresponding decoder uses the above encoder features, denoising coding, global coding, and random coding to calculate the formula for the decoder features as:
[0079]
[0080] Among them Through the formula The final image is obtained Among them, (·)↓ r and (·)↑ r are the corresponding upsampling and downsampling functions, with r as the multiple. f(·) is the feature summation function.
[0081] Among them, the style modulation aggregation layer is used to capture multi-scale feature information while using denoising coding to guide the reconstruction process. Given the input features of the image and the corresponding style hint information (in the encoder of the reconstruction network is in the decoder of the reconstruction network is The expression formula of this style modulation aggregation layer is:
[0082]
[0083] Among them, Mod(·) is the corresponding modulation function, which uses the style vector obtained through transformation to scale the given convolution kernel parameters. The corresponding Demod(·) is used to renormalize the scaled parameters to facilitate model training. Among them, Aggregation(·) is the corresponding convolutional layer, which is used to aggregate the local and global face features extracted from different convolution kernel sizes. The operation definitions of Mod(·) and Demod(·) come from the StyleGAN model (T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, "Analyzing and improving the image quality of styleGAN," in Proc. CVPR, 2020, pp. 8110-8119.)
[0084] Among them, the model of the style encoder is the e4e encoder; the encoding denoiser is a network constructed by combining channel and spatial self-attention mechanisms, which uses the time step t variable, the denoising encoding of the previous step, and the initial encoding as inputs to obtain the denoising encoding result of the next step. The face feature generator comes from the StyleGAN model; the corresponding discriminator is used to train the reconstruction network to make the output result of the reconstruction network sharper.
[0085] Among them, the method needs to be trained in three stages. In the first stage, the style encoder is trained so that the encoder can extract the initial encoding from low-quality images. In the second stage, the denoiser is trained so that the denoiser can sample a high-quality denoising encoding based on the initial encoding. In the third stage, the reconstruction network is trained to reconstruct low-quality images using the predicted denoising encoding, random encoding, and face feature prior.
[0086] The training pseudocode for training the style encoder in the first stage is as follows:
[0087]
[0088] In the second stage, the denoiser is trained. Among them, the denoiser performs backpropagation by calculating the diffusion loss, perceptual loss, and identity loss, and uses the stochastic gradient descent method to update the model parameters. The formula for the diffusion loss is:
[0089]
[0090] where ∈ is the noise added during the diffusion process, is the noise result corresponding to the t-th step. The calculation formula of the perceptual loss is that when the model executes to the t = T step, the corresponding denoising encoding can obtain the corresponding face features by inputting into the pre-trained face feature generator, and the corresponding expression formula is:
[0091]
[0092] where is the corresponding set of face features, and I ve is the corresponding reconstructed image. S(·) is the corresponding pre-trained face feature generator. The calculation formula of the corresponding perceptual loss is:
[0093]
[0094] where VGG(·) is the corresponding pre-trained VGG network. The calculation formula of the identity loss is as follows:
[0095]
[0096] where cos(·) represents the cosine function, and R(·) represents the pre-trained face image recognition model; the denoiser is jointly constrained by the diffusion loss, perceptual loss, and identity loss, and the total objective function is and the loss value is calculated according to the total objective function, and the network parameters θ of the denoiser are updated using the chain rule of differentiation p , and the network parameter update formula is η represents the learning rate in the hyperparameters, represents the parameter gradient of the corresponding denoiser module.
[0097] The training pseudo-code of the denoiser is as follows:
[0098]
[0099] In the third stage, the reconstruction network parameters are trained. The corresponding reconstruction network is jointly constrained by the perceptual loss, identity loss, and generative adversarial loss. The calculation formula of the perceptual loss is The calculation formula of the identity loss is The adversarial loss The formula is where D(·) is the corresponding discriminator model, and γ is the regularization weight; using the perceptual loss The identity loss and the adversarial loss to jointly constrain the reconstruction network and the discriminator network model, and the total objective function is and the loss value is calculated according to the total objective function, and the reconstruction network parameters θ are updated using the chain rule of differentiation gand the network parameters θ of the discriminator d ; where the network parameter update formula is and η represents the learning rate in the hyperparameters and represent the parameter gradients corresponding to the reconstruction network and the discriminator respectively
[0100]
[0101]
[0102] For the training of the above three stages, assuming that the current iteration number is q, the corresponding updated and optimized parameters are At the end of the parameter update stage, it is judged whether the training iteration number q has reached the maximum iteration number Q. If it has reached the maximum iteration number Q, the training stage ends and a trained face image reconstruction model is obtained; otherwise, it will jump to the first step for cyclic iterative training and let q = q + 1
[0103] In summary, first, the style encoder is used to encode the low-quality image, and a relatively rough initial encoding can be obtained. Since the initial encoding cannot accurately predict the encoding of the high-quality representation corresponding to the low-quality image in the generation latent space, through the diffusion denoising process, more accurate denoising and prediction of face attributes in the generation model latent space are realized. At the same time, the style modulation aggregation layer proposed by the present invention is used to capture multi-scale feature information while using denoising encoding and random encoding to guide the reconstruction process, thereby optimizing the perceptual quality of face reconstruction
[0104] Implementing the embodiments of the present invention has the following beneficial effects
[0105] 1. The present invention combines the diffusion model to accurately predict the high-quality representation of a given low-quality face image in the corresponding generation model, thereby generating clear face visual features as prompt information
[0106] 2. The present invention introduces a style modulation aggregation layer to capture multi-scale feature information while using denoising encoding to guide the generation of relevant styles and attributes of the face, thereby strengthening the face reconstruction process
[0107] 3. The present invention introduces a visual style prompt framework for high-quality blind face image reconstruction, effectively using face visual style prompts to predict face attributes, and using the style modulation aggregation layer to improve the perceptual quality of the reconstructed face
[0108] 4. The present invention can be further used to assist downstream tasks, such as improving the accuracy of face feature point detection and the accuracy of face emotion recognition, such asFigure 6 and Figure 7 as shown
[0109] Those of ordinary skill in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc.
[0110] The above-disclosed is only a preferred embodiment of the present invention, and of course, the scope of the rights of the present invention cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A face restoration method based on diffusion model and visual style cue learning, characterized in that: The method comprises the following steps: Step S1: constructing a face reconstruction model and training the model, wherein the face image reconstruction model is composed of a style encoder, a coding denoiser, a face feature generator, a reconstruction network and a discriminator; Step S2: obtaining a low-quality face image as input, and inputting the obtained face image into a trained face reconstruction model to obtain a reconstructed image; The low-quality image enters the style encoder to obtain an initial code, and the code is input into the code denoiser; after multiple denoising calculations, a denoised code is obtained; the denoised code is input into the face feature generator to obtain a rough face feature; At the same time, the low-quality image enters the reconstruction network for reconstruction; during the reconstruction process, the reconstruction network uses the denoising code and facial features generated above to optimize the generation of the reconstruction result.
2. The face restoration method based on diffusion model and visual style cue learning as claimed in claim 1, characterized in that: In step S2, By formula Get the initial code Among them, I de is the low-quality image; By formula In the iterative denoising process, the initial code is used as a guide to optimize the denoising, and the denoising code w is obtained through T optimizations. 0 ; Input the denoising code into the face feature generator to obtain the corresponding face visual features; the corresponding expression formula is: in, is the corresponding set of facial feature maps, represents the facial feature map of the i-th layer; N represents the number of corresponding encoding vectors, the total number is N.
3. The face restoration method based on diffusion model and visual style cue learning as claimed in claim 2, characterized in that: The reconstruction network introduces a novel style modulation aggregation layer to capture multi-scale feature information while using denoising coding and random coding to guide the reconstruction process; the encoder of the reconstruction network obtains multi-scale coding features by calculation in is the inference feature map of the i-th layer, and its calculation process is as follows: Among them, SMART(×) is the style modulation aggregation layer, SC(·,·) is the style convolution; when imod2=1, the corresponding feature map, denoising code and random code are input into this layer for calculation; otherwise, they are input into the style convolution and further input into the downsampling function (·)↓2 for downsampling; [·,·] is the corresponding concatenation function; When i=N, the global encoding can be further obtained, and the corresponding formula is: The corresponding decoder uses the above encoder features, denoising coding, global coding and random coding to calculate the decoder features. The calculation formula is: in By formula Get the final image Among them, (·)↓ r and (·)↑ r The corresponding f(·) is the characteristic summation function.
4. The method for face restoration based on diffusion model and visual style cue learning as claimed in claim 2, characterized in that: The reconstruction network introduces a novel style modulation aggregation layer to capture multi-scale feature information while using denoising coding to guide the reconstruction process; given the input features of the image And the corresponding style prompt information Among them, in the encoder of the reconstruction network is In the decoder of the reconstruction network, The expression formula of the style modulation aggregation layer is: Among them, Mod(×) is the corresponding modulation function, which uses the style vector obtained by transformation to scale the given convolution kernel parameters; and the corresponding Demod(×) is used to renormalize the scaled parameters to facilitate model training; among them, Aggregation(×) is the corresponding convolution layer, which is used to aggregate local and global facial features extracted from different convolution kernel sizes.
5. The face restoration method based on diffusion model and visual style cue learning as claimed in claim 2, characterized in that: The model of the style encoder is an e4e encoder; the coding denoiser is a network constructed by combining channel and spatial self-attention mechanisms, which simultaneously uses the time step variable t, the denoising code of the previous step and the initial code as input to obtain the denoising coding result of the next step; The facial feature generator comes from the StyleGAN model; the corresponding discriminator is used to train the reconstruction network to make the output of the reconstruction network sharper.
6. The method for face restoration based on diffusion model and visual style cue learning as claimed in claim 2, characterized in that: The method performs three-stage training; the first stage trains the style encoder so that the encoder can extract the initial code from the low-quality picture; The second stage trains the denoiser so that it can sample a high-quality denoised code based on the initial code. The third stage trains the reconstruction network to reconstruct low-quality images using the predicted denoising code, random code, and facial feature priors.
7. The face restoration method based on diffusion model and visual style cue learning as claimed in claim 2, characterized in that: During the training process of the denoiser module, back propagation is performed by calculating the diffusion loss, perceptual loss and identity loss, and the model parameters are updated using the stochastic gradient descent method; the formula for the diffusion loss is: Among them, ∈ is the noise added in the diffusion process, is the denoising result corresponding to the tth step; the calculation formula of the perceptual loss is that when the model executes to the t=Tth step, the corresponding denoising code can be input into the pre-trained face feature generator to obtain the corresponding face feature, and the corresponding expression formula is: in, is the corresponding facial feature set, I ve is the corresponding reconstructed image; S(·) is the corresponding pre-trained face feature generator; the corresponding perceptual loss calculation formula is: Among them, VGG(×) is the corresponding pre-trained VGG network; the calculation formula of identity loss is as follows: Where cos(×) represents the cosine function, R(·) represents the pre-trained face image recognition model; the denoiser is jointly constrained by diffusion loss, perceptual loss and identity loss, and the total objective function is: And calculate the loss value according to the total objective function, and use the chain derivation rule to update the network parameters θ of the denoiser p , the network parameter update formula is η represents the learning rate in the hyperparameters, represents the parameter gradient of the corresponding denoiser module; The corresponding reconstruction network is jointly constrained by perceptual loss, identity loss and generative adversarial loss; the calculation formula of perceptual loss is: The formula for calculating identity loss is: Fighting Losses The formula is Among them, D(·) is the corresponding discriminator model, γ is the regularization weight; using perceptual loss Identity loss and combat loss To jointly constrain the reconstruction network and the discriminator network model, the total objective function is And calculate the loss value according to the total objective function, and use the chain derivation rule to update the reconstruction network parameters θ g and the network parameters θ of the discriminator d ;in, The network parameter update formula is: as well as , η represents the learning rate in the hyperparameter, and represents the parameter gradients of the corresponding reconstruction network and the discriminator.