Method for generating face image based on sketch guided by reference image

Through the combination of the global-local fast Fourier convolution module and the adaptive cross-domain attention module, the problem of structural distortion and artifacts in the sketch-generated face image is solved, and more realistic face images are generated.

CN120411283APending Publication Date: 2025-08-01HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510523975.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing sketch-generating face image technology is prone to structural distortion and unnatural artifacts during the generation process, and it is difficult to generate real textures while retaining structural information.

Method used

Using the global-local fast Fourier convolution module and the adaptive cross-domain attention module, combined with the generator and discriminator network, through global context modeling and feature fusion, the generator network uses the mapping features of the reference image for normalization and adjustment. The generator network includes a face sketch encoder, a reference image encoder, an adaptive cross-domain attention module and a decoder. The discriminator network is used to improve the generated image quality.

Benefits of technology

The structural consistency and detail authenticity of the generated face images are improved, and the generated images are more realistic, overcoming the problem of structural distortion in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411283A_ABST
    Figure CN120411283A_ABST
Patent Text Reader

Abstract

The invention discloses a method for generating a face image based on a sketch guided by a reference image. The method comprises the following steps: firstly, respectively extracting face sketch features and reference image features by using a feature extraction network of a generator, wherein the feature extraction network comprises an alternate down-sampling layer and a global-local fast Fourier convolution module; and the extracted sketch and reference image features are fused through an adaptive cross-domain attention module. In the decoding stage, normalization adjustment is carried out by using the mapping features of the reference image, and a face image is generated in combination with the up-sampling layer. And inputting the generated face image and the real face image into a discriminator network. In the training process, the total loss is calculated by using the antagonism loss, the cyclic consistency loss, the perception loss and the learning perception image block similarity loss, the loss function training network is minimized, and finally the sketch and the reference image are input into the trained generator network to generate the realistic face image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, relates to an image generation method, and particularly relates to a method for generating a face image based on a reference image-guided sketch. Background Art

[0002] Different from non-face image synthesis, face image generation has more stringent standards for the consistency of the facial structure and the rationality of the detailed texture of the face. Defects (such as structural missing, artifacts, etc.) in the generated face image are more likely to be detected, making face image synthesis more challenging.

[0003] Generating a face image from a sketch refers to the process of automatically synthesizing a realistic face image from a sketch provided by a user. The technology of generating a face image from a sketch can be divided into traditional methods and deep learning-based methods. The traditional methods for generating a face image from a sketch mainly include the example-based method and the linear regression-based method. The example-based method reconstructs an image by matching the input image with example images, focusing on the correspondence between the two. If the images are misaligned, it will be severely distorted, limiting its practical application. The linear regression-based method focuses on learning the linear mapping between facial images and sketches, significantly reducing the computational complexity. However, due to the limited expressive ability of the linear regression model and the large difference in the feature distributions between the source domain and the target domain, it cannot accurately represent the complex non-linear mapping between the two domains. The deep learning-based method benefits from the excellent representation ability and mapping ability of the deep neural network, and can better handle complex non-linear relationships, having significant advantages compared with traditional methods.

[0004] Currently, certain progress has been made in the task of generating a face image from a sketch. Some methods use a face semantic parsing map to assist in the synthesis of the target domain image and introduce a facial component loss during network training, making the synthesized facial features more distinct and ensuring the stability of the facial structure. However, artifacts may appear unnaturally during the generation process. Other methods adopt a channel attention mechanism (CAM) to enhance the transformation ability of key regions during the image style transfer process and combine it with an adaptive instance normalization layer (AdaIN) to automatically adjust the ratio of instance normalization (IN) and layer normalization (LN) in the normalization, enabling the model to more flexibly control the changes in the image shape and texture. However, structural distortion problems are likely to occur in practical applications.

[0005] In summary, since generating face images from sketches has high requirements for structural consistency and the rationality of detailed textures, it is of great significance to develop a face image synthesis method that can generate realistic textures while retaining structural information. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the present invention proposes a method for generating face images from sketches guided by a reference image. It uses a global-local fast Fourier convolution module to provide a more efficient global context modeling method in the feature extraction stage, and uses an adaptive cross-domain attention module to fuse the features of the extracted sketches and reference images. In the decoder stage of the generator, it uses the mapped features of the reference image for normalization adjustment to maintain the naturalness of the generated face images in local details.

[0007] A method for generating face images from sketches guided by a reference image is as follows:

[0008] Step 1: Use a randomly selected face image as a reference image, which together with the face sketch serves as an input sample, and use the real face image corresponding to the face sketch as a label to form a training set. Figure 1 The real face image corresponding to the face sketch is used as a label to form a training set.

[0009] Step 2: Construct a face image generation network model based on the reference image, including a generator network and a discriminator network.

[0010] The generator network extracts features from the face sketch and the reference image through a face sketch encoder and a reference image encoder respectively, then fuses the face sketch features and the reference image features through an adaptive cross-domain attention module, and finally decodes the features fused by the adaptive cross-domain attention module through a decoder, and uses the mapped features of the reference image for normalization adjustment to output the generated face image. The discriminator network is used to discriminate the face images generated by the generator network to help the generator network gradually improve the quality of the generated images. The face sketch encoder and the reference image encoder have the same structure, including alternating downsampling layers and global-local fast Fourier convolution modules to enhance the global context awareness ability of the model and ensure the facial structure consistency of the generated images. Among them, the downsampling layer generates feature maps with different resolutions through convolution. The global-local fast Fourier convolution module outputs features with a receptive field covering the entire image and local features through a fast Fourier convolution unit and a local feature extraction unit respectively, and then aggregates the useful features of both through a bidirectional gated channel attention module to reduce information redundancy and output the output of the global-local fast Fourier convolution module.

[0011] Step 3: Input the training set data in Step 1 into the generative adversarial network model constructed in Step 2, and use the Adam optimization algorithm to optimize the network model. The loss function includes adversarial loss, cycle consistency loss, perceptual loss, and Learned Perceptual Image Patch Similarity (LPIPS) loss. After training is completed, save the optimized model parameters.

[0012] Step 4: Select a face sketch and a reference image, and input them into the generator network obtained by training in Step 3 to generate a realistic face image.

[0013] The present invention has the following beneficial effects:

[0014] (1) Aiming at the limitation of the traditional convolution-based feature extraction module in understanding the global structure of images, that is, the problem of small underlying receptive fields and difficulty in capturing non-local context relationships, the global-local fast Fourier convolution module is introduced. This module performs convolution in the frequency domain to obtain a global receptive field, combines depthwise separable convolution and standard convolution to extract local features, and then fuses global and local information through a gated channel attention mechanism, so as to efficiently extract local features and model the global context. Compared with other methods based on convolutional networks, it can better retain the overall structure of the sketch.

[0015] (2) This method performs feature pruning and compression on the input sketch features and reference image features through the adaptive cross-domain attention module, generates a mask ratio, divides the features into relevant information and irrelevant information, processes and combines these information respectively, and performs adaptive normalization adjustment through the mapped features of the reference image in the decoder, so that the generated face image is more delicate and realistic in local details. Description of the Drawings

[0016] Figure 1 is a flowchart of the method for generating a face image from a sketch guided by a reference image;

[0017] Figure 2 is a schematic diagram of the generator network structure constructed in the embodiment;

[0018] Figure 3 is a schematic diagram of the global-local fast Fourier convolution module structure constructed in the embodiment;

[0019] Figure 4 is a schematic diagram of the fast Fourier convolution unit structure constructed in the embodiment;

[0020] Figure 5 is a schematic diagram of the local feature extraction unit structure constructed in the embodiment;

[0021] Figure 6 It is a schematic structural diagram of the bidirectional gated channel attention module constructed in the embodiment;

[0022] Figure 7 It is a schematic structural diagram of the adaptive cross-domain attention module constructed in the embodiment;

[0023] Figure 8 It is a schematic structural diagram of the discriminator network constructed in the embodiment. Specific implementation manners

[0024] The present invention will be further explained below with reference to the accompanying drawings;

[0025] As Figure 1 shown, a method for generating a face image from a sketch guided by a reference image specifically includes the following steps:

[0026] Step 1: Use a randomly selected face image as a reference image, together with a face sketch Figure 1 as input samples, and use the real face image corresponding to the face sketch as a label to form a training set.

[0027] Step 2: Construct a sketch-to-face network model based on the reference image, including a generator network and a discriminator network. As Figure 2 shown, the generator network includes a face sketch encoder, a reference image encoder, an adaptive cross-domain attention module, and a decoder.

[0028] s2.1: The face sketch encoder and the reference image encoder have the same structure and are respectively used to extract the features of the sketch and the reference image, including 4 alternately cascaded downsampling layer encoding units and a global-local fast Fourier convolutional module.

[0029] The downsampling layer encoding unit includes a convolutional layer with a convolution kernel of 3, a padding of 1, and a stride of 2, a spectral normalization (SpectralNorm) layer, and an activation function ReLU. Each downsampling layer encoding unit outputs feature maps with different resolutions.

[0030] As Figure 3 shown, the global-local fast Fourier convolutional module includes a fast Fourier convolutional unit, a local feature extraction unit, and a bidirectional gated channel attention module.

[0031] Among them, the fast Fourier convolutional unit is as Figure 4 shown. First, perform a convolution on the feature T output by the encoding unit with a convolution kernel size of 1 and a stride of 1, halve the number of channels, and after BatchNorm and ReLU processing, generate the feature T':

[0032]

[0033] Then, perform a two-dimensional fast Fourier transform on the feature T′ to convert it from the spatial domain to a complex tensor in the frequency domain

[0034]

[0035] Then, for the complex tensor concatenate the real part and the imaginary part to obtain the intermediate tensor F:

[0036]

[0037] Perform a convolution with a kernel size of 1 and a stride of 1 on the intermediate tensor F in the frequency domain, followed by BatchNorm and ReLU processing to obtain the feature F′:

[0038]

[0039] Use the two-dimensional inverse fast Fourier transform to convert the feature back to the spatial domain to obtain the spatial domain feature T″:

[0040]

[0041] Finally, pass through a convolutional layer with a kernel size of 1 and a stride of 1 to restore the original number of channels and output the feature T″′ whose receptive field covers the entire image:

[0042]

[0043] The local feature extraction unit efficiently extracts local features by combining depthwise separable convolution and standard convolution. As Figure 5 shown, for the feature T output by the encoding unit, first perform a depthwise separable convolution with a kernel size of 3, a padding of 1, and independent for each channel to extract local information, and then perform normalization through LayerNorm. Then, perform a 1×1 convolution for channel transformation and pass through GELU activation to enhance the non-linear expression ability. Subsequently, adjust the feature through a 1×1 convolution again, and finally perform a residual connection between the output feature and the input feature.

[0044] The described bidirectional gated channel attention module effectively aggregates the useful features of both by integrating the complementary advantages of global and local information, thereby reducing information redundancy. As Figure 6 shown, the bidirectional gated channel attention module processes the outputs of the local feature extraction unit and the fast Fourier convolution unit through a 1×1 convolution to obtain the local feature F local and the fast Fourier feature F FFC , to enhance the expression ability of the feature. At the same time, the feature Z obtained by adding the outputs of the two units passes through a Gate unit to obtain the weight α:

[0045] α = Gate(Z) = Sigmod(Conv 1×1 (ReLU(Conv 1×1 (Z)))) (7)

[0046] Use the weight α to adaptively weight - fuse the local feature F local and the fast Fourier feature F FFC to obtain the fused feature F out :

[0047]

[0048] Finally, pass F out through a convolutional layer with a kernel size of 3, padding of 1, and stride of 1, and process it through BatchNorm and ReLU as the output of the global - local fast Fourier convolution module.

[0049] s2.2. Prune and compress the face sketch features and reference image features, then generate a mask ratio through the sign function (Sgn) to divide the features into relevant information and irrelevant information. Process and fuse these information respectively to more accurately capture local details and improve the detail authenticity of the generated image:

[0050] As Figure 7 shown, the adaptive cross - domain attention module first performs 1×1 convolution on the face sketch feature F s to generate the query matrix Q, performs 1×1 convolution on the reference image feature F r to generate the key matrix K and the value matrix V, and then generates the attention weight attn:

[0051] Q = Conv 1×1 (F s ), K = Conv 1×1 (F r ), V = Conv 1×1 (F r ) (9)

[0052]

[0053] where d k represents the dimension of the key vector K. Then perform 1×1 convolution on the face sketch feature F s and the reference image feature F r respectively to generate the results Q prune and K prune of feature pruning and compression, and then generate a mask mask through the sign function Sgn() to shield unimportant features:

[0054] Q prune = Conv 1×1 (F s ), K prune = Conv 1×1 (F r ) (11)

[0055] mask = Sgn(Q prune ·K prune T ) (12)

[0056] The superscript T represents transpose.

[0057] Map the reference image feature F r to a new latent space through a mapping network to obtain the mapped feature M r , and input it into the adaptive instance normalization processing module AdaIN(), and normalize the face sketch feature F s to enhance the rationality of local texture features:

[0058]

[0059] where F AdaIN represents the feature after being processed by the AdaIN module. μ() represents calculating the mean of the feature, and σ() represents calculating the variance of the feature.

[0060] The mapping network includes three fully connected layers, and non-linear transformation is performed between each two fully connected layers through the LeakyReLU activation function to obtain the mapped latent vector.

[0061] Finally, use the mask mask and the attention weight attn to generate the masked weight α, and perform adaptive feature fusion on the value matrix V and the feature F AdaIN processed by the AdaIN module, and output the fused feature F out :

[0062]

[0063] s2.3. The decoder includes 4 decoding units, a convolutional layer with a convolutional kernel size of 3, padding of 1, and stride of 1, and an activation function Tanh, which are used to decode the feature F out fused by the adaptive cross-domain attention module, and perform instance adaptive normalization using the mapped feature of the reference image to maintain the naturalness of local details of the face picture, so as to obtain a realistic face image.

[0064] The decoding unit first uses the feature M r mapped by the reference image through the mapping network to the input feature Xi Perform instance adaptive normalization processing to maintain the naturalness of local details:

[0065]

[0066] Then, X AdaINi is concatenated with the output feature map S of the face sketch encoder at the same level i in the channel dimension. The concatenated features are upsampled by a factor of 2 using bilinear interpolation, and then the number of channels is compressed by a convolution with a kernel size of 3, padding of 1, and stride of 1 to obtain the output Z i ′ of the decoding unit:

[0067] Z i ′ = Conv 3×3 (Up(Concat(X AdaINi , S i ))) (16)

[0068] where i = 1, 2, 3, 4, X i = Z′ i-1 , represents the input of the i-th decoding unit, and Z′0 = F out .

[0069] s2.4. As shown in Figure 8 , the discriminator network is used to discriminate the face images generated by the generator, helping the generator gradually improve the quality of the generated images. The discriminator network includes 3 convolutional layers with a size of 4×4, a stride of 2, and a padding of 1, and combines batch normalization layers and a LeakyReLU with a slope of 0.2 for feature extraction, and finally outputs the discrimination result through a convolutional layer with a size of 1×1.

[0070] Step 3. Initialize the generator network and the discriminator network in the face generation network model based on the reference image sketch using the normal distribution (Normal), and use the instance normalization layer (InstanceNorm) to independently calculate the mean and variance for each channel of each sample.

[0071] Input the training set data in Step 1 into the initialized face generation network model based on the reference image sketch, and use the Adam optimization algorithm to optimize the network model. Set the total loss function L total as:

[0072] L total = λ1L adv + λ2L cyc + λ3L per + λ4L LPIPS (18)

[0073]

[0074] where λ1 = 1, λ2 = 1, λ3 = 0.5, λ4 = 5 represent the weight coefficients of the loss function, and L adv and L cyc and L per and L LPIPS represent adversarial loss, cycle consistency loss, perceptual loss, and learned perceptual image patch similarity loss respectively. x represents the sketch image, y represents the real face image, r represents the reference face image, G() represents the generator network, D() represents the discriminator network, F() represents the generator network that converts back to the source domain, and φ j () represents calculating the perceptual distance between features using the intermediate features of the j-th layer of the pre-trained VGG network, C j represents the number of channels of the feature map of the j-th layer of the VGG network, H j represents the width of the feature map of the j-th layer of the VGG network, W j represents the height of the feature map of the j-th layer of the VGG network, and LPIPS() represents the learned perceptual image patch similarity loss. E x [ ] represents taking the expectation with respect to x, and E y [ ] represents taking the expectation with respect to y, and r x represents the reference sketch image, and r y represents the reference real face image, and E x~pdata(x) [ ] represents taking the expectation with respect to the samples of the real data distribution, and ||||1 represents the L1 norm.

[0075] Specify the number of iterations to update the network parameters and save the trained model parameters.

[0076] Step 4: Select a face sketch and a reference image, and input them into the generator network trained in Step 3 to generate a realistic face image.

[0077] To evaluate the performance of this method, comparative experiments were conducted with classical image generation methods on the CUFS dataset, and the complete data was randomly divided into a training set and a test set. This database includes three sub-databases: the CUHK Student Database, the AR Database, and the XM2VTS Database, with a total of 606 face photo pairs. For the CUHK Student Database, 88 data pairs were selected as the training set and 100 data pairs as the test set; for the AR Database, 80 data pairs were selected as the training set and 43 data pairs as the test set; for the XM2VTS Database, 100 data pairs were selected as the training set and 195 data pairs as the test set. LPIPS (Learned Perceptual Image Patch Similarity), FID (Fréchet Inception Distance), and FSIM (Feature Similarity) were used as evaluation metrics for quantitative comparative analysis to evaluate the quality of the generated images. The results are shown in Table 1:

[0078] Table 1

[0079]

[0080] As can be seen from Table 1, this method obtained the best or second-best scores in various evaluation metrics, indicating that this method performs excellently in the quality of the generated images, proving the effectiveness of this method.

Claims

1. A method for generating a face image from a sketch guided by a reference image, characterized in that: The specific method is as follows: Step 1: Use a randomly selected face image as a reference image, together with the face sketch as an input sample, and use the real face image corresponding to the face sketch as a label to form a training set; Step 2: Construct a sketch-to-face network model based on the reference image, including a generator network and a discriminator network; The generator network extracts features from the face sketch and the reference image through a face sketch encoder and a reference image encoder respectively, then fuses the face sketch features and the reference image features through an adaptive cross-domain attention module, and finally decodes the features fused by the adaptive cross-domain attention module through a decoder, and performs normalization adjustment using the mapping features of the reference image to output the generated face image; The face sketch encoder and the reference image encoder have the same structure, including alternating downsampling layer encoding units and global-local fast Fourier convolutional modules; among them, the downsampling layer generates feature maps with different resolutions through convolution, and the global-local fast Fourier convolutional module outputs features with receptive fields covering the entire image and local features through a fast Fourier convolutional unit and a local feature extraction unit respectively, and aggregates them through a bidirectional gated channel attention module as the output of the global-local fast Fourier convolutional module; The discriminator network is used to discriminate the face images generated by the generator network to help the generator network gradually improve the quality of the generated images; Step 3: Input the training set data in Step 1 into the generative adversarial network model constructed in Step 2, and use the Adam optimization algorithm to optimize the network model. After training is completed, save the optimized model parameters; Step 4: Select a face sketch and a reference image, input them into the generator network trained in Step 3, and generate a realistic face image.

2. The method for generating a face image from a sketch based on reference image guidance according to claim 1, wherein: The downsampling layer encoding unit includes a convolutional layer with a convolution kernel of 3, a padding of 1, and a stride of 2, a spectral normalization layer, and an activation function ReLU. Each downsampling layer encoding unit outputs feature maps with different resolutions.

3. The method for generating a face image from a sketch guided by a reference image according to claim 1, characterized in that: The fast Fourier convolutional unit first performs a convolution with a convolution kernel size of 1 and a stride of 1 on the feature T output by the encoding unit, halves the number of channels, and then processes it through BatchNorm and ReLU to generate the feature T′: T′ = ReLU(BN(Conv 1×1 (T)) (1) Then, perform a two-dimensional fast Fourier transform on the feature T′ to convert it from the spatial domain to a complex tensor in the frequency domain Then, for the complex tensors concatenate the real and imaginary parts to obtain the intermediate tensor F: Perform a convolution with a convolution kernel size of 1 and a stride of 1 on the intermediate tensor F in the frequency domain, as well as BatchNorm and ReLU processing, to obtain the feature F′: F′ = ReLU(BN(Conv 1×1 (F)) (4) Use the two-dimensional inverse fast Fourier transform to convert the feature back to the spatial domain to obtain the spatial domain feature T″: Finally, pass through a convolutional layer with a convolution kernel size of 1 and a stride of 1 to restore the original number of channels and output the feature T″′ with receptive fields covering the entire image: T″′ = Conv 1×1 (T″) (6) The local feature extraction unit first performs depthwise separable convolution with a convolution kernel size of 3, padding of 1, and each channel independent on the feature T output by the coding unit to extract local information, and then normalizes it through LayerNorm; then performs channel transformation through 1×1 convolution and enhances the non-linear expression ability through GELU activation; subsequently, adjusts the feature through 1×1 convolution again, and finally performs residual connection on the output feature and the input feature; The bidirectional gated channel attention module processes the outputs of the local feature extraction unit and the fast Fourier convolution unit through 1×1 convolution to obtain the local feature F local and the fast Fourier feature F FFC , and then the feature Z after adding the outputs of the two units passes through a Gate unit to obtain the weight α: α = Gate(Z) = Sigmod(Conv 1×1 (ReLU(Conv 1×1 (Z)))) (7) Use the weight α for the local feature F local and the fast Fourier feature F FFC to perform adaptive weighted fusion to obtain the fused feature F out : F out = α·F local + (1 - α)·F FFC (8) Finally, take F out through a convolutional layer with a convolutional kernel size of out 3, padding of out 1, and stride of out 1, and process it through BatchNorm and ReLU, which serves as the output of the global-local fast Fourier convolution module.

4. The method for generating a face image from a sketch guided by a reference image according to claim 1, wherein: The adaptive cross-domain attention module first processes the face sketch feature F s through a 1×1 convolution to generate a query matrix Q, and then processes the reference image feature F r through a 1×1 convolution to generate a key matrix K and a value matrix V, and generates an attention weight attn: Q=Conv 1×1 (F s ),K=Conv 1×1 (F r ),V=Conv 1×1 (F r ) (9) where d k represents the dimension of the key vector K; Then, the face sketch feature F s and the reference image feature F r are respectively subjected to 1×1 convolution to generate the results Q prune and K prune of feature pruning and compression. Then, the unimportant features are masked by processing through the sign function Sgn() to generate the mask mask: Q prune = Conv 1×1 (F s ), K prune = Conv 1×1 (F r ) (11) mask = Sgn(Q prune ·K prune T ) (12) The superscript T represents transpose; Map the reference image feature F r to a new latent space through a mapping network to obtain the mapped feature M r , and input it into the adaptive instance normalization processing module AdaIN() to normalize the face sketch feature F s to enhance the rationality of local texture features: Among them, F AdaIN represents the feature processed by the AdaIN module; μ() represents calculating the mean of the feature, and σ() represents calculating the variance of the feature; Finally, use the mask and the attention weight attn to generate the masked weight α, and perform adaptive feature fusion on the value matrix V and the feature F processed by the AdaIN module, and output the fused feature F AdaIN : out ​ 5. The method for generating a face image from a sketch guided by a reference image according to claim 4, wherein: The mapping network includes three fully-connected layers, and non-linear transformation is performed between each two fully-connected layers through the LeakyReLU activation function.

6. The method for generating a face image from a sketch guided by a reference image according to claim 1, characterized in that: The decoder includes 4 decoding units, a convolutional layer with a convolutional kernel size of 3, padding of 1, and stride of 1, and an activation function Tanh, which are used to decode the feature F fused by the adaptive cross-domain attention module out and perform instance adaptive normalization using the mapped features of the reference image to maintain the naturalness of the local details of the face image, so as to obtain a realistic face image.

7. The method for generating a face image from a sketch guided by a reference image according to claim 6, wherein: The decoding unit first uses the feature M obtained by mapping the reference image through the mapping network r , and performs instance adaptive normalization processing on the input feature X i to maintain the naturalness of local details: Then X AdaINi is concatenated with the output feature map S of the face sketch encoder at the same level i in the channel dimension. The concatenated features are upsampled by a factor of 2 using bilinear interpolation, and then the number of channels is compressed by a convolution with a kernel size of 3, padding of 1, and stride of 1 to obtain the output Z i ′ of the decoding unit Z i ′ = Conv 3×3 (Up(Concat(X AdaINi , S i ))) (16) where i = 1, 2, 3, 4, X i = Z' i-1 , representing the input of the i-th decoding unit, Z'0 = F out .

8. The method for generating a face image from a sketch based on reference image guidance according to claim 1, characterized in that: The discriminator network includes three convolutional layers with a size of 4×4, a stride of 2, and a padding of 1, and combines batch normalization layers and a LeakyReLU with a slope of 0.2 for feature extraction, and finally outputs the discrimination result through a convolutional layer with a size of 1×1.

9. The method for generating a face image from a sketch guided by a reference image according to claim 1, wherein: During the retraining process, set the total loss function L total as follows: L total = λ1L adv + λ2L cyc + λ3L per + λ4L LPIPS (18) where λ1 to λ4 represent the weight coefficients of the loss function, and L adv , L cyc , L per , L LPIPS represent the adversarial loss, cyclic consistency loss, perceptual loss, and learning perceptual patch similarity loss, respectively.

10. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method according to any one of claims 1 to 9.