A Lightweight Face Image Generation Method Based on Improved StarGAN and Knowledge Distillation

By introducing knowledge distillation technology and designing a variety of loss functions in the StarGAN v2 model, the performance of the student model is optimized, and the problems of high computational complexity and difficulty in lightening are solved, and the effects of high-quality face image generation and model lightening are achieved.

CN119863544BActive Publication Date: 2025-05-30CHANGCHUN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510345420.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-05-30
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

Due to the complex network structure and numerous parameters, the existing StarGAN v2 model has high computational complexity and large model size, making it difficult to deploy on resource-constrained devices, and it is difficult to achieve lightweight without reducing image generation quality.

Method used

Introduce knowledge distillation technology in the StarGAN v2 model, using the pre-trained StarGAN v2 model as the teacher model to guide the student model. By designing a series of loss functions, such as pixel-level MSE loss, perceived loss, triple loss, GKA loss, and wavelet loss, the performance of the student model is optimized so that it reduces the amount of model parameters and computational complexity while maintaining high generation quality.

Benefits of technology

It realizes the reduction of model parameters and calculation complexity, improve inference speed, adapt to resource-constrained equipment deployment, and enhance the generalization ability and adaptability of the model without reducing the image generation quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119863544B_ABST
    Figure CN119863544B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image data processing or generation, and relates to a lightweight face image generation method based on improved StarGAN and knowledge distillation. This method first constructs a student model and a teacher model. The teacher model uses the StarGAN v2 model, and the teacher model is pre-trained. Then, knowledge distillation is carried out. During knowledge distillation, a loss function is set, and the training data is input into the teacher model and the student model. During the knowledge distillation process, the activation values of each layer of the generator in the forward propagation are obtained, and the characteristics output of the pre-trained teacher model is used as the soft label of the student model to minimize the loss function. The loss function includes pixel-level MSE loss, perceptual loss, triplet loss, GKA loss, and feature-based wavelet loss. In this way, while maintaining a high generation quality, the student model can reduce the number of model parameters and computational complexity, thus achieving the lightweight goal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image data processing or generation, and relates to computer vision face image generation. Specifically, it relates to a method for face image generation based on an improved generative adversarial network (StarGAN) and knowledge distillation. Background Art

[0002] With the continuous progress of artificial intelligence technology, significant development has been achieved in the field of computer vision, especially in image processing and generation. In the prior art, generative adversarial (GAN) models, especially high-performance image generation models such as StarGAN, can generate high-quality images. However, in the optimization of GAN models, there is usually a balance problem between generation quality and model size. Larger models usually provide higher generation quality, but overly large models are slower during inference and difficult to meet the requirements of real-time generation.

[0003] StarGAN v2 is an image-to-image translation model developed by the Clova AI team, aiming to address the limitations of existing models in terms of diversity and multi-domain scalability. It achieves diverse image generation and multi-domain scalability through a single framework, significantly surpassing existing baseline models. The generator only receives the input image x and the encoding s of a specific style, and injects s into the generation process through adaptive instance normalization (AdaIN) to generate an image with the style of s. The mapping network transforms the latent code into style encodings for multiple domains. The style encoder extracts the style encoding from the given image. The discriminator adopts a multi-task output branch, and each branch distinguishes real images and fake images from a specific region. This model has strong multi-domain conversion ability and high-quality generated images. However, due to its complex network structure and numerous parameters, it consumes a large amount of computing power during training and inference, with large computational and storage overheads, and is difficult to be deployed on resource-constrained devices, thus limiting its application scope. Therefore, how to further reduce the model volume, reduce the number of model parameters and computational complexity, and improve the inference speed while not reducing the image generation quality (ensuring high generation quality) and fully considering the lightweight requirement has become an urgent problem to be solved.

[0004] Knowledge distillation technology has achieved remarkable success in the field of deep learning, especially in the training of lightweight models. However, in complex image generation models such as StarGAN, the application of knowledge distillation is still relatively rare. Currently, knowledge distillation is usually applied in the classification field, and there has not been enough research and exploration on the optimization of generation tasks, especially in generative adversarial models. Therefore, how to use knowledge distillation technology to improve the lightweight and performance of the StarGAN model is a deficiency in the current technology. Summary of the Invention

[0005] In view of the above technical problems and deficiencies, the present invention provides a lightweight face image generation method based on improved StarGAN and knowledge distillation. This method introduces knowledge distillation technology into the StarGAN v2 model, calls the pre-trained model of StarGANv2 as the teacher model, and designs a series of loss functions to use the trained teacher model to guide the student model. In this way, while maintaining a high generation quality, the student model can reduce the number of model parameters and computational complexity, thus achieving the lightweight goal.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A lightweight face image generation method based on improved StarGAN and knowledge distillation, the method comprising the following steps:

[0008] Step 1. Pre-train the teacher model, and the teacher model uses the StarGAN v2 model;

[0009] Step 2. Construct the student model and then perform knowledge distillation; wherein, the student model is obtained by improvement on the basis of the teacher model. In the student model, depthwise separable convolutions are used to replace the standard convolution operations in the downsampling and upsampling of the teacher model, and movable residual blocks are used to replace the residual blocks in the downsampling and upsampling of the teacher model. Moreover, the number of movable residual blocks in the downsampling and upsampling of the student model is set to 2, reducing the number of convolutional layers and the number of channels in each layer. The max_conv_dim in the student model is 128. In addition, when designing the student model, a movable mapping network is added after the mapping network, and multiple fully connected layers are designed in the middle of the two. The nonlinear representation ability is enhanced through the ReLU activation function. The hidden layer dimension of the movable mapping network is max_hid_dim = 128, and the hidden layer dimension of the mapping network is max_hid_dim = 512;

[0010] When performing knowledge distillation, first set the loss function, and then input the training data into the teacher model and the student model for knowledge distillation. During the knowledge distillation process, obtain the activation values of each layer of the generator during forward propagation, use the characteristics output of the pre-trained teacher model as the soft label of the student model, and minimize the loss function, so that the student model not only has to learn to generate similar output images from the input images, but also learn the intermediate layer representations of each layer of the teacher model, helping the student model to be closer to the feature representation of the teacher model in the feature space, and continuously optimize and improve the performance of the student model itself;

[0011] Among them, the loss function includes pixel-level MSE loss , perceptual loss , triplet loss , GKA loss , feature-based wavelet loss ;

[0012] ;

[0013] Among them, and represent the output images of the student and teacher generators respectively, P is the number of pixels of the image;

[0014]

[0015] Among them, and represent the feature extraction networks of the teacher model and the student model respectively, N is the number of image samples, that is, the number of images in the batch of images; means extracting the high-level features in using the feature extraction network of the teacher model, means extracting the high-level features in using the feature extraction network of the student model, represents the feature difference between the image generated by the student model and the image generated by the teacher model under the feature extraction network of the student model;

[0016] ;

[0017] Among them, represents the feature extraction function, is the real image, Represents the margin of the triplet loss. In the setting of the triplet loss, each training sample consists of a triplet: (Anchor, Positive, Negative). The Anchor is the anchor point, which is a sample in the input data. The Positive is a sample that belongs to the same category as the anchor sample or has similar features. The Negative is a sample that belongs to a different category from the anchor sample or has different features;

[0018] ;

[0019] Among them, X and Y are the outputs of the student and teacher models at a certain layer respectively, N is the number of image samples, and T represents the transpose matrix;

[0020] ;

[0021] Among them, represents the feature extraction function, are the features extracted from the i th images generated by the teacher model and the student model respectively, which are calculated through the feature extraction networks of the teacher model and the student model; is the wavelet loss, is the GKA loss;

[0022]

[0023] Among them, and represent the i th high-frequency details after the wavelet transform of the teacher and student images respectively;

[0024] Step 3. Deploy the trained student model to the actual application scenario for face image generation.

[0025] As a preference of the present invention, when pre-training the teacher model in Step 1, it is necessary to preprocess the original image. The specific steps are as follows: Scale the original image to unify the pixels; then perform data normalization to stabilize the distribution; then crop different regions of the image and scale the cropped regions; finally, randomly flip the image horizontally, swap the left and right pixels during flipping, and then map each pixel value in the image from the integer range of [0, 255] to the floating range of [0.0, 1.0] for normalization processing to adapt it to the training model of the teacher network.

[0026] As a preference of the present invention, the feature extraction networks of the teacher model and the student model adopt the trained VGG network.

[0027] As a further preferred embodiment of the present invention, when cropping an image, it is allowed to crop the image randomly during training, and the cropping probability prob=0.5 is set. If the generated random number is less than the cropping probability, cropping is performed, otherwise cropping is not performed and the original image is returned.

[0028] Advantages and beneficial effects of the present invention:

[0029] (1) The present invention provides a new method for generating facial images based on improved StarGAN and knowledge distillation. The model is distilled on the basis of StarGAN v2. In addition, in order to ensure that the quality of the student model obtained by distillation is higher and the model is more lightweight, the present invention designs a series of loss functions to train the generator. Finally, based on the teacher model, not only the excellent performance of the model is learned, but also the purpose of being more lightweight is achieved.

[0030] (2) When constructing the student model, the present invention replaces the convolution layer of the teacher model and uses depthwise separable convolution instead of traditional convolution to achieve the characteristics of lightweight and fast. In addition, the present invention replaces the traditional modules with more flexible MoblieResBlk, MobileMapping Network, and MoblieStyleEcoder, and designs a series of loss functions, so that the trained student model can reduce the number of model parameters, improve computational efficiency, and enhance robustness while retaining the image generation quality.

[0031] (3) The present invention adds pixel-level MSE loss, perceptual loss, triplet loss, GKA loss, and wavelet loss to the generator and discriminator. The design of the loss function enables the generative model to optimize image quality at multiple levels, including details, texture, and style consistency, thereby improving the authenticity and diversity of the generated facial images.

[0032] (4) By using knowledge distillation, the student model of the present invention not only learns the advantages of the teacher model, achieving the excellent generation technology of the teacher model and the purpose of lightweight, but also enhances the generalization ability of the model by designing a sophisticated loss function and network architecture, making it have stronger generalization ability.

[0033] (5) The present invention has strong adaptability and can be flexibly applied to different types of generation tasks. Due to its lightweight design, the model can perform transfer learning on a variety of different tasks, thereby better adapting to various data distributions. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a flow chart of raw image pre-data processing;

[0035] Figure 2It is a schematic diagram of the student model network structure improved based on the teacher model network structure;

[0036] Figure 3 It is a knowledge distillation flow chart;

[0037] Figure 4 It is a knowledge distillation model diagram;

[0038] Figure 5 It is a partial face image generated by the teacher model (not a real face image);

[0039] Figure 6 It is a partial face image generated by the trained student model (not a real face image). Specific implementation manner

[0040] To enable those skilled in the art to better understand the technical solutions and advantages of the present invention, the present application will be described in detail below with reference to the accompanying drawings, but it is not intended to limit the protection scope of the present invention.

[0041] As Figures 1 to 4 shown, this embodiment provides a lightweight face image generation method based on improved StarGAN and knowledge distillation. This method mainly has the following three stages:

[0042] The first stage: Image data preprocessing, model training, and model evaluation of the teacher model (the existing StarGAN v2), and finally adversarial learning produces the characteristics of the generator and discriminator models.

[0043] The second stage: Construct a student model (improved StarGAN v2), use a lightweight strategy to reduce and replace some layers of the teacher model to achieve the purpose of lightweight; then perform knowledge distillation, set new loss functions (pixel-level MSE loss, perceptual loss, triplet loss, GKA loss, wavelet loss); input the training data into the teacher model and the student model, and the characteristics output of the teacher model in the first stage is used as the soft label of the student model, and learn based on the teacher model. During the learning process, continuously optimize and improve the performance of the student model itself to ensure the quality of the generated images.

[0044] The third stage: Evaluate the performance of the student model and conduct a comparative experiment with the teacher model; then deploy the trained student model to the actual application scenario for face image generation.

[0045] In this embodiment, in the first stage, the existing StarGAN v2 model is directly selected as the teacher model, and the CelebA-HQ dataset is used to train the teacher model. The dataset contains 30,000 high-resolution (1024*1024) face images, and each image is equipped with 40 binary attribute labels; in the preprocessing stage of the original image data, it mainly includes image scaling, data normalization, data augmentation, etc.

[0046] Specifically, as Figure 1 shown, in this embodiment, the preprocessing steps of the original image data are as follows:

[0047] The original image is scaled to unify the pixels; then data normalization is performed to stabilize the distribution; finally, in order to enhance the robustness of the teacher model and increase the diversity of the dataset, different regions of the image are cropped, and the cropped regions are scaled to simulate different shooting scenarios or perspective changes of the target.

[0048] The area of the cropped region is ;

[0049] where is the scaling factor, H and W are the length and width of the input image respectively, H⋅W is the area of the input image, the width and height of the cropped region and the ratio are:

[0050] ;

[0051] ;

[0052] ;

[0053] where .

[0054] In this embodiment, a custom lambda function is applied for the cropping operation during cropping, that is: whether to crop the image is determined according to a certain probability to help the model better adapt to different image inputs, thereby improving the training effect; specifically, if the generated random number is less than the cropping probability prob (prob = 0.5), then cropping is performed, otherwise no cropping is performed and the original image is returned.

[0055] Furthermore, in this embodiment, as a conditional cropping operation, it is wrapped into transforms.Lambda so that it can be added to the transform as part of the preprocessing step. It allows randomly cropping images during training, enhancing the diversity of the data, and thus improving the generalization ability of the model.

[0056] Next, the image is randomly flipped horizontally, and the left and right pixels are exchanged during the flipping. Then, each pixel value in the image is mapped from the integer range of [0, 255] to the floating range of [0.0, 1.0] for normalization to adapt to the training model of the teacher network. The formula is as follows:

[0057] ;

[0058] Among them, , , is the pixel value of the original image, is the converted Tensor value; finally, each channel of the image is standardized. The purpose is to make the image pixel values conform to a certain mean and standard deviation. Assume that the mean of the k-th channel of the input image is , and the standard deviation is . The formula for the standardized pixel is:

[0059]

[0060] Among them, is the pixel value after being converted into a Tensor, is the standardized pixel value.

[0061] After preprocessing the original image data, in this embodiment, the processed picture is input into the teacher model for training. The size of the input image is (batch_size, 3, H, W). The target domain label (Target Domain Label) is usually a multi-hot encoded vector, representing the target domain (different facial styles). During the embedding process of the target domain label, the target domain label maps the multi-hot encoded vector to a low-dimensional space through an Embedding Layer (fully connected layer) to reduce complexity. The Embeding Layer maps each class to a dense low-dimensional vector, representing the information between classes more efficiently and capturing the "semantic" similarity between the corresponding vectors of different data pictures; then the embedded domain label is fused with the input image (occurring in the initial layer of the generator) to form the combined information between the input image and the target domain label, and this information is input into the generator of the teacher model. The generator will use the fused latent vector (containing domain label information) to generate a conditional image. In the network structure of the generator, the latent vector (the result after being embedded and concatenated with the domain label) will be passed as input to the initial convolutional layer, using the initial convolutional layer to map the input image and the latent vector to a higher dimension, and then performing downsampling (encoding), upsampling (decoding), and finally passing through a convolutional layer to map the number of channels back to the RGB image to generate the image.

[0062] Specifically, after the data stream enters the generator of the teacher model, it will pass through 4 residual blocks (ResidualBlock). Each residual block includes convolution, activation, and normalization operations. These operations together complete the downsampling operation, whose main purpose is to gradually reduce the spatial resolution and increase the number of feature map channels, which helps to extract higher-level features while maintaining the content information of the image.

[0063] In this embodiment, the formula for the downsampling convolution operation is: ; where is the feature map, is the convolution kernel, is the bias term, is the output of the convolution operation. The convolution kernel w slides on the input image to perform a dot product operation and calculate the output at each position. The stride s controls the sliding pace of the convolution kernel. During the downsampling process, the stride s is set to 1 to reduce the spatial resolution of the output image and ensure that the convolution operation does not cause too much loss of the input size; the padding is set to 1 to control the edges of the input image to maintain the spatial dimension of the output image.

[0064] The activation function mainly uses the LeakyReLU function, aiming to improve the non-linear expression ability of the network. The formula of the activation function is as follows:

[0065]

[0066] where a = 0.2 is the negative slope. When the input value is less than zero, the output value is 0.2 times the input; this operation allows the gradient of a part of the negative values to flow, avoiding the "dead neuron" problem of ReLU and enabling good gradient propagation during training.

[0067] The normalization operation uses adaptive instance normalization (AdaIN), which is used in the generator to adjust the style of the feature map. The input is normalized and adjusted through two operations, namely instance normalization and style transformation. In adaptive instance normalization, s is the style vector, usually generated by another network. After the input original feature image x is processed by instance normalization, the scaling factor (gamma) and offset (bet) are calculated through the style vector s, and finally the normalized features are scaled and translated.

[0068] In this embodiment, after the teacher model completes the downsampling operation, the generator starts the upsampling stage. During upsampling, the network restores the size of the image by gradually increasing the spatial resolution. Each upsampling stage also further extracts features through residual blocks and uses the same structure as the downsampling stage (convolution, activation function, normalization, etc.). The residual blocks use skip connections (also known as residual connections) to add the feature maps that have gone through convolution and normalization to the original input feature maps. The residual connection alleviates the problem of vanishing gradients during the training process.

[0069] Specifically, transposed convolution (Deconvolution) is used for upsampling. This operation learns the convolutional kernel for upsampling. In fact, it is the inverse process of convolution and is used to map low-resolution feature maps to higher-resolution feature maps. Transposed convolution generates a local region around each pixel to restore spatial details. Each time the size of the feature map is expanded by a factor of 2 until the resolution of the target image is reached to achieve the spatial expansion of the feature map. In this way, the generator can maintain the high-level semantic information of the image while restoring the spatial resolution.

[0070] In addition, to enhance the learning ability of the generator, the last layer of the generator usually uses a convolutional layer to map the output feature map back to the pixel space of the target image. The output of this layer is usually the generated image, with a size of (batch_size, 3, H, W), that is, the output image has the same size as the input image but has the style features of the target domain.

[0071] In this embodiment, the teacher model is trained to enable it to generate high-quality face images, and then the student model is designed based on the pre-trained teacher model. The construction strategy of the student model is as follows: depthwise separable convolution is introduced to replace the traditional convolution in the downsampling and upsampling of the teacher model, significantly reducing the number of model parameters and computational complexity; a lightweight network is designed. The student model designs a more lightweight generator student model by reducing the maximum convolution dimension and the hidden layer dimension.

[0072] Specifically, depthwise separable convolution (Depthwise Separable Convolution) is introduced in the student model to replace the standard convolution operations used in the downsampling and upsampling of the teacher model, and the number of parameters and computational complexity in the convolution operation are optimized. In depthwise separable convolution, the convolution operation is divided into two steps: depthwise convolution and pointwise convolution. Assume the input image is , with a size of (height * width * number of input channels), where depth convolution performs convolution operations on each channel of the image separately. Assuming the convolution kernel size is K*K, for each input channel , the output of depth convolution is:

[0073]

[0074] where, is the pixel value at position (h, w) and channel c in the input image x , is the weight of the depth convolution kernel. For each channel c, there is a K*K convolution kernel, is the output after depth convolution, with a size of . The key to this innovation is that it effectively reduces the amount of computation and the number of parameters by decomposing the traditional convolution operation into two steps: depth convolution and pointwise convolution, which is particularly suitable for resource-constrained devices or scenarios with low computing power.

[0075] In addition, in the traditional StarGAN v2 (Star Generative Adversarial Network), the network structure of the generator is complex, and large convolutional layers are used (large residual blocks (ResBlk) are used in the teacher model for style feature extraction). In the student model (Generator Student) after knowledge distillation, mobile residual blocks (MoblieResBlk) are used instead of traditional residual blocks (ResBlk). The mobile style encoder in the student model of this embodiment further reduces the number of convolutional layers and uses smaller channel numbers at the same time, making the process of extracting style information more efficient.

[0076] Specifically, for the student model, this embodiment adjusts the structures of the style encoder (downsampling) and decoder (upsampling), reducing the number of convolutional layers in the network and the number of channels in each layer. The max_conv_dim used in the teacher model is 512, and this parameter is set to 128 in the student model, which means that each convolutional operation in the student model has fewer output channels, effectively reducing the computational complexity and the number of parameters, reaching a lightweight level. Moreover, the number of convolutional layers in the student model is reduced by 2 layers compared to the teacher model, that is, the number of downsampling and upsampling residual blocks in the student model is reduced to 2, aiming to reduce the model complexity and maintain a certain generation ability, improve the computational efficiency of the student model, and reduce the computational and memory consumption.

[0077] ​Finally, in StarGAN v2, the core task of the generator is to map random noise to different styles. To generate stylized images, vectors in the latent space need to be mapped to the style space through the Mapping Network. However, the efficiency and diversity of the style mapping network directly affect the performance of the model. In this embodiment, when designing the student model, a Mobile Mapping Network is added after the Mapping Network, and multiple (specifically 4) fully connected layers (nn.Linear) are designed in between, and the ReLU activation function (nn.ReLU()) is used to enhance the non-linear representation ability. The Mobile Mapping Network uses a smaller hidden layer dimension (max_hid_dim = 128), while the Mapping Network uses a larger hidden layer dimension (max_hid_dim = 512). This optimization reduces the computational cost and memory consumption by restricting the hidden layer dimension.

[0078] After constructing the student model in this embodiment, knowledge distillation is carried out. During the knowledge distillation process, the activation values of each layer in the forward propagation of the generator need to be obtained. The activation values of each layer are the outputs of each layer, and the get_activations method can be used to obtain the activation values of each layer. Specifically, the input image is converted into the form of the model input; then, the input data (such as images and style information) is propagated forward layer by layer through the encode and decode modules; every time a layer is passed through, the activation value of that layer is saved in the lst list for subsequent loss calculation and distillation process. During the distillation process, the activation values of each layer of the teacher model (t_act) and the activation values of each layer of the student model (s_act) are compared through a loss function (such as GKA or perceptual loss), so that the student model not only has to learn to generate similar output images from the input images, but also has to learn the intermediate layer representations of each layer of the teacher model, helping the student model to be closer to the feature representation of the teacher model in the feature space to enhance the generation ability. For example: by minimizing the perceptual loss, the student model will adjust its internal activations (s_act) to make it as close as possible to the activations of the teacher model (t_act). This process usually causes the student model to adjust its weights, prompting it to learn a representation that is more in line with the characteristics of the teacher model.

[0079] To narrow the gap between the teacher model and the student model and enable the student model to be lightweight and flexibly switchable based on the excellent image generation of the teacher model, a series of loss functions are set in this embodiment to narrow the gap between the two.

[0080] Based on the classical adversarial loss of the student model, the mean squared error (MSE) between the outputs of the student generator and the teacher generator incorporates a pixel-level MSE loss. This loss function measures the pixel differences between the outputs of the student generator and the teacher generator. By minimizing the pixel-level MSE, the student model can directly learn pixel-level details from the teacher model. The pixel-level MSE loss has the following mathematical expression:

[0081]

[0082] where and represent the output images of the student and teacher generators respectively. P is the number of pixels in the image, which helps the output of the student model to be as close as possible to that of the teacher model, thus maintaining consistency with the teacher model in terms of generation quality.

[0083] In the case where the student model is small in scale or the training data is limited, through the perceptual loss function, the student model can more effectively learn important information such as the image style and details captured by the teacher model, compensating for the deficiencies in the capabilities of the student model. The perceptual loss has the following mathematical expression:

[0084]

[0085] where and represent the feature extraction networks of the teacher model and the student model respectively. The feature extraction network is an evaluation tool used to extract image features to measure whether the image quality is similar between the teacher and student models. The output metric is the L2 norm, and a trained VGG network can be used; N is the number of image samples, that is, the number of images in the batch of images; represents extracting the high-level features in using the feature extraction network of the teacher model, represents extracting the high-level features in using the feature extraction network of the student model, represents the feature difference between the image generated by the student model and the image generated by the teacher model under the feature extraction network of the student model;

[0086] Since the quality of the generated image and the perceptual effect on the human eye are crucial, the perceptual loss function encourages the image generated by the student model to be more consistent with the output of the teacher model at the perceptual level, improving the quality and naturalness of the image. It emphasizes the high-level visual features of the image rather than just pixel-level similarity. This is particularly important for image generation tasks.

[0087] To ensure the real image and the teacher-generated image The distance between is less than the real image and the student-generated image To help the student model learn better generation ability to improve the quality of the generated image, this embodiment introduces the triplet loss. The triplet loss The mathematical expression is as follows:

[0088]

[0089] Among them, Represents the feature extraction function, Represents the margin of the triplet loss. In the setting of the triplet loss, each training sample consists of a triplet: (Anchor, Positive, Negative), where the anchor (Anchor, e ) is a sample in the input data, usually a reference sample in the positive samples. The positive sample (Positive, p) is a sample belonging to the same category or having similar features as the anchor sample, aiming to let the model learn that the features of similar samples should be closer. The negative sample (Negative, n) is a sample belonging to a different category or having different features from the anchor sample, aiming to let the model learn that the features of dissimilar samples should be far from the anchor.

[0090] In terms of the student model being able to better learn the teacher network, this embodiment adopts a loss function for knowledge distillation. The main goal is to make the student model better imitate the intermediate features of the teacher model through feature comparison and alignment. It combines the methods of feature distillation and global kernel alignment (GKA) to optimize the process of feature transfer. The mathematical expression of the Transfer Features Distillation GKA Loss function is:

[0091]

[0092] Among them, X and Y are the outputs of the student and teacher models at a certain layer respectively, N is the number of image samples, and T represents the transpose matrix. This method emphasizes transferring features by minimizing the L2 (also called Euclidean distance) difference between the layer outputs. It solves the channel difference through a learnable layer and measures the output similarity through global kernel alignment (GKA), which can be directly compared.

[0093] To enhance the details in the images generated by students, minimize the differences in high-frequency details, and address the limitations of traditional GANs, this embodiment introduces the Discrete Wavelet Transform Loss (DWT). Wavelet transform is a signal processing technique that decomposes a signal (the image in this invention) into components of different frequencies, allowing the analysis of the characteristics of the signal at different scales (image resolutions). The main operations are as follows: First, perform wavelet decomposition. Decompose the input images (the generated image and the real image) into multiple subbands through discrete wavelet transform, usually including low-frequency (LL) and high-frequency (LH, HL, HH) components. The low-frequency subband contains the overall structural information of the image, and the high-frequency subbands contain detail information such as edges and textures. Second, calculate the loss. For the generated image and the real image, calculate the differences between the high-frequency parts after their wavelet transforms respectively. These high-frequency components have an important impact on the details and edges of the image. By minimizing the differences in these high-frequency parts, the generator can retain the details of the image without simply generating a smooth and rough structure. Finally, perform optimization. In this way, during the training process, the generator can focus on improving the detail part, reducing the differences in high-frequency features such as textures and edges between the generated image and the real image, and thus generating more realistic and detailed images. The mathematical expression of the Discrete Wavelet Transform (DWT) is as follows:

[0094]

[0095] Wherein, and respectively represent the i th high-frequency details after wavelet transform of the teacher and student images.

[0096] This embodiment proposes a Wavelet-based Feature Loss, which combines the extraction of intermediate features between frequency components with the GKA metric of high-frequency component similarity. By applying the Discrete Wavelet Transform (DWT) to the generated image and the intermediate feature map, the high-frequency components can be effectively compared, adapting to the differences in the number of channels between models. The mathematical expression of the Wavelet-based Feature Loss is as follows:

[0097] ;

[0098] Wherein, represents the feature extraction function, are the features extracted from the i th images generated by the teacher model and the student model respectively, which are calculated through the feature extraction networks (VGG networks) of the teacher model and the student model; is the wavelet loss, is the GKA loss;

[0099] During the knowledge distillation process, the student model is continuously optimized by minimizing the above loss function. After that, the trained student model is deployed to the actual application scenario, and face images can be generated using the student model.

[0100] In this embodiment, in order to evaluate the correlation between the teacher model and the student model, the traditional Frechet Inception Distance (FID) and Learned Perceptual Image Patch Similarity (LPIPS) are used to evaluate the gap between the two models. Frechet Inception Distance (FID) is used to measure the gap between the generated images and the real images. It is based on the feature space of the Inception network and compares by calculating the distance of the statistical characteristics (mean and covariance) of the generated images and the real images. More specifically, FID calculates the Frechet distance between the feature distributions of two sets of images. The mathematical expression of FID is as follows:

[0101]

[0102] Where, and are the means of the features of the real images and the generated images respectively, and are the covariances between the real images and the generated images respectively, represents the trace of the matrix.

[0103] Learned Perceptual Image Patch Similarity (LPIPS) is not just pixel-level similarity. It trains a deep network to learn the perceptual differences between images, and this network measures the similarity of images by comparing the differences in image features. LPIPS simulates the perceptual process of the human visual system, so it has more advantages than traditional pixel-level metrics (such as MSE or PSNR) when evaluating the subjective quality of images. The lower the LPIPS value, the more similar the two images are perceptually. The specific evaluation results of the student model and the teacher model in this embodiment are shown in Table 1.

[0104] Table 1 Evaluation Results of the Student Model and the Teacher Model

[0105] Evaluation Index Reference Teacher Student Trend FID 23.84±0.03 13.73 11.51 2.22 LPIPS 0.39 0.45 0.418 0.032 Elapsed Time \ 49 min 40.5 min 8.5 min

[0106] In this embodiment, Figure 5 shows a sample image generated by the teacher model, Figure 6It shows a sample image generated by the student model. From the perspective of observation, the student model perfectly possesses the excellent generation performance of the teacher model, and the colors are more vivid, the texture is clear, the saturation is relatively high, and the changes in the face images are more obvious. From the experimental data, the FID score (Frechet distance) indicates the intersection situation between the generated Fake image and the Real image, that is, the larger the FID, the closer the Fake image is to the Real image, and the smaller the face diversity; while the LPIPS parameter explains the perceptual similarity level of the generated images. The lower the LPIPS, the higher the perceptual similarity of the images and the better the quality. The student model decreased by 2.22 in the FID score comparison and decreased by 0.032 in the LPIPS comparison, and the time was shortened by 8.5 minutes, as shown in Table 1. To sum up, it can be seen that the student model trained by the present invention not only achieves the excellent generation technology of the teacher model but also achieves the purpose of lightweight, and it has certain advantages over the teacher model in terms of the quality and speed of generating images.

[0107] The above uses specific examples to elaborate on the present invention, which is only used to help understand the present invention and is not intended to limit the present invention. For those skilled in the technical field to which the present invention pertains, based on the idea of the present invention, several simple deductions, deformations or substitutions can also be made. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A lightweight face image generation method based on improved StarGAN and knowledge distillation, characterized in that: The method comprises the following steps: Step 1. Pre-train the teacher model, where the teacher model adopts the StarGAN v2 model; Step 2. Construct a student model, and then perform knowledge distillation; wherein the student model is obtained after improvement on the basis of the teacher model, and the student model uses depthwise separable convolution to replace the standard convolution operation in downsampling and upsampling in the teacher model, and uses movable residual blocks to replace the residual blocks of downsampling and upsampling in the teacher model, and the movable residual blocks of downsampling and upsampling in the student model are set to 2, which reduces the number of convolutional layers and the number of channels of each layer, and the max_conv_dim in the student model is 128; in addition, when designing the student model, a movable mapping network is added after the mapping network, and multiple fully connected layers are designed between the two, and the nonlinear representation capability is enhanced by the ReLU activation function, the hidden layer dimension of the movable mapping network is max_hid_dim=128, and the hidden layer dimension of the mapping network is max_hid_dim=512; When performing knowledge distillation, first set the loss function, then input the training data into the teacher model and the student model for knowledge distillation. During the knowledge distillation process, obtain the activation values ​​of each layer of the generator in the forward propagation, use the pre-trained teacher model feature output as the soft label of the student model, minimize the loss function, so that the student model not only learns to generate similar output images from the input image, but also learns the intermediate layer representations of each layer of the teacher model, helping the student model to be closer to the feature representation of the teacher model in the feature space, and continuously optimizing and improving the performance of the student model itself; Among them, the loss function includes pixel-level MSE loss , Perceptual Loss , triplet loss , GKA loss , feature-based wavelet loss ; ; in, and denote the output images of the student and teacher generators respectively, P is the number of pixels of the image; ; in, and Represent the feature extraction networks of the teacher model and the student model respectively, N is the number of image samples, that is, the number of images in the batch image; Represents the feature extraction network extracted using the teacher model The high-level features in Represents the feature extraction network extracted using the student model The high-level features in The feature difference between the image generated by the student model and the image generated by the teacher model under the student model feature extraction network; ; in, represents the feature extraction function, is a real image, Represents the margin of triplet loss. In the setting of triplet loss, each training sample consists of a triplet: (Anchor, Positive, Negative), where Anchor is an anchor point, which is a sample in the input data. Positive is a sample that belongs to the same category or has similar features as the anchor point sample. Negative is a sample that belongs to a different category or has different features from the anchor point sample. ; Among them, X and Y are the outputs of the student and teacher models at a certain layer, respectively. N is the number of image samples, T represents the transposed matrix; ; in, represents the feature extraction function, are generated by the teacher model and the student model. i The features extracted from the image are calculated by the feature extraction network of the teacher model and the student model; is the wavelet loss, is the GKA loss; ; in, and Represents the teacher and student images after wavelet transformation. i High-frequency details; Step 3. Deploy the trained student model to the actual application scenario to generate face images.

2. According to claim 1, a lightweight face image generation method based on improved StarGAN and knowledge distillation is characterized in that: Step 1: When pre-training the teacher model, the original image needs to be pre-processed. The specific steps are: scale the original image to unify the pixels; then normalize the data to stabilize the distribution; then crop different areas of the image and scale the cropped areas; finally, randomly flip the image horizontally, swap the left and right pixels during flipping, and then map each pixel value in the image from the integer range of [0, 255] to the floating range of [0.0, 1.0], and normalize it to make it suitable for the training model of the teacher network.

3. The lightweight face image generation method based on improved StarGAN and knowledge distillation according to claim 1, characterized in that: The feature extraction network of the teacher model and the student model uses the trained VGG network.

4. The lightweight face image generation method based on improved StarGAN and knowledge distillation according to claim 2, characterized in that: When cropping images, it is allowed to crop images randomly during training, and the cropping probability prob=0.5 is set. If the generated random number is less than the cropping probability, cropping is performed, otherwise no cropping is performed and the original image is returned.

Citation Information

Patent Citations

  • Zero sample knowledge distillation method and system based on resisting triple loss

    CN114972904A

  • Personalized human body action recognition method based on knowledge distillation

    CN116844225A