An Image Translation Model Compression Method Based on Knowledge Distillation

By introducing the MGFD framework and multi-particle distillation scheme, the student generator is optimized, and the high computational cost and overfitting problems of image translation models are solved, achieving efficient image translation model compression and high-quality image generation.

CN116188863BActive Publication Date: 2025-07-22SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310197492.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-03
Publication Date
2025-07-22
Estimated Expiration
2043-03-03

AI Technical Summary

Technical Problem

The existing image translation model compression technology has the problems of high computational cost, excessive memory usage, complex features and structures not being fully considered, and multi-stage tasks leading to overfitting of students' networks, and lacks a simple and effective model compression framework.

Method used

Using an end-to-end framework MGFD with more generators and fewer discriminators, combining a multi-particle distillation scheme with intermediate layer distillation loss, optimized student generators through structural similarity loss and perceived loss, abandoning multi-stage steps, and optimizing network parameters with dynamic learning rates and adversarial losses.

Benefits of technology

It effectively reduces the complexity of model compression, improves image translation performance, solves the overfitting problem, and realizes high-quality image generation while reducing computational costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116188863B_ABST
    Figure CN116188863B_ABST
Patent Text Reader

Abstract

The present invention proposes an image translation model compression method based on knowledge distillation, belonging to the technical fields of deep learning and computer vision. Generative adversarial networks have achieved excellent performance in image generation tasks. However, due to high computational costs and excessive memory occupancy, their applications are greatly limited. The present invention first introduces a novel network module, achieving good results by adopting the combination of residual connections and traditional architectures. Secondly, in order to reduce the complexity of the model, a framework with more generators and fewer discriminators (MGFD) is proposed. In addition, a multi-granularity distillation scheme and an intermediate layer distillation loss are utilized to improve image quality. The present invention conducts comparative experiments with classical algorithms and the current state-of-the-art algorithms on public datasets, and the experiments prove that the present invention can reduce computational costs while generating high-quality images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of deep learning and computer vision, and particularly relates to an image translation model compression technology.

[0002] Research Background

[0003] As a common method of visual understanding, the key to image translation lies in learning a mapping relationship that can transform between image domains, using a neural network to learn the content of the source domain, and then converting it to the target image domain space. Many problems in human production and life can be transformed into subtasks of image translation. For example, in the field of autonomous driving, converting the street view images captured by in-vehicle cameras into target segmentation maps; in remote sensing monitoring, converting real-scene maps into simple map modes; in entertainment and leisure, people hope to achieve interesting goals such as face cartoonization and old photo restoration.

[0004] Looking back on the past few years, Generative Adversarial Networks (GANs) have made great progress in image synthesis, image-to-image conversion, and generating high-quality, high-resolution high-definition images through the adversarial idea. These technologies have been widely used in commercial editing software. With the birth of generative adversarial networks, more methods with small network scales and good effects have emerged in the image translation task, and the translation work based on this network has become the mainstream of research. However, their applications are severely restricted due to high computational costs and excessive memory usage. Although rich results have been achieved in past work on model compression, such as weight pruning, channel slimming, and neural network architecture search. But the above mainstream model compression algorithms still have the following problems:

[0005] 1) They directly use mature model compression technologies to compress GAN models without considering their complex features and structures. The complex features and designs of GANs often lead to these mainstream compression algorithms failing to achieve the expected effects.

[0006] 2) The compression of GANs is usually a multi-stage task. Such a complex multi-stage task often results in the student network not being able to receive timely guidance from the teacher network and being extremely prone to overfitting of the student network.

[0007] 3) Image translation methods based on generative adversarial networks have occupied an important position in the field of image translation. However, there is still no simple and effective model compression framework for image translation methods based on generative adversarial networks.

[0008] The above points are all problems that urgently need to be solved for the compression of image translation models. Summary of the Invention

[0009] To solve the above problems, the present invention proposes an end-to-end framework with more generators and fewer discriminators (MGFD) to effectively learn GAN. This work focuses on the compression of image translation models, such as CycleGAN. Specifically, the present invention introduces a new network architecture design, which can be used as a teacher network design and a student network architecture. In addition, the present invention also proposes a new distillation strategy that abandons the complex multi-stage compression steps and obtains a compressed model in just one step. Finally, the present invention adopts a multi-granularity distillation scheme and an intermediate layer distillation loss to mine potential information to help better train the model. The technical solutions adopted by the present invention are as follows:

[0010] A method for compressing an image translation model based on knowledge distillation, the method comprising the following steps:

[0011] (1) Prepare a general image translation dataset, randomly select an image from the image translation dataset of the source domain X, and send it to the image translation model compression framework MGFD, which includes a widened teacher generator, a teacher generator, a student generator, and a parameter-sharing discriminator;

[0012] (2) The selected image is respectively generated by the widened teacher generator, the teacher generator and the student generator to generate the corresponding generated image of the target domain Y, where the widened teacher generator and the student generator are optimized by the intermediate layer distillation loss;

[0013] (3) The differences in spatial structure and content style between the output images of the teacher generator and the student generator are measured by structural similarity loss and perceptual loss, thereby optimizing the student generator;

[0014] (4) A multi-granularity distillation mechanism is used to measure the differences in spatial structure and content style between the output images of the widened teacher generator and the student generator through structural similarity loss and perceptual loss, thereby optimizing the student generator;

[0015] (5) The Y-domain image generated by the widened teacher generator and the Y-domain image generated by the teacher generator are sent to the parameter-sharing discriminator for discrimination, and scores between [0, 1] are obtained respectively. The closer the score is to 1, the more realistic the generated Y-domain image is.

[0016] (6) Generative adversarial loss is used to continuously optimize the generator and discriminator. In addition, the generator also uses MSE loss to make the generated images more realistic.

[0017] (7) Through continuous iterations of gradient descent and back propagation, the parameters of the widened teacher generator, teacher generator, student generator, and discriminator networks are updated so that the four networks reach Nash equilibrium.

[0018] Specifically, the widened teacher generator described in step (2) means that each channel in the convolutional layer is multiplied by a channel expansion rate η, and the default value of η is 4.

[0019] Specifically, the discriminator with parameter sharing described in step (5) allows sharing information in the first several layers and obtaining the discriminator outputs of the widened teacher generator and the teacher generator respectively.

[0020] Specifically, during the training process, a learning rate dynamic update strategy is used, that is, different learning rates are used for gradient update, so as to effectively reduce the instability of model training and transfer the information of the best-performing teacher generator to the student generator in real time.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0022] 1. The present invention proposes an end-to-end MGFD framework to effectively learn GAN, introducing a new network architecture design that can not only be designed as a teacher network but also suitable for the student network.

[0023] 2. The MGFD framework abandons complex multi-stage steps and can obtain an optimized student generator in one step without a discriminator, greatly reducing the complexity of model compression.

[0024] 3. A multi-granularity distillation scheme is adopted to obtain more features by using different teacher generators. In addition, an intermediate layer distillation loss is also adopted to obtain more supervision signals. This scheme can not only improve the performance of model image translation from different angles but also effectively solve the overfitting problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a schematic flow chart of the present invention.

[0026] Figure 2 is a model framework diagram proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0027] Referring to Figure 1 and Figure 2 , the present invention provides an end-to-end framework with more generators and fewer discriminators (MGFD) to effectively learn GAN. The method of the present invention will be described in detail below with reference to the drawings and examples, and the method includes the following steps:

[0028] S1: Take an image from the real dataset to be converted in the source domain X, crop it to a resolution of 256*256, and then send it into the widened teacher generator, the teacher generator, and the student generator network respectively.

[0029] S2: As Figure 2As shown, the images fed into the widened teacher generator, teacher generator, and student generator networks undergo convolution and deconvolution to extract features. During the process of feature extraction by the widened teacher generator and student generator, the intermediate layer distillation loss is used to transmit the channel granularity information of the widened teacher generator to the student generator as an additional supervision signal. The weight formula for each channel is as follows:

[0030]

[0031] where w c is the weight of channel c th , H and W are the dimensions of the feature map space, and u c (i,j) is the activation value. The present invention uses a 1×1 convolutional layer to increase the dimension and expands the number of channels by connecting this convolutional layer to the intermediate layer of the student generator. The specific form of the intermediate layer distillation loss is as follows:

[0032]

[0033] where w ij is the channel weight of the j th -th layer feature map. c is the number of channels of the feature map. n is the number of feature maps to be sampled.

[0034] S3: Convert the real image of the source domain X into a generated image of the target domain Y through the widened teacher generator, teacher generator, and student generator networks.

[0035] S4: In the MGFD framework proposed by the present invention, only the teacher network is optimized while the discriminator is discarded for the student generator. Correspondingly, when optimizing the student generator, it only needs to make the student generator learn the output of the teacher generator. By continuously optimizing the student generator and teacher generator through backpropagation, the student generator can gradually learn to imitate the teacher generator. The present invention represents the output of the student generator as G S (x), and represents the output of the teacher generator as G T (x). The difference between the outputs of the teacher generator and student generator is measured by the structural similarity loss and perceptual loss to optimize the student generator.

[0036] S5: In addition to the teacher generator and student generator, the present invention also adopts a multi-granularity distillation scheme, using the output of a wider teacher generator to help the student network learn from multiple aspects and improve the performance of the model for image translation from different angles. The multi-granularity distillation scheme is unique to the present invention and is as follows:

[0037] Most of the previous knowledge distillation methods are aimed at training the student network to obtain effective but single knowledge, ignoring the different abilities of the student network to understand knowledge. In human real life, experienced teachers will refine and summarize knowledge, analyze and generalize knowledge from multiple granularities, so that students can understand and remember it comprehensively. The same is true for neural networks. The performance of a student network trained by a powerful teacher network is often not as good as that of a student network jointly trained by multiple teacher networks. Different from the previous transfer of knowledge from a single teacher network to a student network, the multi-granularity distillation mechanism aims to explore the multi-granularity of the teacher network's knowledge, used to transfer comprehensive knowledge and help the student network learn from multiple aspects. The teacher generator and the student generator are optimized using the structural similarity loss L SSIM and the perceptual loss and to optimize, and the final distillation objective is:

[0038]

[0039] where λ1, λ2, and λ3 are the balance parameters of the loss function.

[0040] Specifically, the present invention uses the widened teacher generator and the teacher generator together to optimize the output of the student generator and continuously update the parameters. The widened generator means that each channel in the convolutional layer is multiplied by a channel expansion rate η. The multi-granularity distillation loss is as follows:

[0041]

[0042] S6: Send the images of the target domain Y generated after converting the widened teacher generator and the teacher generator into the shared discriminator network for discrimination. The discriminator network performs convolution and deconvolution operations on the input images to extract features.

[0043] S7: The feature maps obtained by the discriminator are passed through max pooling and passed through the sigmoid activation function to obtain a discrimination result in the range of [0,1].

[0044] S8: Calculate the generative adversarial network error using the adversarial loss and the MSE loss. Through the optimization methods of backpropagation and gradient descent, the network converges faster and reaches the Nash equilibrium. Among them, the present invention applies the dual-scale and dynamic learning rate update strategy, effectively reducing the instability of model training. The specific method is as follows:

[0045] Dual-scale and dynamic learning rate update strategy

[0046] During the training process of the generative adversarial network, the generator and the discriminator, as two independent networks, continuously compete and update until the model converges. Therefore, each network has its own learning rate during the model training process. In the present invention, the learning rate of the discriminator is fixed at 0.0002, while the learning rate of the generator is first set to 0.0002 and remains unchanged in the first 100 training epochs, and then linearly decays to 0 in the next 100 training epochs. By this method, the instability in model training is effectively reduced, and the regularization process of the discriminator is accelerated.

[0047] Therefore, the final loss function of the present invention can be expressed as:

[0048]

[0049] Where is the adversarial loss and the MSE loss, is the generator distillation and multi-scale distillation loss, is the intermediate layer distillation loss, and λ CD is a hyperparameter.

[0050] S9: To verify the accuracy and generalization of the present invention, the present invention is compared with existing algorithms on two small datasets used in CycleGAN and two small datasets used in Pix2Pix. The FID and mIoU are used to evaluate the model performance. FID represents the distance between the feature vectors of the generated image and the real image. The smaller this metric is, the smaller the distance between the feature vectors of the generated image and the real image, indicating that the performance of the generative model is better. The larger the mIoU, the more the predicted segmentation region coincides with the real label region, and the better the generative ability of the model. The data shown in Tables 1 and 2 are obtained.

[0051] As shown in Table 1, the proposed model consumes the least number of multiply-accumulate operations (MAC). Compared with the state-of-the-art GCC, the computational complexity of the proposed method only accounts for 58.8% of its computational complexity. The proposed method has the least computational complexity and the best performance on the included datasets. For example, on the horse2zebra dataset, the FID of the proposed method is also the lowest among the methods in recent years, decreasing from 61.53 to 54.91. For the summer2winter dataset, the FID of the proposed method also decreases by 5.58 compared with the original model.

[0052] Table 1

[0053]

[0054]

[0055] Table 2

[0056]

[0057] Meanwhile, in order to prove the generalization of the present invention, a model compression experiment was also conducted on the Pix2Pix method. As shown in Table 2, for the Pix2Pix model, the method of the present invention has the best compression rate, and the FID and mIoU of the generated images are also similar to those of the best methods in recent years. On the cityscapes dataset, compared with the state-of-the-art GCC, the mIoU index of the method of the present invention is comparable, but the computational cost is only 45.6% of it. For the edges2shoes dataset, the FID of the images generated by the method of the present invention is similar to that of the state-of-the-art DMAD, but the computational cost is only 32.8% of it. All these indicate that the method of the present invention is a beneficial improvement to the image compression model. Using the framework of the present invention, the computational complexity of the model and the image generation performance can be well balanced.

[0058] In summary, the present invention application provides an image translation model compression method based on knowledge distillation. The innovations of this method are as follows: First, an end-to-end MGFD framework is created to effectively learn GAN, and a new network architecture design is introduced, which can not only be designed as a teacher network but also suitable for the student network. Second, the MGFD framework abandons the complex multi-stage steps and can obtain an optimized student generator in one step without a discriminator. Third, a multi-granularity distillation scheme is adopted to obtain more features by using different teacher generators. In addition, an intermediate layer distillation loss is also adopted to obtain more supervision signals.

Claims

1. A method for compressing an image translation model based on knowledge distillation, characterized in that, The method includes the following steps: Step 1: Prepare a general image translation dataset. Randomly select an image from the image translation dataset in the source domain X and feed it into the image translation model compression framework MGFD. The MGFD framework includes a widened teacher generator, a teacher generator, a student generator, and a discriminator with shared parameters; Step 2: The selected image passes through the widened teacher generator, the teacher generator, and the student generator respectively to generate corresponding target domain Y generated images. Among them, the widened teacher generator and the student generator are optimized through the intermediate layer distillation loss; Step 3: Measure the differences in the spatial structure and content style of the output images of the teacher generator and the student generator through the structural similarity loss and the perceptual loss, so as to optimize the student generator; Step 4: Adopt a multi-granularity distillation mechanism to measure the differences in the spatial structure and content style of the output images of the widened teacher generator and the student generator through the structural similarity loss and the perceptual loss, so as to optimize the student generator; Step 5: Feed the Y-domain images generated by the widened teacher generator and the Y-domain images generated by the teacher generator into the discriminator with shared parameters for discrimination, and obtain scores between [0, 1] respectively. The closer the score is to 1, the more real the generated Y-domain image represents; Step 6: Use the generative adversarial loss to continuously optimize the generator and the discriminator. In addition, the generator also uses the MSE loss to make the generated images more real; Step 7: Continuously iterate through gradient descent and backpropagation to update the network parameters of the widened teacher generator, the teacher generator, the student generator, and the discriminator, so that the four networks reach the Nash equilibrium.

2. The method for compressing an image translation model based on knowledge distillation according to claim 1, wherein The widened generator mentioned in Step 2 means that each channel in the convolutional layer is multiplied by a channel expansion rate η, and the default value of η is 4.

3. The method for compressing an image translation model based on knowledge distillation according to claim 1, wherein The discriminator with shared parameters mentioned in Step 5 allows sharing information in the first several layers and obtaining the discriminator outputs of the widened teacher generator and the teacher generator respectively.

4. The method for compressing an image translation model based on knowledge distillation according to claim 1, wherein During the training process, use a learning rate dynamic update strategy, that is, use different learning rates for gradient update, so as to effectively reduce the instability of model training and transfer the information of the best-performing teacher generator to the student generator in real time.