Model Generation Method, Device, Equipment and Storage Medium Based on Knowledge Distillation Free

The model is generated by the knowledge-free distillation method, and the second model is trained using the image generator and the target loss function, which solves the problems of low accuracy caused by strong data privacy and low sample data, and improves the application accuracy of the model.

CN116306822BActive Publication Date: 2025-07-18PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310393441.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-06
Publication Date
2025-07-18
Estimated Expiration
2043-04-06

AI Technical Summary

Technical Problem

In the case of strong data privacy and small sample data, existing model compression methods cannot effectively improve the accuracy of student networks when applied.

Method used

By generating sample images using a preset image generator, the target loss function is constructed using the knowledgeless distillation method, and the second model is trained to improve accuracy, including generating M sample images, calculating sample weights, constructing the target loss function and adjusting the initial parameters.

Benefits of technology

It improves the accuracy of the generative model when applied, solves the problem of uneven feature distribution caused by small data samples, and enhances the performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116306822B_ABST
    Figure CN116306822B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and in particular to a model generation method, device, equipment and storage medium based on knowledge distillation-free. By using a preset image generator, M sample images are generated, and the M sample images are input into a first model and a second model to obtain a first feature matrix and a second feature matrix. Through a preset kernel function, the second feature matrix is feature decoupled to obtain a multi-feature second feature matrix. The sample weights of each row in the multi-feature second feature matrix are calculated, and according to the sample weights, a target loss function is constructed to train the second model to generate a target model, which solves the problem of fewer data samples. According to the sample weights in each sample image and the difference in the output features of the first model and the second model, a loss function corresponding to the second model is constructed, and the sample weights are calculated, avoiding the problem of uneven distribution of the features of the generated sample images, and improving the accuracy of the generated model in application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a model generation method, device, equipment and storage medium based on knowledge-free distillation. Background Art

[0002] Knowledge distillation, as a model compression method, is currently widely used. Knowledge distillation regards the model to be compressed as the "teacher" and the compressed model as the "student". The teacher network has strong capabilities but a complex structure, making it inconvenient to deploy; the student network has a simple structure, but the directly trained effect is not good. Knowledge distillation is to improve the performance of the student network during application by means of the teacher network assisting the training of the student network, achieving an effect close to that of the teacher network.

[0003] When performing model compression, if the training data can be directly accessed, most of the existing deep neural network compression and acceleration methods are very effective. However, if the training data is inaccessible due to privacy or legal reasons, most of the model compression methods will fail. In the scenario of continual learning, a generative model can be trained as a generator of fake data for old knowledge, and the fake data is mixed with new data and then used to train the student network. Due to the distribution shift of the generated fake data, the accuracy of the trained student network during application is relatively low. Therefore, when the data privacy is strong and the sample data is scarce, how to improve the accuracy of the student model during application has become an urgent problem to be solved. Summary of the Invention

[0004] Based on this, it is necessary to provide a model generation method, device, equipment and storage medium based on knowledge-free distillation for the above technical problems to solve the problem of low accuracy of the model during application.

[0005] The first aspect of the embodiments of the present application provides a model generation method based on knowledge-free distillation, and the method includes:

[0006] Using a preset image generator to generate M sample images, where M is an integer greater than 1;

[0007] Inputting the M sample images into a first model to obtain a first feature tensor of each feature in each sample image, and the first feature tensors of all features in each sample image constitute a first feature matrix corresponding to the sample image;

[0008] Inputting the M sample images into a second model to obtain a second feature tensor of each feature in each sample image, and the second feature tensors of all features in each sample image constitute a second feature matrix corresponding to the sample image;

[0009] By means of a preset kernel function, perform feature decoupling on the second feature matrix to obtain a multi-feature second feature matrix;

[0010] According to the correlation function between each column of features in the multi-feature second feature matrix, calculate the sample weight of each row in the multi-feature second feature matrix;

[0011] According to the sample weight and the difference expression between the first feature matrix and the second feature matrix, construct an objective loss function;

[0012] According to the objective loss function, train the second model to obtain target parameters;

[0013] Use the target parameters to update the initial parameters in the second model to generate a target model.

[0014] The second aspect of the embodiments of the present application provides a model generation device based on knowledge distillation-free, and the device includes:

[0015] A generation module, configured to use a preset image generator to generate M sample images, where M is an integer greater than 1;

[0016] A first feature matrix determination module, configured to input the M sample images into a first model, and obtain a first feature tensor of each feature in each sample image, and the first feature tensors of all features in each sample image constitute the first feature matrix of the corresponding sample image;

[0017] A second feature matrix determination module, configured to input the M sample images into a second model, and obtain a second feature tensor of each feature in each sample image, and the second feature tensors of all features in each sample image constitute the second feature matrix of the corresponding sample image;

[0018] A multi-feature second feature matrix determination module, configured to perform feature decoupling on the second feature matrix by means of a preset kernel function to obtain a multi-feature second feature matrix;

[0019] A sample weight determination module, configured to calculate the sample weight of each row in the multi-feature second feature matrix according to the correlation function between each column of features in the multi-feature second feature matrix;

[0020] An objective loss function construction module, configured to construct an objective loss function according to the sample weight and the difference expression between the first feature matrix and the second feature matrix;

[0021] A target parameter determination module, configured to train the second model according to the objective loss function to obtain target parameters;

[0022] A target model acquisition module, configured to update initial parameters in the second model using the target parameters to generate a target model.

[0023] In a third aspect, an embodiment of the present invention provides a computer device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for generating a model based on knowledge distillation as described in the first aspect is implemented.

[0024] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method for generating a model based on knowledge distillation as described in the first aspect is implemented.

[0025] The beneficial effects of the present invention compared with the prior art are as follows:

[0026] Using a preset image generator to generate M sample images, inputting the M sample images into a first model to output a first feature tensor of each feature in each sample image, and the first feature tensors of all features in each sample image constitute a first feature matrix corresponding to the sample image. M is an integer greater than 1. Inputting the M sample images into a second model to output a second feature tensor of each feature in each sample image, and the second feature tensors of all features in each sample image constitute a second feature matrix corresponding to the sample image. Through a preset kernel function, feature decoupling is performed on the second feature matrix to obtain a multi-feature second feature matrix, and according to the correlation function between each column of features in the multi-feature second feature matrix, the sample weight of each row in the multi-feature second feature matrix is calculated. According to the sample weights and the difference expression of the first feature matrix and the second feature matrix calculated through a preset algorithm, a target loss function is constructed. According to the target loss function, the second model is trained to adjust the initial parameters in the second model to generate a target model, solving the problem of fewer data samples. According to the sample weights in each sample image and the difference between the output features of the first model and the second model, a loss function corresponding to the second model is constructed, and the sample weights are calculated, avoiding the problem of uneven distribution of the generated sample image features and improving the accuracy of the generated model when applied. Description of the Drawings

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other accompanying drawings without creative efforts.

[0028] Figure 1It is a schematic diagram of an application environment of a model generation method based on knowledge distillation according to an embodiment of the present invention;

[0029] Figure 2 It is a schematic flowchart of a model generation method based on knowledge distillation according to an embodiment of the present invention;

[0030] Figure 3 It is a schematic flowchart of a model generation method based on knowledge distillation according to an embodiment of the present invention;

[0031] Figure 4 It is a schematic structural diagram of a model generation device based on knowledge distillation according to an embodiment of the present invention;

[0032] Figure 5 It is a schematic structural diagram of a computer device according to an embodiment of the present invention. Detailed implementation manners

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] It should be understood that when used in the specification and the appended claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0035] It should also be understood that the term "and / or" as used in the specification and the appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0036] As used in the specification and the appended claims of the present invention, the term "if" can be interpreted as "when...", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" according to the context.

[0037] In addition, in the description of the specification and the appended claims of the present invention, the terms "first", "second", "third", etc. are only used for distinguishing descriptions, and cannot be understood as indicating or implying relative importance.

[0038] References to "one embodiment" or "some embodiments" or the like described in the specification of the present invention mean that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants mean "including but not limited to", unless otherwise specifically emphasized.

[0039] Embodiments of the present invention may acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0040] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0041] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution is prior or subsequent. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0042] In order to illustrate the technical solution of the present invention, specific embodiments will be used for illustration below.

[0043] A model generation method based on knowledge distillation-free provided by an embodiment of the present invention can be applied in an application environment such as Figure 1 where the client communicates with the server. Among them, the client includes but is not limited to computer devices such as a palm computer, a desktop computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The server can be implemented by an independent server or a server cluster composed of multiple servers.

[0044] SeeFigure 2 , which is a schematic flowchart of a model generation method based on knowledge-free distillation provided by an embodiment of the present invention. The above-mentioned model generation based on knowledge-free distillation can be applied to Figure 1 the server in Figure 2 . The above-mentioned server is connected to the corresponding client to provide model training services for the client. As

[0045] shown, the model generation method based on knowledge-free distillation may include the following steps.

[0046] In step S201, use a preset image generator to generate M sample images, where M is an integer greater than 1.

[0047] In step S201, use a preset image generator to continuously generate new sample images according to the learned feature distribution of real data. Among them, the new sample images are virtual images similar to real sample images.

[0048] In this embodiment, use a preset image generator to generate M sample images. The preset image generator is a generator that can learn the image features in the preset sample images. In this embodiment, the preset generator is trained by a neural network model with a fully convolutional structure. After inputting the preset sample images into the preset image generator, sample images similar to the preset sample images can be output. The fully convolutional neural network includes a feature extraction network, a feature learning network, and a synthesis network. The feature extraction network is usually the first part of the convolutional neural network model, and its input is the preset sample images, and the output is the preliminarily extracted feature information. The feature learning network is the second part of the convolutional neural network model, and its input is the preliminarily extracted feature information in the feature extraction network, and the output is the learned fine feature information. The synthesis network is the third part of the convolutional neural network model, and its input is the fine feature information, and the output is the sample image. When the feature extraction network in the first part extracts the feature information of the preset sample images, the feature learning network in the second part can obtain the feature information and perform learning and convolutional operation processing. The results of the first part and the second part can be used as the input of the third part, and then the synthesis network in the third part synthesizes the sample images corresponding to the preset sample images.

[0048] S202: Input the M sample images into the first model to obtain the first feature tensors of each feature in each sample image. The first feature tensors of all features in each sample image constitute the first feature matrix corresponding to the sample image.

[0049] In step S202, input the sample images into the first model, and output the first feature tensors of each feature in each sample image, which constitute the first feature matrix corresponding to the sample image. Among them, the first model is a trained deep learning model.

[0050] Input M sample images into the first model. The first model is the teacher model in knowledge distillation. When obtaining the first model that meets the conditions, the sample data in the training set can be manually labeled first, and then the labeled sample data can be used to train the first model. This training process can be understood as a pre-training process. During the training process, the loss function of the first model can be used to calculate the loss value between the actual output of the first model and the labeled result, and backpropagation training can be performed according to the loss value until the loss converges to obtain the first model. For example, the first model is used to extract the region of interest, and the output of the first model can be the coordinate features of the region of interest in the image.

[0051] Input M sample images into the first model, and output the first feature tensors of each feature in each sample image to form the first feature matrix corresponding to the sample image. The first feature matrix includes the feature tensors of M sample images.

[0052] S203: Input the M sample images into the second model, and obtain the second feature tensors of each feature in each sample image. The second feature tensors of all features in each sample image form the second feature matrix corresponding to the sample image.

[0053] In step S203, the second model is an untrained deep learning model. The second model is used to learn the parameters in the first model. Input M sample images into the second model, and output the second feature tensors of each feature in each sample image to form the second feature matrix corresponding to the sample image.

[0054] In this embodiment, the second model is the student model in knowledge distillation. Input M sample images into the second model, and output the second feature tensors of each feature in each sample image. The number of features selected in each sample is the same as the number of features in each sample in the first model. The second feature matrix corresponding to the sample image is formed according to the second feature tensors of each feature in each sample image.

[0055] It should be noted that initial parameters can be set for the second model. For example, when the activation function is the saturation activation function tanh, the Xavier initialization method is used to initialize the parameters in the second model. The Xavier initialization method is beneficial to accelerating convergence and reducing overfitting. When the activation function is ReLU and its variant activation functions, the Kaiming initialization method is used. LeCun initialization is applicable to neural networks with sigmoid activation functions. Its main design idea is to assume that the input of the network follows a Gaussian distribution. During the forward propagation of activation, by controlling the expectation and variance of the sampling distribution of the weight parameters during initialization, the expectation of the activation value of each layer of neurons is 0 and the variance is 1.

[0056] It should be noted that the purpose of initialization is to avoid gradient explosion or gradient disappearance. A higher requirement is to ensure the stability of both the forward propagation of activation and the backward propagation of gradients, that is, to control the means and variances of the neuron activation signals during forward propagation and the error signals during backward propagation to be stable. Therefore, when choosing different initialization methods, try to make the means and variances of the neuron activation signals during forward propagation and the error signals during backward propagation stable.

[0057] S204: Through a preset kernel function, perform feature decoupling on the second feature matrix to obtain a multi-feature second feature matrix.

[0058] In step S2034, the preset kernel function is a non-linear transformation function. Feature decoupling is to extract the multi-scale features of the image background and the shape and appearance of the image foreground in the sample image, map the sample image to a multi-feature space, and obtain the sample weights of the multi-feature second feature matrix, where the correlation is the correlation between any two columns in the multi-feature second feature matrix.

[0059] In this embodiment, the preset kernel function is a random Fourier feature function. The second feature matrix output by the second model is subjected to feature decoupling through the random Fourier feature function. Among them, the random Fourier feature function is as shown in Equation (1):

[0060]

[0061] Among them, x is the feature tensor in the second feature matrix, w follows a Gaussian distribution of N(0,1), follows a uniform distribution of Uni(0, 2π). Using non-linear mapping, the second feature matrix is mapped to a high-dimensional space. When the random Fourier feature function maps, first the second feature matrix is mapped to a randomly selected straight line through a random feature, and then the obtained scalar is transformed through a sine curve to obtain the final result.

[0062] S205: And according to the correlation between each column of features in the multi-feature second feature matrix, calculate the sample weights of each row in the multi-feature second feature matrix.

[0063] In step S205, according to the correlation between each column of features in the multi-feature second feature matrix, calculate each row in the multi-feature second feature matrix.

[0064] In this embodiment, after the second feature matrix is mapped to a high-dimensional space through the random Fourier feature function, a high-dimensional second feature matrix is obtained. The correlation between each dimension in the high-dimensional feature matrix is diluted, and it can be considered that each feature is independent of each other. When the correlation between each column of features in the high-dimensional second matrix is the smallest, the weights of each row in the high-dimensional second matrix are used as the sample weights of each final sample.

[0065] Optionally, according to the correlation function between each column of features in the multi-feature second feature matrix, the sample weights of each row in the multi-feature second feature matrix are calculated, including:

[0066] Construct a correlation function based on the correlation relationship between each column of features in the multi-feature second feature matrix;

[0067] Take the weight with the smallest correlation value of the correlation function as the sample weight. In this embodiment, based on the correlation between each column of features in the multi-feature second feature matrix, a correlation function is constructed. When calculating the correlation, since the sample weight information of each row in the high-dimensional second matrix is related to the feature correlation in each column, when calculating the correlation between each column of features, the sample weight of each row is added to the function for calculating the correlation between each column of features. The function formula of the correlation is as shown in Equation (2):

[0068]

[0069] In Equation (2), Rel AB;w The correlation magnitude between any two columns of features, column A feature and column B feature, in the high-dimensional second matrix, m is the number of rows in the high-dimensional second matrix, w i is the sample weight of the i-th row in the high-dimensional second matrix, w j is the sample weight of the j-th row in the high-dimensional second matrix, A i is the feature tensor of the i-th row in column A, A j is the feature tensor of the j-th row in column A, B i is the feature tensor of the i-th row in column A, B j is the feature tensor of the j-th row in column A, u, v are Fourier functions.

[0070] When the second feature matrix is mapped into the corresponding high-dimensional space to obtain the high-dimensional second feature matrix, there are features that are approximately irrelevant between the high-dimensional second feature matrices. When the correlation value of the correlation function is calculated to take the minimum value, it is considered that the diluted features are approximately uncorrelated, and the weight at this time is used as the final sample weight.

[0071] S206: Construct an objective loss function according to the sample weights and the difference expression between the first feature matrix and the second feature matrix.

[0072] In step S204, the sample weight is the weight of each sample corresponding to each row in the second feature matrix. The sample weight is weighted with the difference between the calculated first feature matrix and the second feature matrix to construct the objective loss function.

[0073] In this embodiment, the sample weight is assigned to the difference expression between the first feature matrix and the second feature matrix to construct the objective loss function.

[0074] Optionally, a target loss function is constructed according to the sample weights and the difference between the first feature matrix and the second feature matrix calculated by a preset algorithm, including:

[0075] Use the KL divergence algorithm to calculate the difference between the first feature matrix and the second feature matrix, and obtain the difference function of the first feature matrix and the second feature matrix;

[0076] Construct a target loss function according to the difference function and the sample weights.

[0077] In this embodiment, the KL divergence algorithm is used to calculate the difference between the first feature matrix and the second feature matrix. The KL divergence represents the calculation of the similarity between the first feature matrix and the second feature matrix. The closer the KL divergence is to 0, the shorter the distance between the first feature matrix and the second feature matrix, that is, the closer the distribution is. The closer the KL divergence is to 1, the greater the distance between the first feature matrix and the second feature matrix, that is, the lower the similarity between the first feature matrix and the second feature matrix.

[0078] Based on the difference between the feature matrix output by the first model that meets the target requirements and the features output by the second model that does not meet the target requirements, and the sample weights of each row in each feature matrix, a loss function is constructed so that the trained second model can learn more features from the first model.

[0079] S207: Train the second model according to the target loss function to obtain target parameters.

[0080] In step S207, train the second model according to the target loss function, and adjust the initial parameters in the second model to obtain target parameters.

[0081] In this embodiment, when training the second model with sample images, for the target loss function, the training result and the feature values in the sample images are used to calculate the loss value through the target loss function, and it is judged whether the loss value meets the preset conditions. When the preset conditions are not met, the second model is updated by backpropagation according to the loss value to obtain the second model with updated model parameters, and the parameters that meet the preset conditions are used as target parameters.

[0082] S208: Update the initial parameters in the second model with the target parameters to generate a target model.

[0083] In step S208, when training each second model, supervised training can be performed, and a corresponding threshold is set. When the difference between the training result and the supervision label is less than the threshold, it is considered that the target loss function converges, and a target model is obtained.

[0084] In this embodiment, when training the second model using the sample images, for the target loss function, the training result obtained and the eigenvalue in the sample images are used to calculate the loss value through the target loss function, and it is determined whether the loss value meets the preset conditions. When the preset conditions are not met, the second model is updated by backpropagation according to the loss value to obtain the second initial model with updated model parameters. Then, the second model with updated model parameters is trained again based on the sample images until the loss value meets the preset conditions, and the target model is obtained in sequence.

[0085] Optionally, training the second model according to the target loss function and adjusting the initial parameters in the second model to generate a target model that meets the target conditions includes:

[0086] Making a preset image generator generate N groups of sample images; N is an integer greater than 1, and each group of sample images includes M sample images;

[0087] Training the second model according to the N groups of sample images and the target loss function to obtain target parameters. In this embodiment, when training the second model, a preset image generator is used to generate N groups of sample images. The N groups of sample images can all be positive sample images or can include negative sample images. When the N groups of sample images include negative sample images, training can be performed using sample pairs. Each sample pair selects the same positive sample image and negative sample image. Since the magnitude of the backpropagated gradient of the positive sample image is determined by the accumulation of the differences of each positive and negative sample image pair, this results in that once the number of samples in the sample image pair is too large, it will be very difficult to find a training separation surface that satisfies all the sample image pairs in the feature space, and the convergence of training will deteriorate accordingly. At the same time, repeatedly calculating the distance difference between two sample images as the backpropagated gradient each time will lead to redundancy in training and an increase in training time. Since the size comparison of a pair of positive and negative sample images in different sample image pairs may occur multiple times, increasing the number of positive and negative sample images in the sample pair too much has little help for training. Therefore, one positive sample image and one negative sample image can be selected in each sample pair. The selected sample pairs are used to train the second model until the target loss function converges to obtain the target model.

[0088] Using a preset image generator, generate M sample images. Input the M sample images into the first model to output the first feature tensors of each feature in each sample image. The first feature tensors of all features in each sample image form the first feature matrix of the corresponding sample image. M is an integer greater than 1. Input the M sample images into the second model to output the second feature tensors of each feature in each sample image. The second feature tensors of all features in each sample image form the second feature matrix of the corresponding sample image. Through a preset kernel function, perform feature decoupling on the second feature matrix to obtain a multi-feature second feature matrix. And according to the correlation function between each column of features in the multi-feature second feature matrix, calculate the sample weights for each row in the multi-feature second feature matrix. According to the sample weights and the difference expression of the first feature matrix and the second feature matrix calculated through a preset algorithm, construct an objective loss function. According to the objective loss function, train the second model to adjust the initial parameters in the second model to generate a target model, which solves the problem of fewer data samples. According to the sample weights in each sample image and the difference between the output features of the first model and the second model, construct a loss function corresponding to the second model, calculate the sample weights, avoid the problem of uneven distribution of the generated sample image features, and improve the accuracy of the generated model in application.

[0089] See Figure 3 , which is a schematic flowchart of a model generation method based on knowledge-free distillation provided by an embodiment of the present invention. As Figure 3 , the model generation method based on knowledge-free distillation may include the following steps:

[0090] S301: Obtain the first model;

[0091] S302: Use the first model to discriminate the similarity between the random images generated by the preset image generator and the real images;

[0092] S303: Train the initial image generator through a pre-constructed adversarial loss function to obtain a trained image generator, and use the trained image generator as the preset image generator. In this embodiment, the first model and the second model are deep learning convolutional neural network models. The first model is trained using a training set to obtain a first model that meets the target conditions. Among them, the first model includes p network layers, and the second model includes q network layers. Both p and q are integers greater than zero, and p is greater than q. The first model is a trained deep learning model, and the second model is an untrained deep learning model.

[0093] It should be noted that the first model and the second model can be neural network models of the same type, that is, the first model and the second model have the same network layer structure, or the first model and the second model can be neural network models of different types, that is, the network layer structures of the first model and the second model are different.

[0094] For example, the first model can be constructed based on the Residual Network ResNet34; the second model can be constructed based on ResNet10. Since larger networks often face the problems of large and redundant deep learning network models and difficulty in meeting the real-time requirements of recognition speed, while small network models are very likely to have insufficient model feature representation ability due to small number of parameters, resulting in low model accuracy. The problem is that although the real-time requirements of online applications are met, the accuracy requirements cannot be satisfied. Therefore, training the second model with the knowledge of the first model can enable the large network model to play a positive role in the small network model, so that the second model can obtain better fitting parameters, thereby improving the accuracy of the second model.

[0095] When training the second model with the first model, when the sample images are small, sample images can be generated according to the generator. In this embodiment, the generator uses the first model for discriminative supervision and trains the initial image generator using a pre-constructed adversarial loss function to obtain a trained image generator.

[0096] In this embodiment, a preset random signal is input to the initial image generator to generate a random image. The first model is used to discriminate the generated random image. The goal is to minimize the difference between the generated random image and the preset sample image. The training process of the initial image generator is to make the discriminator make mistakes as much as possible, while the training process of the discriminator is to improve the ability to distinguish between the preset sample image and the random image generated by the generator. By continuous training, the ability of the initial image generator is improved, and it can generate random images similar to the preset sample image. The abilities of the initial image generator and the discriminator are improved simultaneously through continuous training.

[0097] Optionally, training the initial image generator using a pre-constructed adversarial loss function to obtain the trained image generator includes:

[0098] Obtain a preset cross-entropy loss function, information entropy loss function, and regularization loss function, and construct an adversarial loss function;

[0099] Use the first model that meets the target conditions to discriminate the similarity between the random image generated by the preset image generator based on the preset random signal and the real image, and train the initial image generator using the adversarial loss function to obtain the trained image generator.

[0100] In this embodiment, the preset adversarial loss function is composed of a preset cross-entropy loss function, an information entropy loss function, and a regularization loss function. The preset cross-entropy loss function is used to determine whether the image generated by the preset image generator is real. The preset information entropy loss function calculates the information entropy loss on the output of the first model so that the images generated by the healing image generator are more uniform in category. The preset regularization loss function is the regularization of the activation layer of the first model, making the images generated by the preset image generator closer to the authenticity of the preset sample images. By constructing the adversarial loss function with three loss functions, the adversarial loss function can better supervise the training process of the preset image generator. The preset image generator is trained with the adversarial loss function to obtain a trained image generator.

[0101] Optionally, obtaining the preset cross-entropy loss function, the information entropy loss function, and the regularization loss function, and constructing the adversarial loss function includes:

[0102] Obtaining the proportion coefficients of the preset cross-entropy loss function, the information entropy loss function, and the regularization loss function;

[0103] According to the proportion coefficients, as well as the preset cross-entropy loss function, the information entropy loss function, and the regularization loss function, constructing the adversarial loss function.

[0104] In this embodiment, when obtaining the proportion coefficients of the preset cross-entropy loss function, the information entropy loss function, and the regularization loss function, corresponding proportion coefficients of the cross-entropy loss function, the information entropy loss function, and the regularization loss function are obtained from different combinations of proportion coefficients. For example, in the early stage of training, when the first model is used as the discriminator and contributes more, the proportion coefficient of the cross-entropy loss function can be larger, and the proportion coefficients of the information entropy loss function and the regularization loss function can be smaller. When in the later stage of training, more attention is paid to the quality of the random images generated by the remaining image generator, the information entropy loss function and the regularization loss function can take larger values, and the cross-entropy loss function can take smaller values. According to the proportion coefficients, as well as the cross-entropy loss function, the information entropy loss function, and the regularization loss function, a weighted sum is performed to construct the adversarial loss function.

[0105] S303: Use the preset image generator to generate M sample images, where M is an integer greater than 1;

[0106] S304: Input the M sample images into the first model to obtain the first feature tensors of each feature in each sample image. The first feature tensors of all features in each sample image constitute the first feature matrix corresponding to the sample image;

[0107] S305: Input the M sample images into the second model to obtain the second feature tensors of each feature in each sample image. The second feature tensors of all features in each sample image constitute the second feature matrix corresponding to the sample image.

[0108] S306: Perform feature decoupling on the second feature matrix through a preset kernel function to obtain a multi-feature second feature matrix.

[0109] S307: Calculate the sample weights of each row in the multi-feature second feature matrix according to the correlation function between each column of features in the multi-feature second feature matrix.

[0110] S308: Construct an objective loss function according to the sample weights and the difference expression between the first feature matrix and the second feature matrix.

[0111] S309: Train the second model according to the objective loss function to obtain target parameters.

[0112] S310: Update the initial parameters in the second model with the target parameters to generate a target model.

[0113] Among them, the content of the above steps S303 to S310 is the same as that of the above steps S201 to S208, and reference can be made to the description of the above steps S201 to S208, which will not be elaborated here.

[0114] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a model generation device based on knowledge distillation provided by an embodiment of the present invention. In this embodiment, each unit included in the terminal is used to execute Figures 2 to 3 the corresponding steps in the embodiment. Specifically, please refer to Figures 2 to 3 and Figures 2 to 3 the relevant descriptions in the corresponding embodiments. For the sake of convenience of description, only the parts related to this embodiment are shown. Refer to Figure 4 , the generation device 40 includes: a generation module 41, a first feature matrix determination module 42, a second feature matrix determination module 43, a multi-feature second feature matrix determination module 44, a sample weight determination module 45, an objective loss function construction module 46, a target parameter determination module 47, and a target model acquisition module 48.

[0115] The generation module 41 is used to generate M sample images by using a preset image generator, where M is an integer greater than 1.

[0116] The first feature matrix determination module 42 is configured to input the M sample images into the first model to obtain the first feature tensors of each feature in each sample image, and the first feature tensors of all features in each sample image constitute the first feature matrix corresponding to the sample image;

[0117] The second feature matrix determination module 43 is configured to input the M sample images into the second model to obtain the second feature tensors of each feature in each sample image, and the second feature tensors of all features in each sample image constitute the second feature matrix corresponding to the sample image;

[0118] The multi-feature second feature matrix determination module 44 is configured to perform feature decoupling on the second feature matrix through a preset kernel function to obtain a multi-feature second feature matrix;

[0119] The sample weight determination module 45 is configured to calculate the sample weights of each row in the multi-feature second feature matrix according to the correlation function between each column of features in the multi-feature second feature matrix;

[0120] The target loss function construction module 46 is configured to construct a target loss function according to the sample weights and the difference expression between the first feature matrix and the second feature matrix;

[0121] The target parameter determination module 47 is configured to train the second model according to the target loss function to obtain target parameters;

[0122] The target model acquisition module 48 is configured to update the initial parameters in the second model with the target parameters to generate a target model.

[0123] Optionally, the above sample weight determination module 45 includes:

[0124] The correlation function construction unit is configured to construct a correlation function based on the correlation relationship between each column of features in the multi-feature second feature matrix.

[0125] The sample weight calculation unit is configured to take the weight with the smallest correlation value of the correlation function as the sample weight.

[0126] Optionally, the above target loss function construction module 46 includes:

[0127] The difference function determination unit is configured to use the KL divergence algorithm to calculate the difference between the first feature matrix and the second feature matrix to obtain the difference function between the first feature matrix and the second feature matrix.

[0128] The target loss function determination unit is configured to construct a target loss function according to the difference function and the sample weights.

[0129] Optionally, the above-mentioned target model acquisition module 44 includes:

[0130] A generation unit, configured to generate N groups of sample images by using a preset image generator; N is an integer greater than 1, and each group of sample images includes M sample images.

[0131] A training unit, configured to train the second model according to the N groups of sample images and the target loss function to obtain target parameters.

[0132] Optionally, the above-mentioned generation device 40 further includes:

[0133] An acquisition module, configured to acquire a first model;

[0134] A discrimination module, configured to use the first model to discriminate the similarity between the random images generated by the preset image generator and the real images;

[0135] A preset image generator determination module, configured to train an initial image generator through a pre-constructed adversarial loss function to obtain a trained image generator, and use the trained image generator as the preset image generator.

[0136] Optionally, the above-mentioned preset image generator determination module includes:

[0137] An adversarial loss function determination unit, configured to obtain a preset cross-entropy loss function, an information entropy loss function, and a regularization loss function, and construct an adversarial loss function.

[0138] An image generator training unit, configured to use the first model to discriminate the similarity between the random images generated by the preset image generator based on a preset random signal and the real images, and train the initial image generator through the adversarial loss function to obtain a trained image generator.

[0139] Optionally, the above-mentioned adversarial loss function determination unit includes:

[0140] A proportion coefficient determination subunit, configured to obtain the proportion coefficients of the preset cross-entropy loss function, the information entropy loss function, and the regularization loss function.

[0141] An adversarial loss function construction subunit, configured to construct an adversarial loss function according to the proportion coefficients, the preset cross-entropy loss function, the information entropy loss function, and the regularization loss function.

[0142] It should be noted that for the information interaction, execution process, etc. between the above-mentioned units, since they are based on the same concept as the method embodiment of the present invention, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0143] Figure 5It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. As Figure 5 shown, the computer device of this embodiment includes: at least one processor ( Figure 5 only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, it implements the steps in any of the above-mentioned embodiments of the model generation method based on knowledge distillation.

[0144] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 5 merely examples of computer devices, and do not constitute a limitation on computer devices. A computer device may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include a network interface, a display screen, and an input device, etc.

[0145] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0146] The memory includes a readable storage medium, an internal memory, etc. Among them, the internal memory may be the memory of the computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be the hard disk of the computer device, and in some other embodiments, it may also be an external storage device of the computer device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory may also include both the internal storage unit of the computer device and the external storage device. The memory is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory may also be used to temporarily store data that has been output or will be output.

[0147] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working process of the units and modules in the above device can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiment of the present invention, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiment can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0148] To implement all or part of the processes in the above method embodiment of the present invention, it can also be completed by a computer program product. When the computer program product runs on a computer device, it enables the computer device to execute and implement the steps in the above method embodiment.

[0149] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0150] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0151] In the embodiments provided by the present invention, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0152] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0153] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A model generation method based on knowledge distillation-free, characterized in that, The generation method includes: Using a preset image generator to generate M sample images, where M is an integer greater than 1; Inputting the M sample images into a first model to obtain a first feature tensor for each feature in each sample image, and the first feature tensors of all features in each sample image form a first feature matrix corresponding to the sample image; Inputting the M sample images into a second model to obtain a second feature tensor for each feature in each sample image, and the second feature tensors of all features in each sample image form a second feature matrix corresponding to the sample image; Performing feature decoupling on the second feature matrix through a preset kernel function to obtain a multi-feature second feature matrix; Calculating the sample weights for each row in the multi-feature second feature matrix according to the correlation function between each column of features in the multi-feature second feature matrix; Constructing an objective loss function according to the sample weights and the difference expression between the first feature matrix and the second feature matrix; Training the second model according to the objective loss function to obtain target parameters; Updating the initial parameters in the second model using the target parameters to generate a target model.

2. The model generation method based on knowledge distillation as described in claim 1, wherein Before using the preset image generator to generate M sample images, it further includes: Obtaining a first model; Using the first model to discriminate the similarity between a random image generated by a preset image generator and a real image; Training an initial image generator through a pre-constructed adversarial loss function to obtain the trained image generator, and using the trained image generator as the preset image generator.

3. The model generation method based on knowledge distillation as claimed in claim 2, wherein, The training of the initial image generator through the pre-constructed adversarial loss function to obtain the trained image generator includes: Obtaining a preset cross-entropy loss function, an information entropy loss function, and a regularization loss function, and constructing an adversarial loss function; Using the first model to discriminate the similarity between a random image generated by the preset image generator based on a preset random signal and a real image, and training the initial image generator through the adversarial loss function to obtain the trained image generator.

4. The model generation method based on knowledge distillation-free according to claim 3, wherein The obtaining of the preset cross-entropy loss function, the information entropy loss function, and the regularization loss function, and constructing the adversarial loss function includes: Obtaining the proportion coefficients of the preset cross-entropy loss function, the information entropy loss function, and the regularization loss function; Constructing an adversarial loss function according to the proportion coefficients and the preset cross-entropy loss function, the information entropy loss function, and the regularization loss function.

5. The model generation method based on knowledge distillation-free as claimed in claim 1, wherein The calculating of the sample weights for each row in the multi-feature second feature matrix according to the correlation function between each column of features in the multi-feature second feature matrix includes: Constructing a correlation function based on the correlation relationship between each column of features in the multi-feature second feature matrix; Taking the weight with the smallest correlation value of the correlation function as the sample weight.

6. The model generation method based on knowledge distillation-free as claimed in claim 1, wherein, The constructing of the objective loss function according to the sample weights and the difference expression between the first feature matrix and the second feature matrix includes: Using the KL divergence algorithm, calculate the difference between the first feature matrix and the second feature matrix to obtain the difference function of the first feature matrix and the second feature matrix; Construct a target loss function according to the difference function and the sample weights.

7. The model generation method based on knowledge distillation-free as claimed in claim 1, wherein Training the second model according to the target loss function to obtain target parameters, including: Using a preset image generator to generate N groups of sample images; N is an integer greater than 1, and each group of sample images includes M sample images; Training the second model according to the N groups of sample images and the target loss function to obtain target parameters.

8. A model generation device based on knowledge distillation-free, characterized in that, The model generation device includes: A generation module for using a preset image generator to generate M sample images, where M is an integer greater than 1; A first feature matrix determination module for inputting the M sample images into a first model to obtain a first feature tensor of each feature in each sample image, and the first feature tensors of all features in each sample image form the first feature matrix corresponding to the sample image; A second feature matrix determination module for inputting the M sample images into a second model to obtain a second feature tensor of each feature in each sample image, and the second feature tensors of all features in each sample image form the second feature matrix corresponding to the sample image; A multi-feature second feature matrix determination module for decoupling the features of the second feature matrix through a preset kernel function to obtain a multi-feature second feature matrix; A sample weight determination module for calculating the sample weights of each row in the multi-feature second feature matrix according to the correlation function between each column of features in the multi-feature second feature matrix; A target loss function construction module for constructing a target loss function according to the sample weights and the difference expression between the first feature matrix and the second feature matrix; A target parameter determination module for training the second model according to the target loss function to obtain target parameters; A target model acquisition module for using the target parameters to update the initial parameters in the second model to generate a target model.

9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the model generation method based on knowledge-free distillation according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the model generation method based on knowledge-free distillation according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Model training method and device based on knowledge distillation, equipment and storage medium

    CN115062769A

  • Model compression method and system based on preview mechanism knowledge distillation

    CN115294407A