Image classification method and device based on data-free cloud edge knowledge distillation model, equipment and medium
Through the intelligent communication field in the field of intelligent communication, the data-free cloud edge knowledge distillation model is used to reconstruct the image dataset through the deep inversion diffusion generation module and the teacher model in the cloud, and guide the edge student model for training, solving the problem of low recognition accuracy of lightweight edge deployment networks, and achieving high-precision and real-time response image classification.
Patent Information
- Application Number
- CN202510250579.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-03
AI Technical Summary
In the field of intelligent communications, lightweight edge deployment networks cannot effectively conduct knowledge distillation training due to problems such as transmission restrictions and user privacy, resulting in low recognition accuracy.
The image classification method based on the data-free cloud edge knowledge distillation model is used to reconstruct the original image dataset by deploying the depth inversion diffusion generation module in the cloud and the pre-trained teacher model, thereby guiding the student model at the edge to perform distillation training.
High-precision training of lightweight student models is implemented without the need for original image datasets, ensuring high performance and real-time response of image classification.
Smart Images

Figure CN120088569A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent communication technologies, and in particular, to an image classification method, apparatus, device, and medium based on a data-free cloud-edge knowledge distillation model. Background Art
[0002] Efficiency and security are the core challenges in the field of intelligent communication. Currently, deep neural networks and cloud resources have promoted the rapid development of intelligent communication. With the increase in the capabilities and complexities of deep neural networks, their recognition capabilities have been significantly enhanced. However, due to factors such as inference speed, application flexibility, and reliability, high-performance complex deep neural networks cannot be implemented in practical applications. Instead, lightweight neural networks deployed on mobile edge devices can achieve near-real-time inference, and this ability has accelerated the development of the Internet of Things. However, the inherent computational and storage limitations of lightweight networks often hinder their widespread direct applications. Therefore, deploying lightweight networks (such as image classification models) on mobile devices has become an effective way to promote the development of intelligent communication.
[0003] Lightweight neural networks have accelerated the interpretation of information and communication efficiency, thus promoting the rapid development of the Internet of Things. However, enhancing the recognition capabilities of lightweight neural networks remains challenging. Among them, knowledge distillation is a technique for transferring knowledge from complex models to smaller models, and is commonly used to improve the recognition performance of lightweight networks. Currently, collaborative distillation between complex cloud-deployed networks and lightweight edge-deployed networks has become an effective strategy for guiding the training of lightweight neural networks. However, in a practical environment, when implementing intelligent communication tasks based on neural networks, some challenges will be encountered, including but not limited to: user privacy protection, raw data corruption, edge storage limitations, and transmission constraints. That is to say, practical problems such as transmission limitations and user privacy make it difficult to directly obtain the raw data required for knowledge distillation, resulting in the inability to effectively distill and train the current lightweight edge-deployed networks (such as image classification networks deployed at the edge), leading to relatively low recognition accuracy of lightweight edge-deployed networks. Summary of the Invention
[0004] Based on the above technical problems, the present invention provides an image classification method, apparatus, device, and medium based on a data-free cloud-edge knowledge distillation model, aiming to ensure high-precision image recognition of a student model deployed at the edge without the need for raw training data from the cloud.
[0005] In the first aspect of the present invention, an image classification method based on a data-free cloud-edge knowledge distillation model is provided. The data-free cloud-edge knowledge distillation model at least includes: a deep inverse diffusion generation module deployed in the cloud, a pre-trained teacher model deployed in the cloud, and a student model to be trained deployed at the edge. The teacher model is a model with image classification function trained using the original image dataset. The method includes: Through the deep inverse diffusion generation module, reconstruct the original image dataset from the perspectives of data features and data distribution to obtain a generated image dataset. Among them, the deep inverse diffusion generation module performs inverse training on the teacher model to reconstruct the original image dataset from the perspective of data features, and optimizes the inverse diffusion process through the deep inverse diffusion generation module to reconstruct the original image dataset from the perspective of data distribution. Input the generated image dataset into the teacher model and the student model for distillation learning to obtain a trained student model. Input the image to be classified into the trained student model to obtain an image classification result.
[0006] In the second aspect of the present invention, an image classification device based on a data-free cloud-edge knowledge distillation model is provided. The data-free cloud-edge knowledge distillation model at least includes: a deep inverse diffusion generation module deployed in the cloud, a pre-trained teacher model deployed in the cloud, and a student model to be trained deployed at the edge. The teacher model is a model with image classification function trained using the original image dataset. The device includes: A dataset generation module for reconstructing the original image dataset from the perspectives of data features and data distribution through the deep inverse diffusion generation module to obtain a generated image dataset. Among them, the deep inverse diffusion generation module performs inverse training on the teacher model to reconstruct the original image dataset from the perspective of data features, and optimizes the inverse diffusion process through the deep inverse diffusion generation module to reconstruct the original image dataset from the perspective of data distribution. A distillation training module for inputting the generated image dataset into the teacher model and the student model for distillation learning to obtain a trained student model. An image classification module for inputting the image to be classified into the trained student model to obtain an image classification result.
[0007] In the third aspect of the present invention, an electronic device is provided. The electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the image classification method based on the data-free cloud-edge knowledge distillation model in the first aspect of the embodiments of the present invention.
[0008] In a fourth aspect of the present invention, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the image classification method based on the data-free cloud-edge knowledge distillation model according to the first aspect of the embodiments of the present invention is implemented.
[0009] In the image classification method based on the data-free cloud-edge knowledge distillation model provided by the present invention, first, through a deep inverse diffusion generation module deployed in the cloud, the original image dataset is reconstructed from the perspectives of data features and data distribution to obtain a generated image dataset; then, the generated image dataset is respectively input into a trained teacher model deployed in the cloud and a student model to be trained deployed at the edge for distillation learning to obtain a trained student model; finally, the image to be classified is input into the trained student model to obtain an image classification result. Thus, based on the data-free cloud-edge knowledge distillation model, the present invention combines the inverse training of the trained teacher model deployed in the cloud with the diffusion mechanism through the deep inverse diffusion generation module deployed in the cloud, reconstructs the original image dataset from the perspectives of data features and distribution to guide the distillation training of the student model to be trained deployed at the edge, thereby avoiding the direct access and transmission of the original image dataset, ensuring that the lightweight student model to be trained can perform distillation training without the original image dataset, obtaining a trained student model, and achieving high-performance real-time response, that is, high-precision image category classification. Description of the Drawings
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0011] Figure 1 is a flowchart of the steps of an image classification method based on a data-free cloud-edge knowledge distillation model shown in an embodiment of the present invention; Figure 2 is a simplified schematic diagram of a data-free cloud-edge knowledge distillation model shown in an embodiment of the present invention; Figure 3 is a simplified schematic diagram of the inverse training of a teacher model shown in an embodiment of the present invention; Figure 4 is a comparison schematic diagram of the reverse diffusion process of this embodiment and the original reverse diffusion process shown in an embodiment of the present invention; Figure 5It is a simplified flowchart of a data-free cloud-edge knowledge distillation model shown in an embodiment of the present invention; Figure 6 It is a heatmap representation diagram of CKA between layers of a network with different convolutional neural network architectures shown in an embodiment of the present invention; Figure 7 It is a simplified schematic diagram of a multi-layer feature joint supervision distillation module shown in an embodiment of the present invention; Figure 8 It is a structural block diagram of an image classification device based on a data-free cloud-edge knowledge distillation model provided in an embodiment of the present invention; Figure 9 It is a schematic diagram of an electronic device shown in an embodiment of the present invention. Detailed implementation manners
[0012] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0013] Please refer to Figure 1 , Figure 1 It is a step flowchart of an image classification method based on a data-free cloud-edge knowledge distillation model shown in an embodiment of the present invention. As Figure 1 shown, the image classification method based on the data-free cloud-edge knowledge distillation model provided in this embodiment at least includes the following steps: Step S11: Through the deep inverse diffusion generation module, reconstruct the original image dataset from the perspectives of data features and data distribution to obtain a generated image dataset; wherein, the teacher model is reversely trained through the deep inverse diffusion generation module to reconstruct the original image dataset from the perspective of data features; and the reverse diffusion process is optimized through the deep inverse diffusion generation module to reconstruct the original image dataset from the perspective of data distribution.
[0014] In this embodiment, the data-free cloud-edge knowledge distillation model at least includes: a deep inverse diffusion generation module deployed in the cloud, a pre-trained teacher model deployed in the cloud, and a student model to be trained deployed at the edge. Among them, the teacher model is a model with image classification function trained using the original image dataset. The pre-trained teacher model can be a complex deep neural network model, and the student model to be trained can be a lightweight neural network model. Here, "lightweight" corresponds to "complex" and can be reflected in any one or more aspects such as model parameters and the number of network layers.
[0015] In this embodiment, a pre-trained teacher model deployed in the cloud can be used to reconstruct the original image dataset from the perspectives of data features and data distribution, obtaining a generated image dataset, which is used to train the student model to be trained deployed at the edge. Specifically, in this embodiment, the trained teacher model is reversely trained by a deep inverse diffusion generation module to reconstruct the original image dataset from the perspective of data features; and, the reverse diffusion process is optimized by the deep inverse diffusion generation module to reconstruct the original image dataset from the perspective of data distribution, obtaining a generated image dataset, which has the same or similar data features and data distribution as the original image dataset.
[0016] Step S12: Input the generated image dataset into the teacher model and the student model for distillation learning to obtain the trained student model.
[0017] In this embodiment, after obtaining the generated image dataset, the generated image dataset is transmitted from the cloud to the edge, and the generated image dataset is respectively input into the trained teacher model in the cloud and the student model to be trained in the edge for distillation learning, training the student model to be trained at the edge until the trained student model is obtained, and this trained student model is used for image classification.
[0018] Step S13: Input the image to be classified into the trained student model to obtain the image classification result.
[0019] In this embodiment, after obtaining the trained student model, the image to be classified can be input into this trained student model to obtain the image classification result output by the trained student model, thereby realizing the high-precision image classification function in the edge device.
[0020] In this embodiment, based on the data-free cloud-edge knowledge distillation model, the reverse training of the trained teacher model deployed in the cloud is combined with the diffusion mechanism by the deep inverse diffusion generation module deployed in the cloud to reconstruct the original image dataset from the perspectives of data features and distribution to guide the distillation training of the student model to be trained deployed at the edge, thus avoiding the direct access and transmission of the original image dataset, ensuring that the lightweight student model to be trained can perform distillation training without the original image dataset, obtaining the trained student model to achieve high-performance real-time response, that is, high-precision image category classification.
[0021] In one embodiment, as Figure 2 shown, Figure 2It is a simplified schematic diagram of a data-free cloud-edge knowledge distillation model shown in an embodiment of the present invention. In Figure 2 , the left box represents the model and calculation deployed in the cloud, while the right box represents the model and calculation deployed at the edge (i.e., Figure 2 the edge side in Figure 2 ). The dashed arrow represents the knowledge distillation process, and the parallel arrow represents the batch data transmission process of the generated image dataset. It can be seen from Figure 2 that the data-free cloud-edge knowledge distillation model of this embodiment does not directly access the original image dataset, but completely relies on the trained teacher model deployed in the cloud (i.e.,
[0022] the trained teacher network in to guide the lightweight student model (i.e., the student network) to be trained at the edge. This model avoids the direct access and transmission of data while maintaining the performance of the lightweight student network. Without relying on the original training data, the data-free cloud-edge knowledge distillation model proposed in this embodiment compresses the parameters and floating-point numbers to 1 / 20 and 1 / 12 respectively, only reducing the accuracy by 0.27%.
[0023] In this embodiment, since the deep neural network designed for the image classification task judges the data category by the features extracted and recognized by each layer, the feature expression of the original image dataset is obtained through reverse training. In order to obtain a new data distribution extremely similar to the original image dataset, this embodiment designs an improved diffusion model. This innovative diffusion model uses the guidance of the trained teacher network during the data generation process to reduce the possibility of generating bias. Compared with the reverse diffusion process of the original diffusion model, the improved diffusion model (i.e., the deep inverse diffusion generation module) proposed in this embodiment provides discriminant guidance (DG) and classification guidance (CG) generated by the trained teacher model during the reverse diffusion process (i.e., the inverse diffusion generation process). The trained teacher model can provide richer data information for the diffusion process.
[0024] The deep reverse diffusion generation module can use the trained teacher model as a guide to provide prior information for the improved diffusion generation of the generated image dataset. In this embodiment, the diffusion process is improved through the deep reverse diffusion generation module to reconstruct the original image dataset from the perspective of the data distribution. Specifically, in the process of restoring the data distribution based on the improved diffusion generation, the trained teacher model provides a classification guidance signal and / or a discriminative guidance signal for the improved diffusion process to optimize the reverse diffusion process in the diffusion process and obtain the generated image dataset.
[0025] Among them, the discriminative guidance signal involves using additional discriminative guidance to adjust the score model of the diffusion process, so that the generated image dataset is closer to the original image dataset; the classification guidance signal optimizes and adjusts the diffusion generation by evaluating whether the generated samples are accurately classified.
[0026] Combined with the above embodiments, in one implementation, the present invention also provides an image classification method based on a data-free cloud-edge knowledge distillation model. In this method, "reconstructing the original image dataset from the perspectives of data features and data distribution through the deep reverse diffusion generation module to obtain the generated image dataset" in step S11 above specifically includes step S31 and step S32: Step S31: Input random noise into the teacher model, and through the deep reverse diffusion generation module, perform inverse training on the teacher model to minimize the cross-entropy loss to obtain the target image dataset.
[0027] In this embodiment, the features of the original image dataset are restored by performing inverse training on the trained teacher model. Different from the forward training of the deep neural network, the inverse training of the deep neural network involves optimizing random noise into target data under the supervision of specific labels: input random noise into the trained teacher model (i.e., the pre-trained teacher model), and through the deep reverse diffusion generation module, perform inverse training on the trained teacher model to minimize the cross-entropy loss to obtain the target image dataset, which is the image dataset after feature restoration of the original image dataset.
[0028] Among them, the cross-entropy loss is determined according to the image class label (i.e., the aforementioned specific label) and the output result of the teacher model in the inverse training process. Minimizing the cross-entropy loss is used to minimize the difference between the data features of the target image dataset and the data features of the original image dataset.
[0029] To better understand the feature extraction and recognition mechanisms of deep neural networks, this embodiment proposes an inverse expression of the teacher model. For a deep convolutional neural network, this embodiment is based on a trained teacher model to construct a corresponding inverse representation pattern: so as to achieve to approximate reconstruction. Based on this principle, this embodiment uses the reverse training of the trained teacher model to achieve a preliminary reconstruction of the original image dataset: for an image classification task, this embodiment uses cross-entropy loss as a guiding function to achieve data reconstruction and generation optimization: (1); wherein, is the trained teacher model, X is the random noise to be optimized, is the optimized data, i.e., the target image dataset, Y is the label of the data to be generated, i.e., the image category label, L is the cross-entropy loss, is the height, width, and number of channels of the vector X.
[0030] In one embodiment, as shown in Figure 3 , Figure 3 is a simplified schematic diagram of the reverse training of a teacher model shown in an embodiment of the present invention. In Figure 3 , the thin arrow represents the forward training of the teacher model, indicating the process of layer-by-layer feature extraction during network training, while the thick arrow represents the reverse training of the teacher model, i.e., the inverse training. The label and random Gaussian noise are the inputs for the inverse training of the model, and its goal is to optimize the Gaussian noise into the target image dataset through this inverse training process. That is, the thick arrow represents the reverse training process of optimizing the random noise into the target image dataset. The category label is introduced into the teacher network, and the input noise is transformed into the target image dataset by continuous optimization.
[0031] Step S32: Based on the target image dataset, through the deep inverse diffusion generation module, reconstruct the original image dataset from the perspective of data distribution to obtain the generated image dataset.
[0032] In this embodiment, after obtaining the target image dataset, the target image dataset can be diffused through the deep inverse diffusion generation module to obtain the generated image dataset, so as to reconstruct the original image dataset from the perspective of data distribution.
[0033] Combining the above embodiments, in one implementation, the present invention also provides an image classification method based on a data-free cloud-edge knowledge distillation model. In this embodiment, the above step S32 may specifically include step S41 and step S42: Step S41: Based on the target image dataset, perform the forward diffusion process and the reverse diffusion process through the deep inverse diffusion generation module.
[0034] In this embodiment, after obtaining the target image dataset, based on the target image dataset, the forward diffusion process and the reverse diffusion process can be sequentially performed through the deep inverse diffusion generation module to obtain the generated image dataset.
[0035] The forward diffusion process is to add noise (add iterative Gaussian noise) to the target image dataset to obtain the intermediate image dataset. Among them, the continuous denoising forward diffusion probability model established by using the stochastic differential equation is: (2); where t is a continuous value within the diffusion time interval [0, T], and represent the drift coefficient and the volatility coefficient respectively, is the stochastic differential based on x at time t, x is the target image data, is the standard Wiener process at time t.
[0036] Step S42: Optimize the reverse diffusion process with the classification guidance signal and / or discriminant guidance signal provided by the teacher model to obtain the generated image dataset.
[0037] In this embodiment, the reverse diffusion process is to denoise the intermediate image dataset to obtain the generated image dataset. Among them, in the deep inverse diffusion generation module of this embodiment, the reverse diffusion process is optimized with the classification guidance signal and / or discriminant guidance signal provided by the trained teacher model, so as to obtain the generated image dataset.
[0038] Based on the continuous time model framework, the corresponding reverse diffusion generation process can be expressed as: (3); where, is the probability density in the forward diffusion process, t is the time, is the reverse standard Wiener process at time t, is the diffusion time.
[0039] On the basis of training a scoring network to evaluate the score of real data, correspondingly, Equation (3) can be transformed into: (4); where the score function .
[0040] In Equation (4), when the finally obtained score function When reaching the local optimum, and deviating from the global optimum will lead to generation bias.
[0041] Adjusting the score function can reduce the degree of bias and contribute to the reverse generation of the diffusion model. Based on this, this embodiment conducts discriminative guidance for the reverse diffusion process: by designing a new correction term with the discriminative guidance provided by the teacher model to alleviate this deviation, and the reverse diffusion process at this time is as follows: (5); Among them, assuming is the optimal probability density of reverse diffusion, the correction term cannot be directly calculated because the density ratio is unattainable.
[0042] Train a binary discriminator at each generation point t to distinguish between true and false, which can approximate this density ratio: (6); During the optimization process of this deep neural network discriminator, instead of this correction term : (7); Since the trained teacher model already has the ability to recognize the categories of the original image dataset, the trained teacher model is selected as the guiding discriminator. When fine-tuning the teacher model in the binary classification task, use the fine-tuning process of the discriminator to provide discriminative guidance for the reverse diffusion process in the diffusion process, that is, provide a discriminative guidance signal. This method can minimize the distance between its convergence point and the optimal point. In addition, Equation (6) is transformed into: (8).
[0043] Among them, the discriminative guidance signal provided by this embodiment in the reverse diffusion process is: ; The reverse diffusion process with the discriminative guidance signal added can be expressed as: (9).
[0044] This embodiment conducts classification guidance for the reverse diffusion process: This embodiment provides classification guidance based on the generation process of knowledge distillation without data for the image classification task, which can enhance the classification attributes of the generated data and have a positive impact on the distillation process. In this embodiment, independent of training the classifier, the trained teacher model is used as the classifier to provide classification guidance for diffusion generation. The classification guidance signal provided by this embodiment in the reverse diffusion process is: ; The reverse diffusion process with the added classification guidance signal can be expressed as: (10).
[0045] Based on this, in this embodiment, after adding the discriminant guidance signal and the classification guidance signal, the inverse diffusion based on the stochastic differential equation (the above formula (4)) can be transformed into: (11); Among them, is the classifier in the diffusion time t.
[0046] In one embodiment, as Figure 4 shown, Figure 4 is a schematic diagram comparing the reverse diffusion process of this embodiment with the original reverse diffusion process shown in an embodiment of the present invention. Figure 4 (a) in Figure 4 is the reverse diffusion process of the traditional diffusion model, (b) in is the reverse diffusion process of the improved diffusion model in the deep reverse diffusion generation module proposed in this embodiment. Compared with the reverse diffusion generation process of the original diffusion model, the improved diffusion model proposed in this embodiment provides a discriminant guidance signal (DG) and a classification guidance signal (CG) generated by the trained teacher model at each step (i.e., each feature denoising process) in the reverse diffusion generation process. Among them, the discriminant guidance signal (DG) is: ; The reverse diffusion process with the added discriminant guidance signal is: .
[0047] In one embodiment, as Figure 5 shown, Figure 5 is a simplified flowchart of a data-free cloud-edge knowledge distillation model shown in an embodiment of the present invention. In Figure 5 , the left thick dashed box represents the network model and calculation deployed in the cloud, including the generation of the generated image dataset (i.e., the generated proxy data in Figure 5 ) and the inference and calculation part of the teacher-side distillation process. The thin dashed box in the lower right corner shows the network and calculation deployed at the edge, including the distillation process and the student-side calculation part. The dashed arrows and dotted arrows in the figure represent the generation process of the generated image dataset. Among them, the dashed arrows represent the feature recovery process of the data, and the dotted arrows represent the distribution recovery process of the data. The solid arrows represent the knowledge distillation process.
[0048] Combined with the above embodiments, in one implementation manner, the present invention further provides an image classification method based on a data-free cloud-edge knowledge distillation model. In this embodiment, the above step S12 may specifically include steps S51 to S53: Step S51: Input the generated image dataset into the teacher model and the student model respectively, and add a similarity constraint between the intermediate features output by the corresponding network layers of the teacher model and the student model through the multi-layer feature joint supervision distillation module.
[0049] In this embodiment, the data-free cloud-edge knowledge distillation model further includes: a multi-layer feature joint supervision distillation module deployed at the cloud-edge. In this embodiment, the multi-layer feature joint supervision distillation module being deployed at the cloud-edge means that the input of the multi-layer feature joint supervision distillation module is the relevant features obtained by the pre-trained teacher model in the cloud processing the generated image dataset, and the relevant features obtained by the student model to be trained in the edge processing the generated image dataset. The multi-layer feature joint supervision distillation module is used to effectively guide the student model based on the edge by the teacher model based on the cloud through the feature expression similarity constraint. This module adds a new distillation constraint on the basis of the traditional distillation model, enhancing the guidance of the teacher model to the student model.
[0050] In this embodiment, the obtained generated image dataset can be batch-transmitted to the edge to assist in completing the refinement task. By applying the feature expression similarity constraint to the two networks based on multi-layers of features, the recognition difference between the cloud and the edge network is minimized: the generated image dataset is input into the student model to be trained, and the relevant features (such as the features output by each network layer) obtained by the student model processing the generated image dataset are transmitted to the multi-layer feature joint supervision distillation module. And, the obtained generated image dataset is input into the trained teacher model, and the relevant features (such as the features output by each network layer) obtained by the teacher model processing the generated image dataset are transmitted to the multi-layer feature joint supervision distillation module. A similarity constraint is added between the intermediate features output by the corresponding network layers of the teacher model and the student model through the multi-layer feature joint supervision distillation module, promoting the feature selection of the student model. In this way, the ability of the student model to be trained to select and learn effective features is enhanced, thus promoting the distillation effect.
[0051] Step S52: Perform distillation learning under the similarity constraint to obtain the trained student model.
[0052] When using a deep neural network for knowledge distillation, the optimal result means that the student model can imitate the feature extraction rules of the teacher model under the guidance of the teacher model and learn more discriminative features. To achieve this goal, in this embodiment, a similarity constraint of features is proposed and implemented in the multi-layer feature joint supervision distillation module, focusing on the middle to high-level features of the two models, so as to perform distillation learning under the similarity constraint to obtain the trained student model.
[0053] The similarity constraint of this embodiment is used to enable the student model to continuously select and extract effective features for image classification under the guidance of the teacher model.
[0054] Combined with the above embodiments, in one implementation manner, the present invention also provides an image classification method based on a data-free cloud-edge knowledge distillation model. In this method, the above step S52 may specifically include steps S61 to S63: Step S61: Obtain intermediate features of multiple pairs of network layer outputs with the same output size of the teacher model and the student model.
[0055] In this embodiment, the multi-layer feature joint supervision distillation module can obtain intermediate features of multiple pairs of network layer outputs with the same output size of the teacher model and the student model. Among them, multiple pairs of network layers with the same output size refer to network layers of corresponding layers; for example, the output sizes of the first network layer of the teacher model and the first network layer of the student model are the same, the output sizes of the second network layer of the teacher model and the second network layer of the student model are the same, the output sizes of the third network layer of the teacher model and the third network layer of the student model are the same, and so on.
[0056] Step S62: Calculate the similarity between the intermediate features of the multiple pairs of network layer outputs to obtain a similarity loss.
[0057] In this embodiment, after obtaining the intermediate features of multiple pairs of network layer outputs, the multi-layer feature joint supervision distillation module can calculate the similarity between the intermediate features of multiple pairs of network layer outputs to obtain a similarity loss. For example, based on the intermediate features of each pair of network layer outputs, the similarity between the intermediate features of each pair of network layer outputs can be calculated to obtain a similarity loss.
[0058] In an alternative implementation manner, when evaluating the similarity constraint between networks, the centering kernel alignment (CKA) metric can be selected as the measurement standard for feature similarity. Specifically, the objective function (i.e., similarity loss) of feature similarity supervision based on centering kernel alignment (CKA) is: (12); Among them, the size of n is the number of multi-layer feature joint supervised distillation modules, which is determined by the network structure selected for distillation. For example, if the network structures of the student model and the teacher model have 4 layers, then four similarity constraints are set (each pair of intermediate features output by the network layers corresponds to 1 similarity constraint), so n = 4. and respectively represent the intermediate features learned by the teacher model and the student model in the i-th multi-layer feature joint supervised distillation module. is the similarity loss. CKA is an index used to measure the similarity of feature representations in deep neural networks. By statistically analyzing the similarity between the feature representations of different neural network structures, CKA is used to measure the similarity of feature expressions between deep neural networks.
[0059] The calculation formula of CKA between feature X and feature Y is as follows: (13); Among them, K and L are the kernel functions of feature x and y, and The empirical estimate value of HSIC is: (14); H is the central matrix in, for the linear kernel , HSIC is expressed as: (15). Among them, || · ||∗ represents the nuclear norm, and cov() represents the covariance.
[0060] Step S63: Perform distillation learning based on at least the similarity loss to obtain a trained student model.
[0061] In this embodiment, after obtaining the similarity loss, distillation learning is performed based on at least the similarity loss to update the model parameters of the student model to be trained, so as to obtain a trained student model.
[0062] In one embodiment, in order to effectively monitor the similarity of feature expressions during distillation, this embodiment analyzes the effectiveness of similarity expressions and CKA. A deep convolutional neural network is selected as the distillation model framework, and the CIFAR-10 dataset is used to analyze the representation mode of CKA. Specifically, this embodiment studies two distillation configurations: (WRN40-2 → WRN16-2 and (ResNet34 → ResNet18). These configurations are selected as examples for our analysis.
[0063] Using the CIFAR-10 dataset as input, the similarity (CKA) of feature expressions between different layers of the trained ResNet34 and ResNet18 models is asFigure 6 as shown in (a) of Figure 6 This is a heatmap representation of the CKA between the layers of a network with different convolutional neural network architectures shown in an embodiment of the present invention. In Figure 6 (a) of Figure 6 , the left side represents the layer-by-layer similarity metric within the network, while the right side shows the block-by-block similarity metric. Similarly, Figure 6 (b) of Figure 6 shows the CKA of the trained WRN40-2 and WRN16-2 models. In
[0064] (b) of Figure 7 , the layer-by-layer similarity is shown in detail on the left side, and the block-by-block similarity is presented on the right side. Figure 7 This indicates that there are significant similarities in the feature expressions between networks with different structures. Based on the similarity matrix of the feature expressions of the two networks, significant performance is shown on the sub-diagonal. This strongly similar pattern is more obvious in the shallow layers than in the deep layers. This observation shows that the feature outputs of the corresponding blocks in convolutional neural networks with similar architectures exhibit considerable similarities, and thus a similarity constraint based on CKA is adopted. Figure 7 In an embodiment, as shown in Figure 7 , when implementing multi-layer feature joint supervision, feature similarity measurements are performed on the corresponding modules of two networks with the same output size (such as between block 1 and block 1, block 2 and block 2, etc. in ). Therefore, based on the structural features of the network, four execution points are selected during the supervision process. The number of selected execution points is consistent with n in the above formula (12).
[0065] Combining the above embodiments, in one implementation, the present invention also provides an image classification method based on a data-free cloud-edge knowledge distillation model. In this method, in addition to the above steps, steps S71 and S72 may further be included, and step S63 above may specifically include step S73: Step S71: Obtain the first prediction result output by the teacher model based on the generated image dataset, and obtain the second prediction result output by the student model based on the generated image dataset.
[0066] In this embodiment, after the generated image dataset is input into the teacher model and the student model respectively, the first prediction result output by the teacher model based on the generated image dataset may further be obtained, and the second prediction result output by the student model based on the generated image dataset may be obtained.
[0067] Step S62: Calculate the distillation loss based on the first prediction result and the second prediction result.
[0068] In this embodiment, the distillation loss can be calculated based on the first prediction result and the second prediction result. Specifically, the distillation loss L kd is as follows: (16) where x is the generated image dataset, and are the output probability distributions of the teacher model and the student model respectively, and D kl is the KL divergence.
[0069] Step S63: Update the model parameters of the student model based on the similarity loss and the distillation loss until the trained student model is obtained.
[0070] In this embodiment, after obtaining the similarity loss and the distillation loss, the total loss can be calculated based on the similarity loss and the distillation loss, and then the model parameters of the student model to be trained are updated based on the total loss until the total loss converges, and the model parameters of the student model are fixed to obtain the trained student model.
[0071] In an alternative embodiment, the total loss is: (17) where is the similarity loss, is the distillation loss, and are the weights of the similarity loss and the distillation loss respectively. In an alternative example, = =1, indicating that the two loss terms have equal importance.
[0072] As shown below, the image classification method based on the data-free cloud-edge knowledge distillation model proposed in this embodiment is systematically summarized. Among them, the input is: the given trained teacher network and the output is: the lightweight student network , represents the generated image dataset. and are the output distributions of the penultimate layers of the student model and the teacher model respectively, and are the features output by the teacher network and the student network respectively:
[0073] In summary, currently, completing model compression through cloud-edge knowledge distillation is an effective strategy for implementing deep neural networks in intelligent communication. However, problems such as data unavailability, such as transmission limitations, user privacy, and data corruption, hinder the smooth execution of knowledge distillation. To ensure data security and model efficiency in intelligent communication, an embodiment of the present invention proposes a new cloud-edge knowledge distillation model that does not rely on the original training data, namely, a data-free cloud-edge knowledge distillation model. Instead of directly accessing the original image dataset, this embodiment designs a deep inverse diffusion generation module to mine the original image dataset from a trained teacher model deployed in the cloud, realizing the reconstruction of the original image dataset from features to distribution. In addition, a multi-layer feature joint supervision distillation module is designed to enhance the guidance of the cloud-based teacher model to the edge-based lightweight student model. The data-free cloud-edge knowledge distillation model of this embodiment achieves a win-win situation for the accuracy and speed of the edge lightweight network without relying on the original training data.
[0074] It should be noted that for method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0075] Based on the same inventive concept, an embodiment of the present invention provides an image classification device based on a data-free cloud-edge knowledge distillation model. In this embodiment, the data-free cloud-edge knowledge distillation model at least includes: a deep inverse diffusion generation module deployed in the cloud, a pre-trained teacher model deployed in the cloud, and a student model to be trained deployed at the edge; the teacher model is a model with image classification function trained using the original image dataset. Refer to Figure 8 , Figure 8 is the structural block diagram of an image classification device based on a data-free cloud-edge knowledge distillation model provided by an embodiment of the present invention. As Figure 8 shown, the image classification device based on the data-free cloud-edge knowledge distillation model of this embodiment may include: a dataset generation module, configured to reconstruct the original image dataset from the perspectives of data features and data distribution through the deep inverse diffusion generation module to obtain a generated image dataset; wherein, the deep inverse diffusion generation module is used to perform reverse training on the teacher model to reconstruct the original image dataset from the perspective of data features; the deep inverse diffusion generation module is used to optimize the reverse diffusion process to reconstruct the original image dataset from the perspective of data distribution; A distillation training module for inputting the generated image dataset into the teacher model and the student model for distillation learning to obtain a trained student model; An image classification module for inputting an image to be classified into the trained student model to obtain an image classification result.
[0076] Optionally, the dataset generation module includes: A first guidance module for optimizing the reverse diffusion process through the deep inverse diffusion generation module with the classification guidance signal and / or discrimination guidance signal provided by the teacher model to obtain the generated image dataset.
[0077] Optionally, the dataset generation module includes: A target image generation module for inputting random noise into the teacher model and performing backpropagation training on the teacher model through the deep inverse diffusion generation module to minimize the cross-entropy loss; the cross-entropy loss is determined based on the image class label and the output result of the teacher model during the backpropagation training, and minimizing the cross-entropy loss is used to minimize the difference between the data features of the target image dataset and the data features of the original image dataset; A diffusion generation module for reconstructing the original image dataset from the perspective of data distribution based on the target image dataset through the deep inverse diffusion generation module to obtain the generated image dataset.
[0078] Optionally, the diffusion generation module includes: A diffusion execution module for performing a forward diffusion process and a reverse diffusion process based on the target image dataset through the deep inverse diffusion generation module; A second guidance module for optimizing the reverse diffusion process with the classification guidance signal and / or discrimination guidance signal provided by the teacher model to obtain the generated image dataset.
[0079] Optionally, the data-free cloud-edge knowledge distillation model further includes: a multi-layer feature joint supervision distillation module deployed at the cloud-edge; the distillation training module includes: A similarity constraint module for inputting the generated image dataset into the teacher model and the student model respectively, and adding a similarity constraint between the intermediate features output by the corresponding network layers of the teacher model and the student model through the multi-layer feature joint supervision distillation module; A distillation learning module for performing distillation learning under the similarity constraint to obtain the trained student model, and the similarity constraint is used to enable the student model to select and extract effective features for image classification under the guidance of the teacher model.
[0080] Optionally, the distillation learning module includes: A feature acquisition module for obtaining intermediate features of multiple pairs of network layer outputs with the same output size of the teacher model and the student model; A first loss determination module for calculating the similarity between the intermediate features of the multiple pairs of network layer outputs to obtain a similarity loss; A loss training module for performing distillation learning at least based on the similarity loss to obtain a trained student model.
[0081] Optionally, the apparatus further includes: A prediction result acquisition module for obtaining a first prediction result output by the teacher model based on the generated image dataset, and obtaining a second prediction result output by the student model based on the generated image dataset; A second loss determination module for calculating a distillation loss based on the first prediction result and the second prediction result; The loss training module includes: A loss training sub-module for updating the model parameters of the student model based on the similarity loss and the distillation loss until the trained student model is obtained.
[0082] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the image classification method based on the data-free cloud-edge knowledge distillation model as described in any one of the above embodiments of the present invention.
[0083] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, as Figure 9 shown. Figure 9 is a schematic diagram of an electronic device shown in an embodiment of the present invention. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the steps in the image classification method based on the data-free cloud-edge knowledge distillation model as described in any one of the above embodiments of the present invention.
[0084] For the apparatus embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment.
[0085] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, refer to each other.
[0086] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the embodiments of the present invention can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0087] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0088] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0089] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal devices, such that a series of operation steps are executed on the computer or other programmable terminal devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal devices provide steps for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0090] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the embodiments of the present invention.
[0091] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent in such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the said element.
[0092] The above has introduced in detail an image classification method, apparatus, device and medium based on a data-free cloud-edge knowledge distillation model provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An image classification method based on a data-free cloud-edge knowledge distillation model, characterized in that: The data-free cloud-edge knowledge distillation model at least includes: a deep inversion diffusion generation module deployed in the cloud, a pre-trained teacher model deployed in the cloud, and a student model to be trained deployed in the edge; the teacher model is a model with image classification function obtained by training using the original image data set; the method includes: The original image data set is reconstructed from the perspective of data features and data distribution by the deep inverted diffusion generation module to obtain a generated image data set; wherein the teacher model is reversely trained by the deep inverted diffusion generation module to reconstruct the original image data set from the perspective of data features; the reverse diffusion process is optimized by the deep inverted diffusion generation module to reconstruct the original image data set from the perspective of data distribution; Inputting the generated image data set into the teacher model and the student model to perform distillation learning to obtain a trained student model; The image to be classified is input into the trained student model to obtain the image classification result.
2. The image classification method based on the data-free cloud-edge knowledge distillation model according to claim 1 is characterized in that: The reverse diffusion process is optimized by the deep reverse diffusion generation module to reconstruct the original image data set from the perspective of data distribution, including: The reverse diffusion process is optimized by the deep reverse diffusion generation module using the classification guidance signal and / or the discrimination guidance signal provided by the teacher model to obtain the generated image data set.
3. The image classification method based on the data-free cloud-edge knowledge distillation model according to claim 1 is characterized in that: The original image data set is reconstructed from the perspective of data features and data distribution through the depth inversion diffusion generation module to obtain a generated image data set, including: Inputting random noise into the teacher model, and reversely training the teacher model through the deep inversion diffusion generation module to minimize the cross entropy loss to obtain a target image dataset; the cross entropy loss is determined according to the image category label and the output result of the teacher model in the reverse training process, and minimizing the cross entropy loss is used to minimize the difference between the data features of the target image dataset and the data features of the original image dataset; Based on the target image data set, the original image data set is reconstructed from the perspective of data distribution through the depth inversion diffusion generation module to obtain the generated image data set.
4. The image classification method based on the data-free cloud-edge knowledge distillation model according to claim 3 is characterized in that: Based on the target image data set, the original image data set is reconstructed from the perspective of data distribution through the depth inversion diffusion generation module to obtain the generated image data set, including: Based on the target image data set, performing a forward diffusion process and a reverse diffusion process through the depth inversion diffusion generation module; The back diffusion process is optimized using the classification guidance signal and / or the discrimination guidance signal provided by the teacher model to obtain the generated image data set.
5. The image classification method based on the data-free cloud-edge knowledge distillation model according to claim 1 is characterized in that: The data-free cloud-edge knowledge distillation model also includes: a multi-layer feature joint supervision distillation module deployed on the cloud-edge end; inputting the generated image data set into the teacher model and the student model to perform distillation learning to obtain a trained student model, including: Inputting the generated image dataset into the teacher model and the student model respectively, and adding similarity constraints between the intermediate features output by the corresponding network layers of the teacher model and the student model through the multi-layer feature joint supervision distillation module; Distillation learning is performed under the similarity constraint to obtain the trained student model, and the similarity constraint is used to enable the student model to select and extract effective features for image classification under the guidance of the teacher model.
6. The image classification method based on the data-free cloud-edge knowledge distillation model according to claim 5 is characterized in that: Performing distillation learning under the similarity constraint to obtain the trained student model includes: Obtaining intermediate features of multiple pairs of network layer outputs of the same output size of the teacher model and the student model; Calculating the similarity between the intermediate features output by the multiple pairs of network layers to obtain a similarity loss; At least based on the similarity loss, distillation learning is performed to obtain a trained student model.
7. The image classification method based on the data-free cloud-edge knowledge distillation model according to claim 6 is characterized in that: The method further comprises: Obtain a first prediction result output by the teacher model based on the generated image dataset, and obtain a second prediction result output by the student model based on the generated image dataset; Calculating distillation loss based on the first prediction result and the second prediction result; At least based on the similarity loss, distillation learning is performed to obtain a trained student model, including: Based on the similarity loss and the distillation loss, the model parameters of the student model are updated until the trained student model is obtained.
8. An image classification device based on a data-free cloud-edge knowledge distillation model, characterized in that: The data-free cloud-edge knowledge distillation model at least includes: a deep inversion diffusion generation module deployed in the cloud, a pre-trained teacher model deployed in the cloud, and a student model to be trained deployed in the edge; the teacher model is a model with image classification function obtained by training using the original image data set; the device includes: A data set generation module, configured to reconstruct the original image data set from the perspective of data features and data distribution through the deep inversion diffusion generation module to obtain a generated image data set; wherein the teacher model is reversely trained through the deep inversion diffusion generation module to reconstruct the original image data set from the perspective of data features; and the reverse diffusion process is optimized through the deep inversion diffusion generation module to reconstruct the original image data set from the perspective of data distribution; A distillation training module, used for inputting the generated image data set into the teacher model and the student model to perform distillation learning to obtain a trained student model; The image classification module is used to input the image to be classified into the trained student model to obtain the image classification result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by the processor, the image classification method based on the data-free cloud-edge knowledge distillation model is implemented as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image classification method based on the data-free cloud-edge knowledge distillation model is implemented as described in any one of claims 1 to 7.