Model compression method, system and equipment based on self-distillation and storage medium

By automatically generating soft targets as teacher signals during model training, self-distillation model compression is achieved, the problem of relying on external teacher models in the existing technology is solved, which significantly reduces the computational complexity and storage requirements, and improves the flexibility and applicability of model compression.

CN119940457AInactive Publication Date: 2025-05-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510412939.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing knowledge distillation method needs to rely on external teacher models, increasing computational costs and limiting the flexibility and applicability of model compression.

Method used

The self-distillation model compression is achieved by automatically generating soft targets as teacher signals during model training. This method does not rely on external teacher models, and generates a smooth probability distribution through temperature scaling technology, which is used to optimize model parameters.

Benefits of technology

It significantly reduces the computational complexity and storage requirements of the model, reduces the computational cost, improves the flexibility and applicability of model compression, and can achieve model compression while maintaining model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940457A_ABST
    Figure CN119940457A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, and discloses a model compression method, system and device based on self-distillation and a storage medium, the method comprises the following steps: obtaining sample data, inputting the sample data into a preset neural network model, carrying out classification processing to obtain a corresponding classification result, and determining classification loss; adjusting parameters in a preset neural network model according to the classification loss, and recalculating the classification loss; when the classification loss obtained by recalculation meets a preset condition, performing classification processing on the sample data based on a self-distillation module in the correspondingly adjusted neural network model to obtain a corresponding soft target; determining self-distillation loss based on the classification result corresponding to the sample data and the soft target corresponding to the sample data; and according to the classification loss and the self-distillation loss, continuing adjustment to obtain a trained model. According to the invention, the calculation cost is reduced, the flexibility and applicability of model compression are improved, and efficient model compression is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine learning technology, for example, to a model compression method, system, device and storage medium based on self-distillation. Background Art

[0002] With the rapid development of deep learning technology, the size and complexity of models are increasing, resulting in the limitation of computing resources and storage resources in practical applications. In order to enable deep learning models to be deployed in resource-constrained environments (such as mobile devices, edge computing devices, etc.), model compression technology has gradually become a research hotspot.

[0003] Existing model compression technologies mainly include pruning, quantization, knowledge distillation and other methods. Among them, knowledge distillation is an effective model compression technology that achieves model compression by transferring the knowledge of a large model (teacher model) to a small model (student model).

[0004] However, existing knowledge distillation methods need to rely on external teacher models, which not only increases the computational cost but also limits the flexibility and applicability of model compression.

[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application. Summary of the invention

[0006] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0007] The model compression method, system, device and storage medium based on self-distillation disclosed in the present invention reduce the computing cost and improve the flexibility and applicability of model compression.

[0008] The present disclosure provides a model compression method based on self-distillation, the method comprising: Obtain sample data, and input the sample data into a preset neural network model for classification processing to obtain corresponding classification results; Determine the classification loss based on the classification result corresponding to the sample data and the sample label corresponding to the sample data; According to the classification loss, the parameters in the preset neural network model are adjusted, and the classification loss is recalculated based on the adjusted neural network model; When the recalculated classification loss meets the preset conditions, the sample data is classified based on the self-distillation module in the corresponding adjusted neural network model to obtain the corresponding soft target; Determine the self-distillation loss based on the classification result corresponding to the sample data and the soft target corresponding to the sample data; According to the classification loss and self-distillation loss, the parameters in the adjusted neural network model continue to be adjusted until the stopping condition is met to obtain the trained model.

[0009] In some embodiments, the classification loss is determined based on the classification result corresponding to the sample data and the sample label corresponding to the sample data. , satisfying the formula: , in, is the sample label corresponding to the i-th sample data, indicating the true label of the c-th category of the i-th sample data; is the classification result corresponding to the i-th sample data output by the neural network model, indicating the probability of the c-th category corresponding to the i-th sample data; N is the number of sample data.

[0010] In some embodiments, the self-distillation loss is determined based on the classification result corresponding to the sample data and the soft target corresponding to the sample data. , satisfying the formula: , in, is the soft target corresponding to the i-th sample data; is the classification result corresponding to the i-th sample data output by the neural network model, indicating the probability of the c-th category corresponding to the i-th sample data; N is the number of sample data.

[0011] In some embodiments, It is determined by temperature scaling technique and satisfies the formula: , Where T is the temperature parameter, which is used to control the smoothness of the soft target.

[0012] In some embodiments, according to the classification loss and the self-distillation loss, the parameters in the adjusted neural network model are adjusted until the stop condition is met to obtain the trained model, including: Determine the total loss function based on the classification loss and self-distillation loss , satisfying the formula: , in, is the classification loss, is the loss from distillation, It is a dynamically adjusted weight parameter that gradually decreases with training to balance the task objectives and knowledge distillation objectives; Based on the total loss function, the parameters in the adjusted neural network model are adjusted until the stopping condition is met to obtain the trained model.

[0013] In some embodiments, during the initial phase of training, =1, adjust the parameters in the preset neural network model and only optimize the classification loss ; When the recalculated classification loss meets the preset conditions, it gradually decreases The value of .

[0014] The present disclosure provides a model compression system based on self-distillation, the system comprising: The acquisition module is used to acquire sample data and input the sample data into a preset neural network model for classification processing to obtain corresponding classification results; A processing module, used for determining a classification loss based on a classification result corresponding to the sample data and a sample label corresponding to the sample data; The processing module is further used to adjust the parameters in the preset neural network model according to the classification loss, and recalculate the classification loss based on the adjusted neural network model; The processing module is further used to classify the sample data based on the self-distillation module in the corresponding adjusted neural network model to obtain the corresponding soft target when the recalculated classification loss meets the preset condition; The processing module is further used to determine the self-distillation loss based on the classification result corresponding to the sample data and the soft target corresponding to the sample data; The processing module is also used to adjust the parameters in the adjusted neural network model according to the classification loss and the self-distillation loss until the stopping condition is met to obtain the trained model.

[0015] An embodiment of the present disclosure provides an electronic device, the device comprising at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor to enable the at least one processor to perform the above-mentioned self-distillation-based model compression method.

[0016] An embodiment of the present disclosure provides a storage medium storing program instructions, which, when run, execute the above-mentioned self-distillation-based model compression method.

[0017] The model compression method, system, device, and storage medium based on self-distillation provided by the embodiments of the present disclosure can achieve the following technical effects: The disclosed embodiment enables the model to automatically generate soft targets during the training process as its own teacher signal, thereby achieving efficient model compression. While maintaining the model performance, it can significantly reduce the model's computational complexity and storage requirements, reduce computational costs, improve the flexibility and applicability of model compression, and can be applied to a variety of deep learning tasks and scenarios.

[0018] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] One or more embodiments are exemplarily described by corresponding drawings, which do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements, and the drawings do not constitute a scale limitation, and wherein: Figure 1 is a flow chart of a model compression method based on self-distillation provided in an embodiment of the present disclosure; Figure 2 is a structural schematic diagram of a model compression system based on self-distillation provided in an embodiment of the present disclosure; Figure 3 It is a structural schematic diagram of a model compression device based on self-distillation provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0020] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and systems can be simplified for display.

[0021] The terms "first", "second", etc. in the embodiments of the present disclosure are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so as to describe the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.

[0022] Unless otherwise stated, the term "plurality" means two or more.

[0023] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B indicates: A or B.

[0024] The term "and / or" is a description of the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.

[0025] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.

[0026] The following describes the self-distillation-based model compression method, system, device and storage medium provided by the embodiments of the present disclosure in conjunction with the accompanying drawings.

[0027] Figure 1 It is a flow chart of a model compression method based on self-distillation provided in an embodiment of the present disclosure.

[0028] Combination Figure 1 As shown, the model compression method based on self-distillation may include: S101, obtaining sample data, and inputting the sample data into a preset neural network model to perform classification processing to obtain corresponding classification results; S102, determining a classification loss based on a classification result corresponding to the sample data and a sample label corresponding to the sample data; S103, adjusting parameters in a preset neural network model according to the classification loss, and recalculating the classification loss based on the adjusted neural network model; S104, when the recalculated classification loss meets the preset conditions, the sample data is classified based on the self-distillation module in the corresponding adjusted neural network model to obtain the corresponding soft target; S105, determining a self-distillation loss based on a classification result corresponding to the sample data and a soft target corresponding to the sample data; S106, according to the classification loss and the self-distillation loss, the parameters in the adjusted neural network model are continuously adjusted until the stopping condition is met to obtain the trained model.

[0029] In some embodiments, the classification loss is determined based on the classification result corresponding to the sample data and the sample label corresponding to the sample data. , satisfying the formula: , in, is the sample label corresponding to the i-th sample data, indicating the true label of the c-th category of the i-th sample data; is the classification result corresponding to the i-th sample data output by the neural network model, indicating the probability of the c-th category corresponding to the i-th sample data; N is the number of sample data.

[0030] In some embodiments, the self-distillation loss is determined based on the classification result corresponding to the sample data and the soft target corresponding to the sample data. , satisfying the formula: , in, is the soft target corresponding to the i-th sample data; is the classification result corresponding to the i-th sample data output by the neural network model, indicating the probability of the c-th category corresponding to the i-th sample data; N is the number of sample data.

[0031] In some embodiments, It is determined by temperature scaling technique and satisfies the formula: , in, is the optimized c, and T is the temperature parameter used to control the smoothness of the soft target.

[0032] In some embodiments, according to the classification loss and the self-distillation loss, the parameters in the adjusted neural network model are adjusted until the stop condition is met to obtain the trained model, including: Determine the total loss function based on the classification loss and self-distillation loss , satisfying the formula: , in, is the classification loss, is the loss from distillation, It is a dynamically adjusted weight parameter that gradually decreases with training to balance the task objectives and knowledge distillation objectives; Based on the total loss function, the parameters in the adjusted neural network model are adjusted until the stopping condition is met to obtain the trained model.

[0033] In some embodiments, during the initial phase of training, =1, adjust the parameters in the preset neural network model and only optimize the classification loss ; When the recalculated classification loss meets the preset conditions, it gradually decreases The value of .

[0034] Specifically, in terms of model structure design, the present disclosure adopts a deep neural network as the basic model, and the model structure includes an input layer, multiple hidden layers and an output layer. In order to realize the self-distillation mechanism, a self-distillation module is introduced into the model to generate soft targets.

[0035] The input of the model is raw data, and the output is the task target (such as classification probability). The self-distillation module converts the output of the model into a smooth probability distribution as a soft target through temperature scaling technology.

[0036] In terms of loss function design, the loss function of the present disclosure consists of two parts: traditional classification loss and self-distillation loss. The specific formula is as follows: Classification loss: The task objective used to optimize the model. The formula is: , in, is the sample label corresponding to the i-th sample data, indicating the true label of the c-th category of the i-th sample data; is the classification result corresponding to the i-th sample data output by the neural network model, indicating the probability of the c-th category corresponding to the i-th sample data; N is the number of sample data.

[0037] Self-distillation loss: The knowledge distillation target used to optimize the model, the formula is: , in, is the soft target corresponding to the i-th sample data; is the classification result corresponding to the i-th sample data output by the neural network model, indicating the probability of the c-th category corresponding to the i-th sample data; N is the number of sample data.

[0038] in, It is determined by temperature scaling technique and satisfies the formula: , in, is the optimized c, and T is the temperature parameter used to control the smoothness of the soft target.

[0039] Total loss function: Combining classification loss and self-distillation loss, a total loss function that dynamically adjusts weights is designed: , in, is the classification loss, is the loss from distillation, It is a dynamically adjusted weight parameter that gradually decreases with training to balance the task objectives and knowledge distillation objectives; Based on the total loss function, the parameters in the adjusted neural network model are adjusted until the stopping condition is met to obtain the trained model.

[0040] In terms of training strategy, the present disclosure adopts a phased training strategy, the specific steps are as follows: Phase 1: Optimizing mission objectives In the initial stage of training, only the classification loss is optimized , that is, by setting = 1. The goal of this stage is to let the model learn the task objectives and ensure the performance of the model on the task.

[0041] Stage 2: Introducing self-distillation losses After the model performance reaches a certain level, that is, when the recalculated classification loss meets the preset conditions, the self-distillation loss is gradually introduced. , that is, gradually decreases The goal of this stage is to achieve model compression through the self-distillation mechanism while maintaining model performance.

[0042] Phase 3: Stabilization Training In the later stages of training, keep The stability of the value ensures that the model achieves a balance between task objectives and knowledge distillation objectives.

[0043] In terms of the self-distillation mechanism, the core of the self-distillation mechanism in this disclosure is to generate soft targets through temperature scaling technology. The specific implementation steps are as follows: Generate soft targets: At each training iteration, the probability distribution of the model output Generating soft targets through temperature scaling techniques .

[0044] Calculating self-distillation loss: Based on soft targets And the model output Calculating self-distillation losses .

[0045] Update model parameters: according to the total loss function Update model parameters and optimize model performance.

[0046] In a specific example, a code example based on the PyTorch framework is used to show the specific implementation of the model compression method based on self-distillation: import torch import torch.nn as nn import torch.optim as optim class SelfDistillationModel(nn.Module): def __init__(self, input_size, hidden_size, output_size, temperature=3.0): super(SelfDistillationModel, self).__init__() self.fc1 = nn.Linear(input_size, hidden_size) self.fc2 = nn.Linear(hidden_size, output_size) self.temperature = temperature def forward(self, x): x = torch.relu(self.fc1(x)) output = self.fc2(x) return output def generate_soft_targets(self, output): # Generate soft targets by temperature scaling return torch.softmax(output / self.temperature, dim=1) def train_model(model, train_loader, val_loader, epochs=100, alpha=1.0, alpha_decay=0.99): criterion_cls = nn.CrossEntropyLoss() criterion_distill = nn.KLDivLoss(reduction='batchmean') optimizer = optim.Adam(model.parameters(), lr=0.001) for epoch in range(epochs): model.train() total_loss = 0.0 for inputs, targets in train_loader: optimizer.zero_grad() outputs = model(inputs) # Generate soft targets soft_targets = model.generate_soft_targets(outputs) # Calculate classification loss loss_cls = criterion_cls(outputs, targets) # Calculate self-distillation loss loss_distill = criterion_distill(outputs.softmax(dim=1), soft_targets) # Total loss loss = alpha * loss_cls + (1 - alpha) * loss_distill loss.backward() optimizer.step() total_loss += loss.item() # Adjust alpha alpha = max(alpha * alpha_decay, 0.1) # Validation stage model.eval() val_loss = 0.0 with torch.no_grad(): for inputs, targets in val_loader: outputs = model(inputs) soft_targets = model.generate_soft_targets(outputs) loss_cls = criterion_cls(outputs, targets) loss_distill = criterion_distill(outputs.softmax(dim=1), soft_targets) loss = alpha * loss_cls + (1 - alpha) * loss_distill val_loss += loss.item() print(f"Epoch {epoch+1}, Train Loss: {total_loss / len(train_loader)},Val Loss: {val_loss / len(val_loader)}") if __name__ == "__main__": # Example Dataset input_size = 784 hidden_size = 256 output_size = 10 model = SelfDistillationModel(input_size, hidden_size, output_size) # Assume train_loader and val_loader are already defined train_model(model, train_loader, val_loader) The above embodiments are described below in conjunction with specific scenarios.

[0047] Scenario 1: The developer describes the requirements through voice: "Create a user login page." The system automatically generates the corresponding low-code components and recommends related APIs.

[0048] Scenario 2: The developer draws a flowchart by hand, and the system automatically converts it into a workflow and generates the corresponding code.

[0049] Scenario 3: Multiple developers edit the same module at the same time, and the system automatically merges the code and resolves conflicts.

[0050] This paper discloses a model compression method based on self-distillation. The core innovation is to achieve model compression through the self-distillation mechanism without relying on an external teacher model. Specific innovations include: Self-distillation mechanism: The model automatically generates soft targets during training as its own teacher signal. Through temperature scaling technology, the model is able to generate a smooth probability distribution, thereby achieving self-transfer of knowledge.

[0051] Loss function design: Combining traditional classification loss and self-distillation loss, a loss function with dynamic weight adjustment is designed. This loss function can balance the task objectives and knowledge distillation objectives, ensuring that the performance of the model does not decrease during the compression process.

[0052] Training strategy: A phased training strategy is adopted, which first optimizes the task objectives and then gradually introduces self-distillation losses. This strategy can ensure the stability of the model during training and avoid performance fluctuations.

[0053] The disclosed model compression method based on self-distillation achieves efficient model compression by allowing the model to automatically generate soft targets as its own teacher signal during the training process. While maintaining the model performance, it can significantly reduce the model's computational complexity and storage requirements, reduce the computational cost, improve the flexibility and applicability of model compression, and can be applied to a variety of deep learning tasks and scenarios.

[0054] and Figure 1 Corresponding to the model compression method based on self-distillation in the present disclosure, the present disclosure also provides a model compression system based on self-distillation, such as Figure 2 As shown, the system may specifically include: The acquisition module 201 is used to acquire sample data and input the sample data into a preset neural network model to perform classification processing to obtain corresponding classification results; A processing module 202, configured to determine a classification loss based on a classification result corresponding to the sample data and a sample label corresponding to the sample data; The processing module 202 is further used to adjust the parameters in the preset neural network model according to the classification loss, and recalculate the classification loss based on the adjusted neural network model; The processing module 202 is further used to classify the sample data based on the self-distillation module in the corresponding adjusted neural network model to obtain the corresponding soft target when the recalculated classification loss meets the preset condition; The processing module 202 is further used to determine the self-distillation loss based on the classification result corresponding to the sample data and the soft target corresponding to the sample data; The processing module 202 is also used to adjust the parameters in the adjusted neural network model according to the classification loss and the self-distillation loss until the stopping condition is met to obtain the trained model.

[0055] In some embodiments, the classification loss is determined based on the classification result corresponding to the sample data and the sample label corresponding to the sample data. , satisfying the formula: , in, is the sample label corresponding to the i-th sample data, indicating the true label of the c-th category of the i-th sample data; is the classification result corresponding to the i-th sample data output by the neural network model, indicating the probability of the c-th category corresponding to the i-th sample data; N is the number of sample data.

[0056] Determine the self-distillation loss based on the classification result corresponding to the sample data and the soft target corresponding to the sample data , satisfying the formula: , in, is the soft target corresponding to the i-th sample data; is the classification result corresponding to the i-th sample data output by the neural network model, indicating the probability of the c-th category corresponding to the i-th sample data; N is the number of sample data.

[0057] In some embodiments, It is determined by temperature scaling technique and satisfies the formula: , in, is the optimized c, and T is the temperature parameter used to control the smoothness of the soft target.

[0058] In some embodiments, according to the classification loss and the self-distillation loss, the parameters in the adjusted neural network model are adjusted until the stop condition is met to obtain the trained model, including: Determine the total loss function based on the classification loss and self-distillation loss , satisfying the formula: , in, is the classification loss, is the loss from distillation, It is a dynamically adjusted weight parameter that gradually decreases with training to balance the task objectives and knowledge distillation objectives; Based on the total loss function, the parameters in the adjusted neural network model are adjusted until the stopping condition is met to obtain the trained model.

[0059] In some embodiments, during the initial phase of training, =1, adjust the parameters in the preset neural network model and only optimize the classification loss ; When the recalculated classification loss meets the preset conditions, it gradually decreases The value of .

[0060] Combination Figure 3As shown, the embodiment of the present disclosure also provides a model compression device 300 based on self-distillation, including a processor 304 and a memory 301. Optionally, the system may also include a communication interface 302 and a bus 303. Among them, the processor 304, the communication interface 302, and the memory 301 can communicate with each other through the bus 303. The communication interface 302 can be used for information transmission. The processor 304 can call the logic instructions in the memory 301 to execute the model compression method based on self-distillation of the above embodiment.

[0061] In addition, the logic instructions in the memory 301 described above can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0062] The memory 301, as a computer-readable storage medium, can be used to store software programs, computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor 304 executes the function application and data processing by running the program instructions / modules stored in the memory 301, that is, the model compression method based on self-distillation in the above embodiment is implemented.

[0063] The memory 301 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 301 may include a high-speed random access memory and may also include a non-volatile memory.

[0064] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured as a model compression method based on self-distillation.

[0065] The computer-readable storage medium mentioned above may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.

Claims

1. A model compression method based on self-distillation, characterized in that: The method comprises: Obtaining sample data, and inputting the sample data into a preset neural network model to perform classification processing to obtain corresponding classification results; Determining a classification loss based on a classification result corresponding to the sample data and a sample label corresponding to the sample data; According to the classification loss, adjusting the parameters in the preset neural network model, and recalculating the classification loss based on the adjusted neural network model; When the recalculated classification loss meets the preset conditions, the sample data is classified based on the self-distillation module in the corresponding adjusted neural network model to obtain the corresponding soft target; Determining a self-distillation loss based on a classification result corresponding to the sample data and a soft target corresponding to the sample data; According to the classification loss and the self-distillation loss, the parameters in the adjusted neural network model are continuously adjusted until a stopping condition is met to obtain a trained model.

2. The method according to claim 1, characterized in that The classification loss is determined based on the classification result corresponding to the sample data and the sample label corresponding to the sample data. , satisfying the formula: , in, is the sample label corresponding to the i-th sample data, indicating the true label of the c-th category of the i-th sample data; is the classification result corresponding to the i-th sample data output by the neural network model, indicating the probability of the c-th category corresponding to the i-th sample data; N is the number of sample data.

3. The method according to claim 1, characterized in that The self-distillation loss is determined based on the classification result corresponding to the sample data and the soft target corresponding to the sample data. , satisfying the formula: , in, is the soft target corresponding to the i-th sample data; is the classification result corresponding to the i-th sample data output by the neural network model, indicating the probability of the c-th category corresponding to the i-th sample data; N is the number of sample data.

4. The method according to claim 3, characterized in that Said It is determined by temperature scaling technique and satisfies the formula: , in, is the optimized c, and T is the temperature parameter used to control the smoothness of the soft target.

5. The method according to claim 1, characterized in that The step of adjusting the parameters in the adjusted neural network model according to the classification loss and the self-distillation loss until a stop condition is met to obtain a trained model includes: According to the classification loss and the self-distillation loss, the total loss function is determined , satisfying the formula: , in, is the classification loss, is the loss from distillation, It is a dynamically adjusted weight parameter that gradually decreases with training to balance the task objectives and knowledge distillation objectives; Based on the total loss function, the parameters in the adjusted neural network model are adjusted until a stopping condition is met to obtain a trained model.

6. The method according to claim 5, characterized in that In the initial stage of training, =1, adjust the parameters in the preset neural network model and only optimize the classification loss ; When the recalculated classification loss meets the preset conditions, it gradually decreases The value of .

7. A model compression system based on self-distillation, characterized in that: The system comprises: An acquisition module is used to acquire sample data and input the sample data into a preset neural network model to perform classification processing to obtain corresponding classification results; A processing module, configured to determine a classification loss based on a classification result corresponding to the sample data and a sample label corresponding to the sample data; The processing module is further used to adjust the parameters in the preset neural network model according to the classification loss, and recalculate the classification loss based on the adjusted neural network model; The processing module is further configured to classify the sample data based on the self-distillation module in the corresponding adjusted neural network model to obtain a corresponding soft target when the recalculated classification loss meets a preset condition; The processing module is further used to determine the self-distillation loss based on the classification result corresponding to the sample data and the soft target corresponding to the sample data; The processing module is also used to adjust the parameters in the adjusted neural network model according to the classification loss and the self-distillation loss until the stopping condition is met to obtain the trained model.

8. The system according to claim 7, characterized in that The classification loss is determined based on the classification result corresponding to the sample data and the sample label corresponding to the sample data. , satisfying the formula: , in, is the sample label corresponding to the i-th sample data, indicating the true label of the c-th category of the i-th sample data; is the classification result corresponding to the i-th sample data output by the neural network model, indicating the probability of the c-th category corresponding to the i-th sample data; N is the number of sample data.

9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Classification detection method for distributed small-scale medical data set

    CN113239985A

  • Long-tail image recognition method based on self-supervision and self-distillation

    CN113837238A

  • Self-distillation method based on tolerable label, computer equipment and storage medium

    CN116403069A

  • Self-distillation method and system based on progressive associative learning

    CN117892841A

  • Self-distillation method and system based on mixed sample, electronic equipment and medium

    CN118051848A