Knowledge distillation method and apparatus based on spatial distance alignment
By calculating the central feature loss and feature distance alignment loss of the teacher model and the student model, the student model parameters are optimized, and the problem of low efficiency of traditional knowledge distillation algorithm is solved, achieving more efficient knowledge distillation effect.
Patent Information
- Application Number
- PCT/CN2024/114729
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-08-27
- Publication Date
- 2025-07-03
AI Technical Summary
Traditional knowledge distillation algorithm has low efficiency and poor effect, and has failed to effectively utilize the relationship between the teacher model and the student model output features.
By calculating the central feature loss and feature distance alignment loss of the teacher model and the student model, the parameters of the student model are optimized to achieve knowledge distillation from the teacher model to the student model.
It improves the efficiency and effect of knowledge distillation and improves the training accuracy of student models.
Smart Images

Figure CN2024114729_03072025_PF_FP_ABST
Abstract
Description
Knowledge distillation method and device based on spatial distance alignment
[0001] This disclosure is based on and claims the priority of Chinese patent application with application number 202311788159X and application date December 25, 2023. The entire content of the Chinese patent application is hereby incorporated into this application by reference. Technical Field
[0002] The present disclosure relates to the technical field of knowledge distillation, and in particular to a knowledge distillation method and device based on spatial distance alignment. Background Art
[0003] Knowledge distillation algorithms use a trained teacher model to constrain the student model's outputs during training (effectively, using the teacher model to optimize the student model's parameters). Traditional knowledge distillation algorithms achieve this by simply comparing the output features of the teacher and student models. These algorithms fail to consider the relationships between the teacher and student model output features, resulting in low efficiency and poor results.
[0004] Summary of the Invention
[0005] In view of this, the embodiments of the present disclosure provide a knowledge distillation method, device, electronic device and computer-readable storage medium based on spatial distance alignment to solve the problem of low efficiency and poor effect of knowledge distillation algorithms in the prior art.
[0006] A first aspect of an embodiment of the present disclosure provides a knowledge distillation method based on spatial distance alignment, including: obtaining training data, inputting multiple training samples in the training data into a teacher model and a student model respectively according to batches, and outputting the teacher model features and student model features of each training sample in each batch; respectively calculating the teacher model center features and student model center features corresponding to the teacher model features and student model features of all training samples in each batch, wherein the training data is an image of a detection object; calculating the center feature loss between the teacher model center features and the student model center features corresponding to each batch; respectively calculating the teacher model feature distance and the student model feature distance corresponding to the teacher model features and the student model features of any two training samples in each batch; calculating the feature distance alignment loss between the teacher model feature distance and the student model feature distance corresponding to any two training samples in each batch; optimizing the model parameters of the student model based on the center feature loss corresponding to each batch and the feature distance alignment loss corresponding to any two training samples in each batch, so as to complete the knowledge distillation from the teacher model to the student model.
[0007] According to a second aspect of an embodiment of the present disclosure, a knowledge distillation device based on spatial distance alignment is provided, comprising: an acquisition module configured to acquire training data, input multiple training samples in the training data into a teacher model and a student model respectively in batches, and output teacher model features and student model features of each training sample in each batch, wherein the training data is an image of a detection object; a first calculation module configured to respectively calculate the teacher model center features and student model center features corresponding to the teacher model features and student model features of all training samples in each batch; a second calculation module configured to calculate the center feature loss between the teacher model center features and the student model center features corresponding to each batch; a third calculation module configured to respectively calculate the teacher model feature distance and the student model feature distance corresponding to the teacher model features and the student model features of any two training samples in each batch; a fourth calculation module configured to calculate the feature distance alignment loss between the teacher model feature distance and the student model feature distance corresponding to any two training samples in each batch; and an optimization module configured to optimize the model parameters of the student model based on the center feature loss corresponding to each batch and the feature distance alignment loss corresponding to any two training samples in each batch, so as to complete the knowledge distillation from the teacher model to the student model.
[0008] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0009] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.
[0010] Compared with the prior art, the disclosed embodiment has the following advantages: because the disclosed embodiment obtains training data, inputs multiple training samples in the training data into the teacher model and the student model respectively in batches, and outputs the teacher model features and the student model features of each training sample in each batch, wherein the training data is an image of the detection object; respectively calculates the teacher model center features and the student model center features corresponding to all the training samples in each batch; calculates the center feature loss between the teacher model center features and the student model center features corresponding to each batch; respectively calculates the teacher model feature distance and the student model feature distance corresponding to any two training samples in each batch; calculates the feature distance alignment loss between the teacher model feature distance and the student model feature distance corresponding to any two training samples in each batch; optimizes the model parameters of the student model based on the center feature loss corresponding to each batch and the feature distance alignment loss corresponding to any two training samples in each batch, so as to complete the knowledge distillation from the teacher model to the student model. The above technical means can solve the problem of low efficiency and poor effect of the knowledge distillation algorithm in the prior art, thereby improving the efficiency of knowledge distillation and enhancing the effect of knowledge distillation. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] FIG1 is a schematic diagram of a process of a knowledge distillation method based on spatial distance alignment provided by an embodiment of the present disclosure;
[0013] FIG2 is a schematic diagram of a flow chart of another knowledge distillation method based on spatial distance alignment provided by an embodiment of the present disclosure;
[0014] FIG3 is a schematic structural diagram of a knowledge distillation device based on spatial distance alignment provided by an embodiment of the present disclosure;
[0015] FIG4 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0016] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present disclosure with unnecessary detail.
[0017] A knowledge distillation method and apparatus based on spatial distance alignment according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0018] FIG1 is a flow chart of a knowledge distillation method based on spatial distance alignment provided by an embodiment of the present disclosure. The knowledge distillation method based on spatial distance alignment in FIG1 can be executed by a computer or server, or software on the computer or server. As shown in FIG1 , the knowledge distillation method based on spatial distance alignment includes:
[0019] S101, obtaining training data, inputting multiple training samples in the training data into the teacher model and the student model in batches, and outputting the teacher model features and the student model features of each training sample in each batch;
[0020] S102, respectively calculating the teacher model central features and the student model central features corresponding to the teacher model features and the student model features of all training samples in each batch;
[0021] S103, calculating the central feature loss between the central features of the teacher model and the central features of the student model corresponding to each batch;
[0022] S104, respectively calculating the teacher model feature distance and the student model feature distance corresponding to the teacher model feature and the student model feature of any two training samples in each batch;
[0023] S105, calculating the feature distance alignment loss between the teacher model feature distance and the student model feature distance corresponding to any two training samples in each batch;
[0024] S106: Optimize the model parameters of the student model based on the central feature loss corresponding to each batch and the feature distance alignment loss corresponding to any two training samples in each batch to complete the knowledge distillation from the teacher model to the student model.
[0025] The embodiments of the present disclosure can be applied to the field of target detection, such as face recognition. Both the teacher model and the student model are face recognition models. The difference is that the teacher model is a trained model and the student model is a model to be trained; the training data includes facial images of multiple people. There can also be differences between the teacher model and the student model, for example, the teacher model is heavyweight and the student model is lightweight. The teacher model and the student model can be the same type of face recognition model or different types. There are many common face recognition models, and the face recognition model used in the embodiments of the present disclosure can be any common face recognition model, such as a deep convolutional neural network.
[0026] It should be noted that during training, the training samples in the training data are divided into multiple batches, and one batch of training samples is used to train the student model each time. The number of training samples in a batch is the batch size, and the batch size can be set by yourself.
[0027] According to the technical solution provided by the embodiment of the present disclosure, training data is obtained, multiple training samples in the training data are input into the teacher model and the student model respectively according to batches, and the teacher model features and student model features of each training sample in each batch are output; the teacher model central features and student model central features corresponding to the teacher model features and student model features of all training samples in each batch are calculated respectively; the central feature loss between the teacher model central features and the student model central features corresponding to each batch is calculated; the teacher model feature distance and the student model feature distance corresponding to the teacher model features and the student model features of any two training samples in each batch are calculated respectively; the feature distance alignment loss between the teacher model feature distance and the student model feature distance corresponding to any two training samples in each batch is calculated; the model parameters of the student model are optimized according to the central feature loss corresponding to each batch and the feature distance alignment loss corresponding to any two training samples in each batch to complete the knowledge distillation from the teacher model to the student model. The above technical means can solve the problem of low efficiency and poor effect of the knowledge distillation algorithm in the existing technology, thereby improving the efficiency of knowledge distillation and enhancing the effect of knowledge distillation.
[0028] Furthermore, multiple training samples in the training data are input into the teacher model and the student model respectively according to batches, and the teacher model features and the student model features of each training sample in each batch are output, including: inputting multiple training samples in the training data into the teacher model according to batches, and outputting the teacher model features of each training sample in each batch through the second-to-last layer network in the teacher model; inputting multiple training samples in the training data into the student model according to batches, and outputting the student model features of each training sample in each batch through the second-to-last layer network in the student model.
[0029] For example, if both the teacher model and the student model are face recognition models, and the last network layer in the face recognition model is a classification layer, the teacher model features for each training sample in each batch are output through the network layer above the classification layer in the teacher model, and the student model features for each training sample in each batch are output through the network layer above the classification layer in the student model.
[0030] Furthermore, the teacher model central features and student model central features corresponding to the teacher model features and student model features of all training samples in each batch are calculated respectively, including: averaging the teacher model features of all training samples in each batch to obtain the teacher model central features corresponding to the teacher model features of all training samples in each batch; averaging the student model features of all training samples in each batch to obtain the student model central features corresponding to the student model features of all training samples in each batch.
[0031] Calculate the average value of the teacher model features of all training samples in a batch, and use the average value as the central feature of the teacher model corresponding to the batch; calculate the average value of the student model features of all training samples in a batch, and use the average value as the central feature of the student model corresponding to the batch.
[0032] Furthermore, the central feature loss between the central features of the teacher model and the central features of the student model corresponding to each batch is calculated, including: calculating the Euclidean distance between the central features of the teacher model and the central features of the student model corresponding to each batch; and using the Euclidean distance corresponding to each batch as the central feature loss corresponding to each batch.
[0033] Furthermore, the teacher model feature distance and student model feature distance corresponding to the teacher model features and student model features of any two training samples in each batch are calculated respectively, including: calculating the Euclidean distance between the teacher model features of any two training samples in each batch, and taking the Euclidean distance between the teacher model features of any two training samples in each batch as the teacher model feature distance corresponding to any two training samples in each batch; calculating the Euclidean distance between the student model features of any two training samples in each batch, and taking the Euclidean distance between the student model features of any two training samples in each batch as the student model feature distance corresponding to any two training samples in each batch.
[0034] In practice, it is found that the relationship between the features output by the teacher model is correlated with the relationship between the features output by the student model. Therefore, the disclosed embodiment finds the relationship between the features output by the teacher model by calculating the Euclidean distance between any two features output by the teacher model, finds the relationship between the features output by the student model by calculating the Euclidean distance between any two features output by the student model, and finally uses the relationship between the features output by the teacher model to constrain the relationship between the features output by the student model, thereby improving the efficiency and effectiveness of knowledge distillation.
[0035] Furthermore, the feature distance alignment loss between the feature distance of the teacher model and the feature distance of the student model corresponding to any two training samples in each batch is calculated, including: calculating the mean square error between the feature distance of the teacher model and the feature distance of the student model corresponding to any two training samples in each batch; and using the mean square error corresponding to any two training samples in each batch as the feature distance alignment loss corresponding to any two training samples in each batch.
[0036] Calculate the mean square error between the teacher model feature distance and the student model feature distance corresponding to any two training samples in a batch, and use the mean square error as the feature distance alignment loss corresponding to the any two training samples in the batch.
[0037] Furthermore, after respectively calculating the teacher model feature distance and the student model feature distance corresponding to the teacher model features and the student model features of any two training samples in each batch, the method also includes: determining the teacher model feature distance vector corresponding to each batch according to the teacher model feature distance corresponding to any two training samples in each batch; determining the student model feature distance vector corresponding to each batch according to the student model feature distance corresponding to any two training samples in each batch; calculating the Euclidean distance between the teacher model feature distance vector and the student model feature distance vector corresponding to each batch, and using the Euclidean distance between the teacher model feature distance vector and the student model feature distance vector corresponding to each batch as the feature distance alignment loss corresponding to each batch; optimizing the model parameters of the student model based on the central feature loss and the feature distance alignment loss corresponding to each batch, so as to complete the knowledge distillation from the teacher model to the student model.
[0038] For example, if a batch has 10 training samples, then there are 45 possible combinations of the 10 training samples in a batch. A batch corresponds to 45 teacher model feature distances. Concatenating these 45 teacher model feature distances yields the teacher model feature distance vector for the batch. A batch corresponds to 45 student model feature distances. Concatenating these 45 student model feature distances yields the student model feature distance vector for the batch.
[0039] FIG2 is a schematic diagram of another knowledge distillation method based on spatial distance alignment provided by an embodiment of the present disclosure. As shown in FIG2 , the method includes:
[0040] S201, calculating the Euclidean distance between the teacher model features and the student model features of each training sample in each batch, and using the Euclidean distance corresponding to each training sample in each batch as the sample feature loss corresponding to each training sample in each batch;
[0041] S202, complete the knowledge distillation from the teacher model to the student model by multi-stage training of the student model:
[0042] S203, performing a first phase training on the student model: optimizing the model parameters of the student model based on the sample feature loss corresponding to each training sample in each batch, and ending the first phase training when the accuracy of the student model is greater than a first threshold;
[0043] S204, performing a second phase of training on the student model: optimizing the model parameters of the student model based on the central feature loss corresponding to each batch, and ending the second phase of training when the accuracy of the student model is greater than a second threshold;
[0044] S205, performing the third stage training on the student model: optimizing the model parameters of the student model based on the feature distance alignment loss corresponding to any two training samples in each batch, and ending the third stage training when the accuracy of the student model is greater than the third threshold.
[0045] The thresholds are, from small to large, the first threshold, the second threshold, and the third threshold.
[0046] In some embodiments, after the third stage of training is completed, it includes: calculating the Euclidean distance between the teacher model features of any two training samples in any two batches, and using the Euclidean distance between the teacher model features of any two training samples in any two batches as the teacher model feature distance corresponding to any two training samples in any two batches; calculating the Euclidean distance between the student model features of any two training samples in any two batches, and using the Euclidean distance between the student model features of any two training samples in any two batches as the teacher model feature distance corresponding to any two training samples in any two batches; calculating the mean square error between the teacher model feature distance and the student model feature distance corresponding to any two training samples in any two batches; using the mean square error corresponding to any two training samples in any two batches as the feature distance alignment loss corresponding to any two training samples in any two batches; performing the fourth stage of training on the student model: optimizing the model parameters of the student model based on the feature distance alignment loss corresponding to any two training samples in any two batches, and ending the fourth stage of training when the accuracy of the student model is greater than the fourth threshold.
[0047] The fourth threshold is greater than the third threshold.
[0048] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.
[0049] The following are embodiments of the apparatus disclosed herein, which can be used to implement the method embodiments disclosed herein. For details not disclosed in the apparatus embodiments disclosed herein, please refer to the method embodiments disclosed herein.
[0050] FIG3 is a schematic diagram of a knowledge distillation device based on spatial distance alignment provided by an embodiment of the present disclosure. As shown in FIG3 , the knowledge distillation device based on spatial distance alignment includes:
[0051] An acquisition module 301 is configured to acquire training data, input multiple training samples in the training data into the teacher model and the student model in batches, and output the teacher model features and the student model features of each training sample in each batch;
[0052] The first calculation module 302 is configured to calculate the teacher model central features and the student model central features corresponding to the teacher model features and the student model features of all training samples in each batch;
[0053] The second calculation module 303 is configured to calculate the central feature loss between the central features of the teacher model and the central features of the student model corresponding to each batch;
[0054] The third calculation module 304 is configured to respectively calculate the teacher model feature distance and the student model feature distance corresponding to the teacher model feature and the student model feature of any two training samples in each batch;
[0055] The fourth calculation module 305 is configured to calculate the feature distance alignment loss between the teacher model feature distance and the student model feature distance corresponding to any two training samples in each batch;
[0056] The optimization module 306 is configured to optimize the model parameters of the student model based on the central feature loss corresponding to each batch and the feature distance alignment loss corresponding to any two training samples in each batch to complete the knowledge distillation from the teacher model to the student model.
[0057] According to the technical solution provided by the embodiment of the present disclosure, training data is obtained, multiple training samples in the training data are input into the teacher model and the student model respectively according to batches, and the teacher model features and student model features of each training sample in each batch are output; the teacher model central features and student model central features corresponding to the teacher model features and student model features of all training samples in each batch are calculated respectively; the central feature loss between the teacher model central features and the student model central features corresponding to each batch is calculated; the teacher model feature distance and the student model feature distance corresponding to the teacher model features and the student model features of any two training samples in each batch are calculated respectively; the feature distance alignment loss between the teacher model feature distance and the student model feature distance corresponding to any two training samples in each batch is calculated; the model parameters of the student model are optimized according to the central feature loss corresponding to each batch and the feature distance alignment loss corresponding to any two training samples in each batch to complete the knowledge distillation from the teacher model to the student model. The above technical means can solve the problem of low efficiency and poor effect of the knowledge distillation algorithm in the existing technology, thereby improving the efficiency of knowledge distillation and enhancing the effect of knowledge distillation.
[0058] In some embodiments, the acquisition module 301 is further configured to input multiple training samples in the training data into the teacher model in batches, and output the teacher model features of each training sample in each batch through the penultimate network layer in the teacher model; input multiple training samples in the training data into the student model in batches, and output the student model features of each training sample in each batch through the penultimate network layer in the student model.
[0059] In some embodiments, the first computing module 302 is further configured to average the teacher model features of all training samples in each batch to obtain the teacher model central features corresponding to the teacher model features of all training samples in each batch; and average the student model features of all training samples in each batch to obtain the student model central features corresponding to the student model features of all training samples in each batch.
[0060] In some embodiments, the second calculation module 303 is further configured to calculate the Euclidean distance between the central features of the teacher model and the central features of the student model corresponding to each batch; and use the Euclidean distance corresponding to each batch as the central feature loss corresponding to each batch.
[0061] In some embodiments, the third calculation module 304 is further configured to calculate the Euclidean distance between the teacher model features of any two training samples in each batch, and use the Euclidean distance between the teacher model features of any two training samples in each batch as the teacher model feature distance corresponding to any two training samples in each batch; calculate the Euclidean distance between the student model features of any two training samples in each batch, and use the Euclidean distance between the student model features of any two training samples in each batch as the student model feature distance corresponding to any two training samples in each batch.
[0062] In some embodiments, the fourth calculation module 305 is further configured to calculate the mean square error between the teacher model feature distance and the student model feature distance corresponding to any two training samples in each batch; and use the mean square error corresponding to any two training samples in each batch as the feature distance alignment loss corresponding to any two training samples in each batch.
[0063] In some embodiments, the optimization module 306 is further configured to determine the teacher model feature distance vector corresponding to each batch based on the teacher model feature distance corresponding to any two training samples in each batch; determine the student model feature distance vector corresponding to each batch based on the student model feature distance corresponding to any two training samples in each batch; calculate the Euclidean distance between the teacher model feature distance vector and the student model feature distance vector corresponding to each batch, and use the Euclidean distance between the teacher model feature distance vector and the student model feature distance vector corresponding to each batch as the feature distance alignment loss corresponding to each batch; optimize the model parameters of the student model based on the central feature loss and feature distance alignment loss corresponding to each batch to complete the knowledge distillation from the teacher model to the student model.
[0064] In some embodiments, the optimization module 306 is further configured to calculate the Euclidean distance between the teacher model features and the student model features of each training sample in each batch, and use the Euclidean distance corresponding to each training sample in each batch as the sample feature loss corresponding to each training sample in each batch; complete the knowledge distillation from the teacher model to the student model by performing multi-stage training on the student model: perform the first stage training on the student model: optimize the model parameters of the student model based on the sample feature loss corresponding to each training sample in each batch, and end the first stage training when the accuracy of the student model is greater than the first threshold; perform the second stage training on the student model: optimize the model parameters of the student model based on the central feature loss corresponding to each batch, and end the second stage training when the accuracy of the student model is greater than the second threshold; perform the third stage training on the student model: optimize the model parameters of the student model based on the feature distance alignment loss corresponding to any two training samples in each batch, and end the third stage training when the accuracy of the student model is greater than the third threshold.
[0065] In some embodiments, the optimization module 306 is further configured to calculate the Euclidean distance between the teacher model features of any two training samples in any two batches, and use the Euclidean distance between the teacher model features of any two training samples in any two batches as the teacher model feature distance corresponding to any two training samples in any two batches; calculate the Euclidean distance between the student model features of any two training samples in any two batches, and use the Euclidean distance between the student model features of any two training samples in any two batches as the teacher model feature distance corresponding to any two training samples in any two batches; calculate the mean square error between the teacher model feature distance and the student model feature distance corresponding to any two training samples in any two batches; use the mean square error corresponding to any two training samples in any two batches as the feature distance alignment loss corresponding to any two training samples in any two batches; perform fourth-stage training on the student model: optimize the model parameters of the student model based on the feature distance alignment loss corresponding to any two training samples in any two batches, and end the fourth-stage training when the accuracy of the student model is greater than the fourth threshold.
[0066] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure.
[0067] Figure 4 is a schematic diagram of an electronic device 4 provided in an embodiment of the present disclosure. As shown in Figure 4, electronic device 4 in this embodiment includes a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable by the processor 401. When the processor 401 executes the computer program 403, the steps in the aforementioned method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of the modules / units in the aforementioned apparatus embodiments are implemented.
[0068] Electronic device 4 can be a desktop computer, laptop, PDA, cloud server, or other electronic device. Electronic device 4 can include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art will appreciate that FIG4 is merely an example of electronic device 4 and does not limit the scope of electronic device 4. Electronic device 4 can include more or fewer components than shown, or different components.
[0069] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0070] Memory 402 can be an internal storage unit of electronic device 4, such as a hard disk or memory of electronic device 4. Memory 402 can also be an external storage device of electronic device 4, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on electronic device 4. Memory 402 can also include both an internal storage unit of electronic device 4 and an external storage device. Memory 402 is used to store computer programs and other programs and data required by the electronic device.
[0071] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0072] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present disclosure implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. The computer program may include computer program code, which may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0073] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure, and should all be included in the scope of protection of the present disclosure.
Claims
1. A knowledge distillation method based on spatial distance alignment, applied to the field of object detection, characterized in that, Including: Obtain training data, input multiple training samples in the training data into a teacher model and a student model respectively by batches, and output the teacher model features and student model features of each training sample in each batch, where the training data is an image of a detection object; Calculate the teacher model center features and student model center features corresponding to the teacher model features and student model features of all training samples in each batch respectively; Calculate the center feature loss between the teacher model center features and student model center features corresponding to each batch; Calculate the teacher model feature distances and student model feature distances corresponding to the teacher model features and student model features of any two training samples in each batch respectively; Calculate the feature distance alignment loss between the teacher model feature distances and student model feature distances corresponding to any two training samples in each batch; Optimize the model parameters of the student model according to the center feature loss corresponding to each batch and the feature distance alignment loss corresponding to any two training samples in each batch, so as to complete the knowledge distillation from the teacher model to the student model.
2. The method according to claim 1, wherein Inputting multiple training samples in the training data into a teacher model and a student model respectively by batches, and outputting the teacher model features and student model features of each training sample in each batch, including: Input multiple training samples in the training data into the teacher model by batches, and output the teacher model features of each training sample in each batch through the penultimate layer network in the teacher model; Input multiple training samples in the training data into the student model by batches, and output the student model features of each training sample in each batch through the penultimate layer network in the student model.
3. The method according to claim 1, wherein Calculate the teacher model center features and student model center features corresponding to the teacher model features and student model features of all training samples in each batch respectively, including: Average the teacher model features of all training samples in each batch to obtain the teacher model center features corresponding to the teacher model features of all training samples in each batch; Average the student model features of all training samples in each batch to obtain the student model center features corresponding to the student model features of all training samples in each batch.
4. The method according to claim 1, wherein Calculate the center feature loss between the teacher model center features and student model center features corresponding to each batch, including: Calculate the Euclidean distance between the teacher model center features and student model center features corresponding to each batch; Take the Euclidean distance corresponding to each batch as the center feature loss corresponding to each batch.
5. The method according to claim 1, characterized in that, Calculate the teacher model feature distances and student model feature distances corresponding to the teacher model features and student model features of any two training samples in each batch respectively, including: Calculate the Euclidean distance between the teacher model features of any two training samples in each batch, and take the Euclidean distance between the teacher model features of any two training samples in each batch as the teacher model feature distance corresponding to any two training samples in each batch; Calculate the Euclidean distance between the student model features of any two training samples in each batch, and use the Euclidean distance between the student model features of any two training samples in each batch as the student model feature distance corresponding to any two training samples in each batch.
6. The method according to claim 1, wherein Calculate the feature distance alignment loss between the teacher model feature distance and the student model feature distance corresponding to any two training samples in each batch, including: Calculate the mean square error between the teacher model feature distance and the student model feature distance corresponding to any two training samples in each batch; Use the mean square error corresponding to any two training samples in each batch as the feature distance alignment loss corresponding to any two training samples in each batch.
7. The method according to claim 1, wherein After calculating the teacher model feature distance and the student model feature distance corresponding to the teacher model features and the student model features of any two training samples in each batch respectively, the method further includes: Determine the teacher model feature distance vector corresponding to each batch according to the teacher model feature distance corresponding to any two training samples in each batch; Determine the student model feature distance vector corresponding to each batch according to the student model feature distance corresponding to any two training samples in each batch; Calculate the Euclidean distance between the teacher model feature distance vector and the student model feature distance vector corresponding to each batch, and use the Euclidean distance between the teacher model feature distance vector and the student model feature distance vector corresponding to each batch as the feature distance alignment loss corresponding to each batch; Optimize the model parameters of the student model according to the center feature loss and the feature distance alignment loss corresponding to each batch to complete the knowledge distillation from the teacher model to the student model.
8. A knowledge distillation device based on spatial distance alignment, which is applied to the field of object detection, is characterized in that Including: An acquisition module, configured to acquire training data, input multiple training samples in the training data into the teacher model and the student model batch by batch, and output the teacher model features and the student model features of each training sample in each batch, where the training data is an image of a detection object; A first calculation module, configured to calculate the teacher model center features and the student model center features corresponding to the teacher model features and the student model features of all training samples in each batch respectively; A second calculation module, configured to calculate the center feature loss between the teacher model center feature and the student model center feature corresponding to each batch; A third calculation module, configured to calculate the teacher model feature distance and the student model feature distance corresponding to the teacher model features and the student model features of any two training samples in each batch respectively; A fourth calculation module, configured to calculate the feature distance alignment loss between the teacher model feature distance and the student model feature distance corresponding to any two training samples in each batch; An optimization module, configured to optimize the model parameters of the student model according to the center feature loss corresponding to each batch and the feature distance alignment loss corresponding to any two training samples in each batch to complete the knowledge distillation from the teacher model to the student model.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method described in claim 1 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to claim 1.
Citation Information
Patent Citations
Model training method, face recognition method and device
CN114611672A
Long-tail remote sensing image target identification method based on dynamic relation distillation
CN115272881A
Model training method and device, equipment and storage medium
CN116976428A
Knowledge distillation method and device based on spatial distance alignment
CN117474037A
Learning efficient object detection models with knowledge distillation
US20180268292A1