A multi-teacher distillation method for neural network architecture search

CN122635486BActive Publication Date: 2026-09-25WUXI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611125602.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-25
Estimated Expiration
2046-07-28

AI Technical Summary

Technical Problem

单一教师模型或高精度教师模型直接参与蒸馏时,教师知识难以匹配学生模型的学习能力,导致训练不稳定或负迁移

Benefits of technology

本申请提高NAS学生模型的知识吸收能力和分类性能:本申请围绕神经架构搜索得到的NAS学生模型开展教师选择、蒸馏训练和校准训练,使教师模型知识能够更加匹配学生模型的结构特点和学习能力。通过适配教师组合的基础蒸馏、中间神经网络结构单元表征引导以及高性能教师模型校准训练,学生模型能够更充分地学习教师模型中的类别关系和暗知识,从而提高最终分类性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122635486B_ABST
    Figure CN122635486B_ABST
Patent Text Reader

Abstract

The application belongs to the field of artificial intelligence, and discloses a multi-teacher distillation method for neural network architecture search, which comprises the following steps: step 1, determining an image classification dataset used for model training and evaluation, constructing a student model, a teacher pool, and screening a teacher model from the teacher pool; step 2, calculating a comprehensive score of a teacher combination and selecting an optimal teacher combination in combination with the adaptability of the teacher combination and the complementary dark knowledge of the teacher combination; step 3, based on the optimal teacher combination, performing student model distillation training based on a unit-level knowledge support; and step 4, performing high-performance teacher model calibration training on the distilled student model to obtain final student model parameters and a student model. The application can improve the knowledge absorption capacity and classification performance of the student model without changing the structure of the NAS student model and increasing the final inference complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence, specifically relating to a multi-teacher distillation method for searching neural network architectures. Background Technology

[0002] With the widespread application of deep neural networks in tasks such as image classification, large-scale models can typically achieve high classification accuracy. However, their large number of parameters and high computational complexity make them difficult to deploy directly on resource-constrained devices or in real-time inference scenarios. To reduce model deployment costs, knowledge distillation is often used to transfer knowledge from large models or teacher models to lightweight student models, enabling student models to achieve good classification performance with lower computational overhead. Meanwhile, neural network architecture search can automatically search for lightweight network structures suitable for the target task. NAS student models obtained through neural network architecture search typically feature automated structure, controllable scale, and flexible deployment.

[0003] Current knowledge distillation methods for NAS student models still have significant shortcomings, mainly in the following aspects: 1) After the student model structure is fixed, there is a tendency for insufficient adaptation between the teacher model's knowledge and the student model. After the NAS student model search is completed, its structure and capacity are basically determined, and the representation capabilities of different neural network structural units vary. When a single teacher model or a high-precision teacher model directly participates in distillation, the teacher's knowledge is difficult to match the learning ability of the student model, leading to unstable training or negative transfer. 2) Teacher knowledge mainly affects the final output layer of the student model and is difficult to fully guide the representation of intermediate neural network structural units in the NAS student model. For NAS student models formed by stacking multiple neural network structural units, relying solely on the final class prediction score (logits) for distillation will result in a long knowledge transfer path, making it difficult for intermediate neural network structural units to directly obtain guidance from the teacher's soft knowledge. 3) The timing of high-performance teacher models participating in distillation training lacks control. High-performance teacher models have complex knowledge; if they directly participate in distillation in the early stages of student training, the unstable student representation and the large gap in teacher and student capabilities will lead to unstable training. Summary of the Invention

[0004] To address the aforementioned technical issues, this application provides a multi-teacher distillation method for neural network architecture search. This method uses a NAS student model as the optimization target. Based on a fixed student model structure, a pre-trained teacher pool is constructed. The method comprehensively evaluates the class discrimination ability of the teacher models, the fit between teacher and student representations, and the complementary relationship of hidden knowledge among teachers to select the teacher combination most suitable for the current NAS student model's learning. Simultaneously, a training-period knowledge scaffold is temporarily connected after the later neural network structural units in the NAS student model, allowing the soft label supervision provided by the teacher combination to directly affect the intermediate representations of the student model. The strength of this intermediate supervision is controlled at different training stages through weight scheduling. After the student model completes knowledge scaffold distillation, the student model with the best performance on the validation set is selected for initialization, and a high-performance teacher model is introduced for calibration training. This method can improve the knowledge absorption capacity and classification performance of the student model without changing the NAS student model structure or increasing the final inference complexity.

[0005] To achieve the above objectives, this application is implemented through the following technical solution.

[0006] This application presents a multi-teacher distillation method for searching neural network architectures, which specifically includes the following steps: Step 1, Data Preparation, Basic Model and Teacher Pool Initialization: Determine the image classification dataset to be used for model training and evaluation, construct student models and teacher pools, and select teacher models from the teacher pools; Step 2: Select the optimal teacher combination: Based on the student model and teacher pool constructed in Step 1, calculate the single teacher model fit score, calculate the teacher group fit score, calculate the complementarity of the two teacher models in the teacher combination, calculate the teacher combination implicit knowledge complementarity by combining the complementarity and the single teacher model fit score, calculate the teacher combination comprehensive score by combining the teacher group fit score and the teacher combination implicit knowledge complementarity, and select the optimal teacher combination. Step 3: Based on the optimal teacher combination selected in Step 2, conduct student model distillation training based on unit-level knowledge scaffolding; Step 4: Perform high-performance teacher model calibration training on the distilled student model to obtain the final student model parameters. ; Step 5: Determine the final student model: Based on the final student model parameters obtained in Step 4. The final student model is obtained as follows .

[0007] A further improvement in this application is that step 1 specifically includes the following steps: Step 1.1, Dataset Preparation: Given an image classification dataset and the image classification dataset It is divided into training set, validation set and test set, where, Indicates the first One input image sample, Indicates the first The true class label corresponding to each input image sample This represents the total number of image samples.

[0008] Step 1.2: Constructing the student model: Obtaining the student model The student model include The nth sequentially stacked neural network structural units, the nth Each neural network structural unit is denoted as . ,in The transfer process of intermediate features in the student model is represented as follows: , Indicates the first The intermediate features output by each neural network structural unit The intermediate features represented by the output of the previous neural network structural unit are denoted as logits, and the final main classification output of the student model is denoted as... ,in, This represents the trainable parameters of the student model. This represents the input image sample.

[0009] Step 1.3, Teacher Pool and High-Performance Teacher Model Initialization: Establishing a Pre-trained Teacher Pool , This represents the total number of teacher models in the teacher pool, for each input image sample. , No. Teacher Model The output logits is denoted as Meanwhile, a high-performance teacher model will be prepared separately. This high-performance teacher model The classification accuracy on the CIFAR-100 dataset meets the requirements. , Not part of the teacher pool .

[0010] Step 1.4, Teacher Combination Settings: Set up a pool of teachers. Select size as Candidate teacher combination , This represents the teacher model selected from the teacher pool.

[0011] A further improvement in this application is that, in step 2, calculating the single-teacher model fit score specifically includes the following steps: Step 2.1.1: Student model matches the real category The predicted probability is The teacher model is for the real category The predicted probability is , No. Teacher Model For the real category Predictive advantage is defined as: , in, Indicates the first Teacher Model The predictive advantage of the current student model over the true category, if the teacher model... Predicted probability on a given input image sample If the input image sample contribution is lower than that of the student model, the contribution is truncated to 0.

[0012] Step 2.1.2: Let the feature matrix output of the intermediate layer corresponding to the student model be... The feature matrix output by the intermediate layer of the teacher model is: Then use CKA to represent and Feature compatibility between: , in, This represents the feature matrix of the student model after centering. , Represents the feature matrix of the student model The mean vector calculated along the sample dimension This represents the feature matrix of the teacher model after centralization. , Represents the feature matrix of the teacher model The mean vector calculated along the sample dimension Student model feature matrix The transpose of the matrix, Teacher model feature matrix The transpose of .

[0013] Step 2.1.3, place the first Teacher Model For the real category Predictive advantage and feature compatibility Combined, the single-teacher model fit score is obtained. .

[0014] A further improvement in this application is that, in step 2, the teacher group suitability is calculated based on the single-teacher model fit score, specifically including the following steps: Step 2.2.1: For candidate teacher combinations According to the candidate teacher combination Each single-teacher model adaptation score Calculate teacher weights : ; in, Indicates the first Teacher Model In the candidate teacher combination Weight within, Indicates teacher weight Calculated temperature parameters, Indicates candidate teacher combination The Middle Teacher Model The fit score.

[0015] Step 2.2.2: Based on teacher weights Calculate the suitability of teacher groups: ; in, Indicates candidate teacher combination Adaptability to student models.

[0016] A further improvement in this application is that, in step 2, calculating the complementarity of teachers' combined implicit knowledge specifically includes the following steps: Step 2.3.1: Calculate temperature given dark knowledge. , No. Teacher Model Compared to the student model in the first Input image samples The hidden knowledge residuals provided above: ; in, Calculating temperature for student models using dark knowledge The soft probability distribution under the following conditions For the teacher model, temperature is calculated using hidden knowledge. Under the soft probability distribution, to avoid overlap between the real class supervision and the cross-entropy objective, the real class position is set to zero. , Only retain the dark knowledge residuals on non-target categories.

[0017] Step 2.3.2: For candidate teacher combinations The first in Teacher Model and the Teacher Model Teacher model Teacher Model After expanding the dark knowledge residuals on non-target categories into vectors, the teacher model is calculated. Teacher Model Complementarity between them: ; in, Teacher model The non-target dark knowledge residual vector Teacher model The non-target dark knowledge residual vector This represents the cosine similarity.

[0018] Step 2.3.3: Based on the teacher model Teacher Model complementarity between Combined with single-teacher model adaptation scores To obtain the teacher model Teacher Model Weighted complementarity of students' tacit knowledge : ; in, Teacher model The fit score for the student model, Teacher model Fit scores for the student model.

[0019] Step 2.3.4: Weighted complementarity based on dark knowledge Calculate the complementary nature of teacher combinations' implicit knowledge for: ; in, For candidate teacher combinations All teachers in the group form a set.

[0020] A further improvement in this application is that, in step 2, the comprehensive score of the teacher combination is calculated by combining the suitability of the teacher group and the complementarity of the teacher combination's implicit knowledge, and the optimal teacher combination is selected. This specifically includes the following steps: Step 2.4.1: For candidate teacher combinations Teacher combination composite score Defined as: ; in, Indicates candidate teacher combination Adaptability to the student model This indicates the complementary nature of the teachers' implicit knowledge. This represents the coefficient for enhancing the complementarity of dark knowledge.

[0021] Step 2.4.2: Select the teacher combination with the highest overall score. .

[0022] A further improvement in this application is that step 3 specifically includes the following steps: Step 3.1: Temporarily connect the training period knowledge scaffold header after the later neural network structure units in the student model, so that the teacher's knowledge enters the intermediate features of the student model. Let the student model have a total of The insertion position of the knowledge scaffold head during the training phase is defined as follows: (The original text contains several typographical errors and inconsistencies. A more accurate translation would require the full context.) ; in, This indicates the sequence number of the neural network structure unit inserted into the knowledge scaffold during the training period. This indicates the relative depth position of the knowledge support head during the training period.

[0023] Step 3.2, Constructing the Knowledge Scaffold Header for the Training Period: In the... Intermediate features output by each neural network structural unit A lightweight training branch is then added to process the intermediate features output by the neural network architecture units. Mapped to category logits, and subject to soft-label distillation supervision by a combination of teachers: ; in, This indicates the knowledge support head during the training period. This represents the intermediate logits output by the knowledge scaffold head during the training period.

[0024] Step 3.3: Constructing the Distillation Training Objective: For the selected teacher combination, choose the teacher combination with the highest overall score. The main output layer of the student model receives soft-label distillation supervision from the teacher combination, resulting in the main output distillation loss. Meanwhile, the intermediate logits output by the knowledge scaffold head during the training period That is, the knowledge scaffold head during the training period is paired with intermediate features. The class prediction logits obtained after mapping are subjected to soft-label distillation supervision by the same teacher combination, resulting in the knowledge scaffold distillation loss. The total loss during distillation training is defined as: ; in, This represents the true label cross-entropy loss of the student model's main output. This indicates the distillation loss in the main output layer; This indicates the knowledge scaffold distillation loss; This indicates the weight corresponding to the distillation loss of the main output layer. This represents the weight of the knowledge scaffold distillation loss.

[0025] Step 3.4: Set up knowledge scaffold distillation loss weight scheduling: Let the first... The knowledge scaffold distillation loss weights for each epoch. Three-stage scheduling is adopted: ; ; ; in, Indicates the current training round. This indicates the end of the warm-up training round. Indicates the cycle in which decay begins. This indicates the end of the distillation phase training round. This represents the basic weights of knowledge scaffold distillation.

[0026] Step 3.5: Complete distillation training and remove the knowledge scaffold head during training: optimize the total loss. The student model parameters were obtained after adaptive teacher combination distillation and unit-level knowledge scaffolding training. During this training process, the knowledge scaffolding head during the training period Soft-label distillation supervision is used to receive and adapt the combined output of teachers, enabling teacher model knowledge to be applied in advance to the representations of later neural network structural units in the student model. After distillation training is completed, the training-period knowledge scaffold header is deleted. Only the feature extraction backbone and final classifier of the NAS student model are retained.

[0027] A further improvement in this application is that step 4 specifically includes the following steps: Step 4.1: Determine the initial student model: During the distillation training phase, the student model will acquire a series of parameters at different epochs. After each epoch, the student model is evaluated using the validation set to obtain the validation set accuracy. The student model with the best performance on the validation set during the distillation stage is selected as the initial student model. This initializes the student model. As a high-performance teacher model Initialization parameters during calibration training phase .

[0028] Step 4.2: Construct a high-performance teacher model and calibrate the training loss: Initialize the student model as determined in Step 4.1. Based on this, high-performance teacher model knowledge is introduced in a low-intensity manner, i.e., from the high-performance teacher model The output knowledge information is used to guide the student model calibration training, enabling the high-performance teacher model to serve as a later calibration signal rather than re-dominant student training. The total loss during the high-performance teacher model calibration training phase is defined as: ; in, This represents the distillation loss of the high-performance teacher model; This represents the distillation weights of the high-performance teacher model. The student model's main output represents the true label cross-entropy loss, while the high-performance teacher model's distillation loss employs a high-temperature soft-label distillation method. , in, This indicates the temperature during the calibration training phase. This indicates the temperature of the high-performance teacher model during the calibration and training phase. The distribution of soft labels below, This indicates the temperature of the student model during the calibration and training phase. The distribution of soft labels below; The refinement stage meets the following requirements: , This indicates the temperature during the basic distillation stage. This indicates the weight corresponding to the distillation loss of the main output layer in step 3. This represents the learning rate of the student model distillation training in step 3, which is used to control the parameter update magnitude during the overall distillation training process.

[0029] Step 4.3: Complete the high-performance teacher model calibration training: Use the high-performance teacher model to calibrate the total loss during the training phase. The student model is trained over a short period of time to obtain the final student model parameters. .

[0030] The beneficial effects of this application are: This application improves the knowledge absorption capacity and classification performance of the NAS student model: It focuses on teacher selection, distillation training, and calibration training of the NAS student model obtained through neural architecture search, enabling the teacher model's knowledge to better match the structural characteristics and learning ability of the student model. Through basic distillation adapted to the teacher combination, guidance from intermediate neural network structural unit representations, and calibration training of the high-performance teacher model, the student model can more fully learn the class relationships and hidden knowledge in the teacher model, thereby improving the final classification performance.

[0031] This application improves the stability of distillation training and reduces the cost of ineffective training: Before complete distillation training, this application performs an adaptation evaluation on candidate teacher combinations, reducing the computational cost of blindly selecting teacher models or training different teacher combinations one by one. Simultaneously, the high-performance teacher model is not directly involved in the early stages of training, but rather undergoes low-intensity calibration after the student model has completed the knowledge scaffolding distillation training. This helps reduce training oscillations and the risk of negative transfer caused by excessive gaps in teacher and student abilities.

[0032] This application balances performance enhancement with lightweight deployment requirements: the candidate teacher model, high-performance teacher model, and training-phase knowledge scaffold head are used only during the training phase. In the final deployment phase, only the NAS student model backbone and the final classifier are retained; the training-phase knowledge scaffold head is not retained, and no teacher model is invoked. Therefore, this method can improve student model performance without increasing the number of parameters, computational load, or inference latency of the final model, making it suitable for lightweight image classification model deployment scenarios. Attached Figure Description

[0033] Figure 1 This is a framework diagram of the multi-teacher distillation method in this application.

[0034] Figure 2 This is a schematic diagram of the distillation method of this application.

[0035] Figure 3 This is a flowchart illustrating step 2 of the application process for selecting the optimal teacher combination.

[0036] Figure 4 This is a flowchart illustrating the student model distillation training process in step 3 of this application.

[0037] Figure 5 This is a flowchart illustrating the high-performance teacher model calibration training process in step 4 of this application. Detailed Implementation

[0038] The embodiments of this application will be disclosed below with reference to the drawings. For clarity, many practical details will be described in the following description. However, it should be understood that these practical details should not be used to limit this application. That is, in some embodiments of this application, these practical details are not necessary. In addition, for the sake of simplicity, some conventional structures and components will be shown in the drawings in a simple schematic manner.

[0039] like Figure 1-2 As shown, this application presents a multi-teacher distillation method for neural network architecture search, which specifically includes the following steps:

[0040] Step 1: Data Preparation, Basic Model and Teacher Pool Initialization: Determine the image classification dataset used for model training and evaluation, and specify the input samples, class labels, and number of classes. Construct student models and a teacher pool, and select teacher models from the teacher pool. This includes the following steps: Step 1.1, Dataset Preparation: Given an image classification dataset and the image classification dataset It is divided into training set, validation set and test set, where, Indicates the first One input image sample, Indicates the first The true class label corresponding to each input image sample This represents the total number of image samples; the training set is used for student model training, the validation set is used for teacher combination scoring, model selection, and performance evaluation during training, and the test set is used for the final performance evaluation of student models.

[0041] Step 1.2: Construct the student model: Obtain the student model using DARTS or other neural architecture search methods. The student model include The nth sequentially stacked neural network structural units, the nth Each neural network structural unit is denoted as . ,in The transfer process of intermediate features in the student model is represented as follows: , Indicates the first The intermediate features output by each neural network structural unit The intermediate features represented by the output of the previous neural network structural unit are denoted as logits, and the final main classification output of the student model is denoted as... ,in, This represents the trainable parameters of the student model. This represents the input image sample.

[0042] Step 1.3, Teacher Pool and High-Performance Teacher Model Initialization: Establishing a Pre-trained Teacher Pool , This represents the total number of teacher models in the teacher pool, for each input image sample. , No. Teacher Model The output logits is denoted as Different teacher models have different structures, output distributions, and hidden knowledge representation capabilities, so it is necessary to determine their suitability for the current student model in subsequent stages; at the same time, a high-performance teacher model should be prepared separately. This high-performance teacher model The classification accuracy on the CIFAR-100 dataset meets the requirements. , Not part of the teacher pool .

[0043] Step 1.4, Teacher Combination Settings: Based on task requirements, set the number of teacher models included in each teacher combination, starting from the teacher pool. Select size Candidate teacher combination , This represents the teacher model selected from the teacher pool.

[0044] Step 2: Select the optimal teacher combination: such as Figure 3 As shown, based on the student model and teacher pool constructed in step 1, the single teacher model fit score is calculated, the teacher group fit score is calculated, the complementarity of the two teacher models in the teacher group is calculated, the teacher group implicit knowledge complementarity is calculated by combining the complementarity and the single teacher model fit score, and the teacher group comprehensive score is calculated by combining the teacher group fit score and the teacher group implicit knowledge complementarity, and the optimal teacher group is selected.

[0045] Step 2, calculating the single-teacher model fit score, specifically includes the following steps: Step 2.1.1: Calculate the teacher model Compared to the student model In the real category The predictive advantage of the student model over the true category. The predicted probability is The teacher model is for the real category The predicted probability is , No. Teacher Model For the real category The predictive advantage is defined as: ; in, Indicates the first Teacher Model The predictive advantage of the current student model over the true category, if the teacher model... Predicted probability on a given input image sample If the input image sample contribution is lower than that of the student model, the contribution is truncated to 0, thus retaining only the positive gain of the teacher model relative to the student model.

[0046] Step 2.1.2: Calculate the feature compatibility between the teacher model and the student model. Let the feature matrix output by the intermediate layer corresponding to the student model be... The feature matrix output by the intermediate layer of the teacher model is: Then use CKA to represent and Feature compatibility between: , in, This represents the feature matrix of the student model after centering. , Represents the feature matrix of the student model The mean vector calculated along the sample dimension This represents the feature matrix of the teacher model after centralization. , Represents the feature matrix of the teacher model The mean vector calculated along the sample dimension Student model feature matrix The transpose of the matrix, Teacher model feature matrix The transpose of the matrix; the higher the value, the closer the feature spaces of the teacher model and the student model are, and the easier it is for the student model to learn the knowledge of the teacher model.

[0047] Step 2.1.3, place the first Teacher Model For the real category Predictive advantage and feature compatibility Combined, the single-teacher model fit score is obtained. This score reflects both whether the teacher model is superior to the student model and whether the knowledge of the teacher model is easily learned by the student model, thus avoiding the selection of a teacher model solely based on its own accuracy.

[0048] Step 2 involves calculating the teacher group suitability based on the single-teacher model fit score, specifically including the following steps: Step 2.2.1: Based on the single-teacher model fit score, further evaluate whether a teacher combination as a whole is suitable for the current student model learning. For candidate teacher combinations... According to the candidate teacher combination Each single-teacher model adaptation score Calculate teacher weights : ; in, Indicates the first Teacher Model In the candidate teacher combination Weight within, Indicates teacher weight Calculated temperature parameters, Indicates candidate teacher combination The Middle Teacher Model The fit score.

[0049] Step 2.2.2: Based on teacher weights Calculate the suitability of teacher groups: ; in, Indicates candidate teacher combination Adaptability to student models.

[0050] Through the above calculations, the combined scoring will be more biased towards the teacher models that are truly effective for the student models, rather than simply averaging the contributions of all teacher models within the combination. This can reduce the interference of mismatched teacher models on the combined scoring.

[0051] In step 2, calculating the complementarity of teachers' implicit knowledge includes the following steps: Step 2.3.1: This step determines whether the teacher models in the combination can provide complementary non-target category knowledge, avoiding multiple teacher models providing highly repetitive dark knowledge. Given dark knowledge, calculate the temperature. , No. Teacher Model Compared to the student model in the first Input image samples The hidden knowledge residuals provided above: ; in, Calculating temperature for student models using dark knowledge The soft probability distribution under the following conditions For the teacher model, temperature is calculated using hidden knowledge. Under the soft probability distribution, to avoid overlap between the real class supervision and the cross-entropy objective, the real class position is set to zero. ,so, Only retain the dark knowledge residuals on non-target categories.

[0052] Step 2.3.2: For candidate teacher combinations The first in Teacher Model and the Teacher Model Teacher model Teacher Model After expanding the dark knowledge residuals on non-target categories into vectors, the teacher model is calculated. Teacher Model Complementarity between them: ; in, Teacher model The non-target dark knowledge residual vector Teacher model The non-target dark knowledge residual vector Cosine similarity is expressed as the similarity between two teacher models. and If the directions of the dark knowledge residuals are similar, it indicates that the non-target category information they provide is similar and their complementarity is low; if the directions of the residuals are significantly different, it indicates that they provide different dark knowledge relationships and their complementarity is high.

[0053] Step 2.3.3: Based on the teacher model Teacher Model complementarity between Combined with single-teacher model to adapt scores To obtain the teacher model Teacher Model Weighted complementarity of students' tacit knowledge : ; in, Teacher model The fit score for the student model, Teacher model Fit scores for the student model.

[0054] Step 2.3.4: Weighted complementarity based on dark knowledge Calculate the complementary nature of teacher combinations' implicit knowledge for: ; in, For candidate teacher combinations All teachers in the group form a set.

[0055] Step 2 involves calculating the comprehensive score of teacher combinations and selecting the optimal teacher combination by combining the suitability of teacher group pairings and the complementarity of teachers' implicit knowledge. This includes the following steps: Step 2.4.1: For candidate teacher combinations Teacher combination composite score Defined as: ; in, Indicates candidate teacher combination Adaptability to the student model This indicates the complementary nature of teachers' implicit knowledge. This represents the coefficient for enhancing the complementarity of dark knowledge.

[0056] Step 2.4.2: Select the teacher combination with the highest overall score. .

[0057] This step allows for the selection of teacher combinations before distillation training, reducing the training costs associated with verifying different combinations through individual distillations. Furthermore, because the scoring process considers both teacher model suitability and the complementarity of hidden knowledge, it mitigates the risk of negative transfer caused by blindly selecting high-accuracy teacher models or simply averaging multiple teacher models.

[0058] Step 3: Based on the optimal teacher combination selected in Step 2, conduct student model distillation training using unit-level knowledge scaffolding. For example... Figure 4 As shown, step 3 specifically includes the following steps: Step 3.1: Temporarily connect the training period knowledge scaffold header after the later neural network structure units in the student model, so that the teacher's knowledge enters the intermediate features of the student model. It applies not only to the final classification output, but also to the student model. The insertion position of the knowledge scaffold head during the training phase is defined as follows: (The original text contains several typographical errors and inconsistencies. A more accurate translation would require the full context.) ; in, This indicates the sequence number of the neural network structure unit inserted into the knowledge scaffold during the training period. This indicates the relative depth of the knowledge scaffold head during the training period. Mid-to-late layer neural network structures were chosen as the knowledge scaffold location because this position already possesses a certain semantic expressive power but has not yet reached the final classification layer, making it suitable as an intermediate injection point for the teacher model's soft labels. Compared to shallower layers, mid-to-late layer features are closer to classification semantics; compared to the final output layer, this position allows the teacher model's knowledge to influence the representation learning of subsequent neural network structures earlier.

[0059] Step 3.2, Constructing the Knowledge Scaffold Header for the Training Period: In the... Intermediate features output by each neural network structural unit A lightweight training branch is then added to process the intermediate features output by the neural network architecture units. Mapped to category logits, and subject to soft-label distillation supervision by a combination of teachers: ; in, This indicates the knowledge support head during the training period. This represents the intermediate logits output by the knowledge scaffold head during the training period.

[0060] The training-phase knowledge scaffold head can include a feature adaptation module, a feature aggregation module, and a class mapping module. These modules are used to convert the output features of intermediate neural network structural units into logits consistent with the number of task classes. The system consists of convolutional layers, batch normalization, nonlinear activation, global average pooling, and a linear classification layer. Its function is to map the features of intermediate neural network structural units to class logits, enabling the student's later-layer features to directly receive teacher soft-label supervision. This training-period knowledge scaffold does not alter the student's main network structure; it exists only as a temporary knowledge injection branch during the training phase. After training, the training-period knowledge scaffold is deleted, and only the student model's main backbone and main classification head are retained during final inference. This paper proposes a training-period knowledge scaffold mechanism for NAS student models. By introducing temporary knowledge injection branches at the output positions of specific neural network structural units, combined with teacher combinatorial distillation supervision and staged weight scheduling, it achieves intermediate representation enhancement and deletes the scaffold after training.

[0061] Step 3.3: Constructing the Distillation Training Objective: Integrate the real label supervision, the student model's main output layer distillation supervision, and the knowledge scaffold distillation supervision into a unified training objective, enabling the student model to simultaneously learn task label knowledge, the teacher model's final output knowledge, and the knowledge injected into the teacher model. The teacher combination with the highest overall score among the selected teacher combinations is then selected. The main output layer of the student model receives soft-label distillation supervision from the teacher combination, resulting in the main output distillation loss. Meanwhile, the intermediate logits output by the knowledge scaffold head during the training period That is, the knowledge scaffold head during the training period is paired with intermediate features. The class prediction logits obtained after mapping are subjected to soft-label distillation supervision by the same teacher combination, resulting in the knowledge scaffold distillation loss. In this invention, the knowledge scaffold head during training does not use the true label cross-entropy as the primary supervision, i.e. This setting is used to prevent intermediate features from prematurely undertaking the complete classification task, thereby interfering with the feature evolution of subsequent neural network structural units. This application, targeting the characteristics of the NAS student model, performs adaptability screening on the teacher combinations participating in soft-label supervision, ensuring that the main output layer receives not arbitrary teacher knowledge, but rather combined teacher knowledge determined by the teacher combination evaluation mechanism that is more suitable for the current student model. The total loss of distillation training is defined as: ; in, This represents the true label cross-entropy loss of the student model's main output. This indicates the distillation loss in the main output layer; This indicates the knowledge scaffold distillation loss; This indicates the weight corresponding to the distillation loss of the main output layer. This represents the weight of the knowledge scaffold distillation loss.

[0062] Step 3.4: Set up knowledge scaffold distillation loss weight scheduling: Implement phased control of the knowledge scaffold distillation loss weights, ensuring that the knowledge scaffold heads function as knowledge scaffolds in the middle of training and gradually withdraw in the later stages, thereby improving the representation learning effect of intermediate neural network structural units. Let the first step be... The knowledge scaffold distillation loss weights for each epoch. Three-stage scheduling is adopted: ; ; ; in, Indicates the current training round. This indicates the end of the warm-up training round. Indicates the cycle in which decay begins. This indicates the end of the distillation phase training round. This represents the basic weights of knowledge scaffold distillation.

[0063] Step 3.5: Complete distillation training and remove the knowledge scaffold head during training: optimize the total loss. The student model parameters were obtained after adaptive teacher combination distillation and unit-level knowledge scaffolding training. During this training process, the knowledge scaffolding head during the training period Soft-label distillation supervision is used to receive and adapt the combined output of teachers, enabling teacher model knowledge to be applied in advance to the representations of later neural network structural units in the student model. After distillation training is completed, the training-period knowledge scaffold header is deleted. Only the feature extraction backbone and final classifier of the NAS student model are retained.

[0064] Step 4: Perform high-performance teacher model calibration training on the distilled student model to obtain the final student model parameters. ;like Figure 5 As shown, step 4 specifically includes the following steps: Step 4.1: Determine the initial student model: After distillation training, select the student model with the best performance on the validation set from the training process and use it as the high-performance teacher model. The initialization of the student model is performed during the calibration training phase. During the distillation training phase, the student model acquires a series of parameters at different epochs. After each epoch, the student model is evaluated using a validation set to obtain the validation accuracy. The student model with the best performance on the validation set during the distillation stage is selected as the initial student model. The student model corresponding to this parameter has been trained by distillation in step 3 and has relatively stable intermediate representation and knowledge absorption capabilities. Subsequently, the initial student model was initialized. As a high-performance teacher model Initialization parameters during calibration training phase Through the above settings, a high-performance teacher model Instead of participating in distillation from the initial random initialization or early training stages of the student model, it intervenes only after the student model has completed adaptation to the teacher portfolio and intermediate representation training. This step can reduce the risk of negative transfer caused by excessively difficult knowledge, overly complex output distribution, or large gap between teachers and students in the high-performance teacher model, allowing the high-performance teacher model to mainly play a role in late-stage boundary calibration and fine-grained dark knowledge supplementation.

[0065] Step 4.2: Construct a high-performance teacher model and calibrate the training loss: Initialize the student model as determined in Step 4.1. Based on this, high-performance teacher model knowledge is introduced in a low-intensity manner, allowing the high-performance teacher model to serve as a calibration signal in the later stages, rather than re-dominant student training. The total loss of the high-performance teacher model calibration training phase is defined as: ; in, This represents the distillation loss of the high-performance teacher model; This represents the distillation weights of the high-performance teacher model. The student model's main output represents the true label cross-entropy loss, while the high-performance teacher model's distillation loss employs a high-temperature soft-label distillation method. , in, This indicates the temperature during the calibration training phase. This indicates the temperature of the high-performance teacher model during the calibration and training phase. The distribution of soft labels below, This indicates the temperature of the student model during the calibration and training phase. The distribution of soft labels below; The refinement stage meets the following requirements: , This indicates the temperature during the basic distillation stage. This indicates the weight corresponding to the distillation loss of the main output layer in step 3. The learning rate represents the student model's distillation training in step 3, used to control the magnitude of parameter updates during the overall distillation training process. By constraining the influence of the high-performance teacher model through high temperature, low weights, and low learning rate, the high-performance teacher model primarily provides fine-grained boundary calibration without destroying the student model representation already formed during the distillation stage.

[0066] Step 4.3: Complete the high-performance teacher model calibration training: Use the high-performance teacher model to calibrate the total loss during the training phase. The student model is trained over a short period of time to obtain the final student model parameters. After this step is completed, the student model, while maintaining the stability of the representation in the distillation stage of step 3, further absorbs the fine-grained dark knowledge provided by the high-performance teacher model.

[0067] Step 5: Determine the final student model: Based on the final student model parameters obtained in Step 4. The final student model is obtained as follows .

[0068] This application uses the CIFAR-100 image classification task as an example. In this example, the student model is the NAS student model obtained by DARTS search, the teacher pool is set to 4 teacher models, the number of teacher combinations is set to 2, and the high-performance teacher model is prepared separately and is not included in the teacher pool.

[0069] Step 1: Dataset Preparation. This example uses the CIFAR-100 image classification dataset, with image dimensions of [missing information]. The dataset contains 100 classes. The CIFAR-100 dataset includes 50,000 training images and 10,000 test images. The 50,000 training images are divided into a 45,000-image training set and a 5,000-image validation set. The training set is used for student model distillation training, the validation set is used for teacher-combined scoring and initializing student model selection, and the test set is used for final student model performance evaluation.

[0070] Step 2: Student Model Construction. The NAS student model obtained through DARTS search is used as the student model, denoted as... The student model is composed of multiple neural network structural units (cells) stacked sequentially. This NAS student model has a total of 6 neural network structural units.

[0071] Step 3: Teacher Pool Initialization. Set the teacher pool to... ,in , , , The aforementioned teacher models are all commonly used in image classification distillation tasks. Additionally, a high-performance teacher model, Pyramid272, is prepared, which achieves a classification accuracy of [missing information] on the CIFAR-100 dataset. .

[0072] Step 4: Teacher Group Setup. This embodiment sets each teacher group to include two teacher models, i.e. There are 4 teacher models in the teacher pool, therefore the number of teacher combinations is... The teacher combinations are as follows: .

[0073] Step 5: Calculate the single-teacher model fit score. For ease of explanation, this embodiment uses 4 samples from the validation set for numerical calculation. The predicted probability of the student model on the true class of the 4 samples is... The predicted probabilities of the four teacher models on the true class of the same sample are as follows: , , , Teacher model For example, teacher model The advantage of the student model in predicting the true class is: ; Other teachers calculated in the same way and obtained... The feature compatibility between teachers and students, calculated using CKA, is as follows: This yields the single-teacher fit score, using the teacher model. For example: ; Other teachers calculated in the same way and obtained... , , .

[0074] Step 6: Calculate teacher group suitability. Based on the individual teacher model fit score, further assess whether a teacher group as a whole is suitable for the current student model learning. In this embodiment, the combination weight calculation temperature is... , to combine For example: ; ; Therefore, the suitability of the teacher group is: .

[0075] Step 7: Calculate the tacit knowledge complementarity of teacher combinations. In this embodiment, the combination... Non-target dark knowledge residual cosine similarity is ,so and The complementarity of tacit knowledge is The weighted tacit knowledge complementarity is: ; In this embodiment, each candidate teacher combination contains two teacher models, therefore the combination... There is only one teacher in the room. ,therefore ,combination The overall complementarity of tacit knowledge is .

[0076] Step 8: Calculate the overall score of the teacher combination and select the optimal teacher combination. The final overall score of the teacher combination is obtained by combining the suitability of the teacher combination with the complementarity of implicit knowledge. In this embodiment, the enhancement coefficient of implicit knowledge complementarity is [value missing]. , to combine For example: .

[0077] Other teacher combinations were calculated directly, and the results are shown in Table 1 below.

[0078] Table 1 Calculation results for other teacher combinations

[0079] Due to the combination The overall score is the highest, therefore In this embodiment, ResNet32×4 and WRN-40-2 are selected as the optimal teacher combination for subsequent basic distillation training.

[0080] Step 9: Determine the insertion position of the training-period knowledge scaffold head and construct the training-period knowledge scaffold head. The student model has a total of... Set the relative insertion depth of the knowledge scaffold head during the training period to [value]. The insertion position of the knowledge scaffold head during the training period is... Therefore, in this embodiment, the training-phase knowledge scaffold head is connected to the fourth neural network structure unit. Features are then output in the middle of the fourth neural network structure. The training-phase knowledge scaffolding head is then connected, and this training-phase knowledge scaffolding head is composed of... It consists of convolution, batch normalization, nonlinear activation, global average pooling, and linear classification layers, and is used only during the training phase; it is removed during the inference phase.

[0081] Step 10: Construct distillation training objectives. In this embodiment, the adapted teacher combination is as follows: Specifically, ResNet32×4 and WRN-40-2. During training, the knowledge scaffold head does not use real-label cross-entropy supervision. Let the basic distillation temperature be... At the 20th epoch, the KL divergence between the student model's main output layer and the adapted teacher's combined soft labels is: Then the distillation loss of the main output layer is Let the KL divergence between the knowledge scaffold head output and the adapted teacher combination soft labels during the training period be . The knowledge distillation loss during the training period is... The student's main output true label cross-entropy loss is Let the distillation weight of the main output layer be... The training scaffold head distillation weights for the 20th epoch are: The total loss during the distillation stage is ; Therefore, the total loss of the CKS stage in the 20th epoch is .

[0082] Step 11: Configure scaffold distillation weight scheduling. This embodiment sets... The weight scheduling for knowledge scaffold head distillation during the training period is as follows: ; when hour: ; when hour: ; when hour: ; Therefore, the knowledge support head during the training period is turned off in the early stage of training, activated in the middle stage of training, and gradually weakened and withdrawn in the later stage of training.

[0083] Step 12: Complete distillation training and remove the knowledge scaffold head during training. This embodiment uses the training set to perform 50 epochs of distillation training on the NAS student model. The training objective is... After training, the student model parameters are obtained after optimal teacher combination distillation and unit-level knowledge scaffolding training. After the distillation training is complete, delete the knowledge scaffold head from the training period. Only the feature extraction backbone and final classifier of the NAS student model are retained.

[0084] Step 13: Determine the initial student model. During distillation training, the student model performance is evaluated using a validation set after each epoch. Let the validation set accuracies at epochs 45, 47, and 50 be respectively... , , ,but ; Subsequently, these student parameters were used as initialization parameters for the calibration training phase of the high-performance teacher model. .

[0085] Step 14: Construct a high-performance teacher model and calibrate the training loss. Let the temperature during the calibration training phase of the high-performance teacher model be . The KL divergence between the soft label distributions of the high-performance teacher model and the student model is: The distillation loss of the high-performance teacher model is The true label cross-entropy loss during the calibration training phase is The high-performance teacher model distillation weights are The total loss during the calibration and training phase of the high-performance teacher model is then... In this embodiment, the basic distillation stage temperature, principal KD weights, and learning rate are set as follows: The high-performance teacher model calibration training phase is set as follows: ,satisfy: ; ; ; Therefore, the high-performance teacher model is calibrated and trained using a high-temperature, low-distillation-weight, and low-learning-rate approach.

[0086] Step 15: Complete the high-performance teacher model calibration training. In this embodiment, the student model undergoes 15 epochs of high-performance teacher model calibration training. After the calibration training is completed, the final student model parameters are obtained. The final student model is obtained as follows .

[0087] This application targets NAS student models and comprehensively evaluates teacher combinations before distillation by considering the predictive advantages of teacher models for the true categories, teacher-student feature compatibility, and the complementarity of non-target dark knowledge residuals among teachers. This mechanism can screen out teacher combinations that are more suitable for the current student model's learning before formal distillation training, reducing the training cost of ineffective teacher combinations.

[0088] This application temporarily inserts a training-period knowledge scaffold header after the later neural network structural units in the NAS student model. This allows the soft-label distillation supervision from the teacher combination to directly affect the intermediate representations of the student model, and its activation, decay, and deactivation are controlled through knowledge scaffold distillation weight scheduling. This mechanism enhances the knowledge absorption capacity of the intermediate structural units in the NAS student model while ensuring that the training-period knowledge scaffold header is removed after training, without increasing the inference overhead of the final deployed model.

[0089] This application's high-performance teacher model possesses strong classification capabilities, providing finer-grained class relationships and soft-label distribution information. However, its knowledge complexity is high, and premature participation in distillation can easily affect training stability. This application introduces the high-performance teacher model for calibration training only after the knowledge scaffold distillation is complete and the student model has formed a relatively stable representation. This mechanism, while maintaining the stability of the initial distillation training, utilizes the high-performance teacher model to further improve the output distribution and classification boundary quality of the student model.

[0090] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.

Claims

1. A multi-teacher distillation method for neural network architecture search, characterized in that: The multi-teacher distillation method specifically includes the following steps: Step 1, Data Preparation, Basic Model and Teacher Pool Initialization: Determine the image classification dataset to be used for model training and evaluation, construct student models and teacher pools, and select teacher models from the teacher pools; Step 2: Select the optimal teacher combination: Based on the student model and teacher pool constructed in Step 1, calculate the single-teacher model fitness score, calculate the teacher group fitness score based on the single-teacher model fitness score, calculate the complementarity of the two teacher models in the teacher combination, calculate the teacher combination implicit knowledge complementarity by combining the complementarity and the single-teacher model fitness score, and calculate the teacher combination comprehensive score by combining the teacher group fitness score and the teacher combination implicit knowledge complementarity to select the optimal teacher combination; wherein, calculating the teacher combination implicit knowledge complementarity specifically includes the following steps: Step 2.3.1: Calculate temperature given dark knowledge. , No. Teacher Model Compared to the student model in the first Input image samples The hidden knowledge residuals provided above: ; in, Calculating temperature for student models using dark knowledge The soft probability distribution under the following conditions For the teacher model, temperature is calculated using hidden knowledge. Under the soft probability distribution, to avoid overlap between the real class supervision and the cross-entropy objective, the real class position is set to zero. , Only retain the hidden knowledge residuals on non-target categories; Step 2.3.2: For candidate teacher combinations The first in Teacher Model and the Teacher Model Teacher model Teacher Model After expanding the dark knowledge residuals on non-target categories into vectors, the teacher model is calculated. Teacher Model Complementarity between them: ; in, Teacher model The non-target dark knowledge residual vector Teacher model The non-target dark knowledge residual vector Indicates cosine similarity; Step 2.3.3: Based on the teacher model Teacher Model complementarity between Combined with single-teacher model adaptation scores To obtain the teacher model Teacher Model Weighted complementarity of students' tacit knowledge : ; in, Teacher model The fit score for the student model, Teacher model Fit scores for the student model; Step 2.3.4: Weighted complementarity based on dark knowledge Calculate the complementary nature of teacher combinations' implicit knowledge for: ; in, For candidate teacher combinations All teachers within the group constitute a set; Step 3: Based on the optimal teacher combination selected in Step 2, conduct student model distillation training based on unit-level knowledge scaffolding; Step 4: Perform high-performance teacher model calibration training on the distilled student model to obtain the final student model parameters. ; Step 5: Determine the final student model: Based on the final student model parameters obtained in Step 4. The final student model is obtained as follows .

2. The multi-teacher distillation method for neural network architecture search according to claim 1, characterized in that: Step 1 specifically includes the following steps: Step 1.1, Dataset Preparation: Given an image classification dataset and the image classification dataset It is divided into training set, validation set and test set, where, Indicates the first One input image sample, Indicates the first The true class label corresponding to each input image sample Indicates the total number of image samples; Step 1.2: Constructing the student model: Obtaining the student model The student model include The nth sequentially stacked neural network structural units, the nth Each neural network structural unit is denoted as . ,in The transfer process of intermediate features in the student model is represented as follows: , Indicates the first The intermediate features output by each neural network structural unit The intermediate features represent the output of the previous neural network structural unit, and the final main classification output logits of the student model are denoted as . ,in, This represents the trainable parameters of the student model. Indicates the input image sample; Step 1.3, Teacher Pool and High-Performance Teacher Model Initialization: Establishing a Pre-trained Teacher Pool , This represents the total number of teacher models in the teacher pool, for each input image sample. , No. Teacher Model The output logits is denoted as At the same time, prepare a high-performance teacher model. This high-performance teacher model The classification accuracy on the CIFAR-100 dataset meets the requirements. , Not part of the teacher pool ; Step 1.4, Teacher Combination Settings: Set up a pool of teachers. Select size Candidate teacher combination , This represents the teacher model selected from the teacher pool.

3. The multi-teacher distillation method for neural network architecture search according to claim 2, characterized in that: Step 2, calculating the single-teacher model fit score, specifically includes the following steps: Step 2.1.1: Student model matches the real category The predicted probability is The teacher model is for the real category The predicted probability is , No. Teacher Model For the real category Predictive advantage is defined as: ; in, Indicates the first Teacher Model The predictive advantage of the current student model over the true category, if the teacher model... Predicted probability on a given input image sample If the input image sample contribution is lower than that of the student model, the contribution is truncated to 0. Step 2.1.2: Let the feature matrix output of the intermediate layer corresponding to the student model be... The feature matrix output by the intermediate layer of the teacher model is: Using CKA to represent and Feature compatibility between: , in, This represents the feature matrix of the student model after centering. , Represents the feature matrix of the student model The mean vector calculated along the sample dimension This represents the feature matrix of the teacher model after centralization. , Represents the feature matrix of the teacher model The mean vector calculated along the sample dimension Student model feature matrix The transpose of the matrix, Teacher model feature matrix The transpose of the matrix; Step 2.1.3, place the first Teacher Model For the real category Predictive advantage and feature compatibility Combined, the single-teacher model fit score is obtained. .

4. The multi-teacher distillation method for neural network architecture search according to claim 3, characterized in that: Step 2 involves calculating the teacher group suitability based on the single-teacher model fit score, specifically including the following steps: Step 2.2.1: For candidate teacher combinations According to the candidate teacher combination Each single-teacher model adaptation score Calculate teacher weights : ; in, Indicates the first Teacher Model In the candidate teacher combination Weight within, Indicates teacher weight Calculated temperature parameters, Indicates candidate teacher combination The Middle Teacher Model The fit score; Step 2.2.2: Based on teacher weights Calculate the suitability of teacher groups: ; in, Indicates candidate teacher combination Adaptability to student models.

5. The multi-teacher distillation method for neural network architecture search according to claim 1, characterized in that: Step 2 involves calculating the comprehensive score of teacher combinations and selecting the optimal teacher combination by combining the suitability of teacher group pairings and the complementarity of teachers' implicit knowledge. This includes the following steps: Step 2.4.1: For candidate teacher combinations Teacher combination composite score Defined as ; in, Indicates candidate teacher combination Adaptability to the student model This indicates the complementary nature of the teachers' implicit knowledge. This represents the coefficient for enhancing the complementarity of dark knowledge. Step 2.4.2: Select the teacher combination with the highest overall score. 。 6. The multi-teacher distillation method for neural network architecture search according to claim 5, characterized in that: Step 3 specifically includes the following steps: Step 3.1: Temporarily connect the training period knowledge scaffold header after the later neural network structure units in the student model, so that the teacher's knowledge enters the intermediate features of the student model. Let the student model have a total of The insertion position of the knowledge scaffold head during the training phase is defined as follows: (The original text contains several typographical errors and inconsistencies. A more accurate translation would require the full context.) ; in, This indicates the sequence number of the neural network structure unit inserted into the knowledge scaffold during the training period. Indicates the relative depth position of the knowledge scaffold head during the training period; Step 3.2, Constructing the Knowledge Scaffold Header for the Training Period: In the... Intermediate features output by each neural network structural unit A lightweight training branch is then added to process the intermediate features output by the neural network architecture units. Mapped to category logits, and subject to soft-label distillation supervision by a combination of teachers: ; in, This indicates the knowledge support head during the training period. This represents the intermediate logits output by the knowledge scaffold head during the training period; Step 3.3: Constructing the Distillation Training Objective: For the selected teacher combination, choose the teacher combination with the highest overall score. The main output layer of the student model receives soft-label distillation supervision from the teacher combination, resulting in the main output distillation loss. Meanwhile, the intermediate logits output by the knowledge scaffold head during the training period That is, the knowledge scaffold head during the training period is paired with intermediate features. The class prediction logits obtained after mapping are subjected to soft-label distillation supervision by the same teacher combination, resulting in the knowledge scaffold distillation loss. The total loss during distillation training is defined as: ; in, This represents the true label cross-entropy loss of the student model's main output. This indicates the distillation loss in the main output layer; This indicates the knowledge scaffold distillation loss; This indicates the weight corresponding to the distillation loss of the main output layer. Indicates the weight of knowledge scaffold distillation loss; Step 3.4: Set up knowledge scaffold distillation loss weight scheduling: Let the first... The knowledge scaffold distillation loss weights for each epoch. Three-stage scheduling is adopted: ; ; ; in, Indicates the current training round. This indicates the end of the warm-up training round. Indicates the cycle in which decay begins. This indicates the end of the distillation phase training round. This represents the basic weights of knowledge scaffold distillation; Step 3.5: Complete distillation training and remove the knowledge scaffold head during training: optimize the total loss. The student model parameters were obtained after adaptive teacher combination distillation and unit-level knowledge scaffolding training. During this training process, the knowledge scaffolding head during the training period Soft-label distillation supervision is used to receive and adapt the combined output of teachers, enabling teacher model knowledge to be applied in advance to the representations of later neural network structural units in the student model. After distillation training is completed, the training-period knowledge scaffold header is deleted. Only the feature extraction backbone and final classifier of the NAS student model are retained.

7. The multi-teacher distillation method for neural network architecture search according to claim 6, characterized in that: Step 4 specifically includes the following steps: Step 4.1: Determine the initial student model: During the distillation training phase, the student model will acquire a series of parameters at different epochs. After each epoch, the student model is evaluated using the validation set to obtain the validation set accuracy. The student model with the best performance on the validation set during the distillation stage is selected as the initial student model. This initializes the student model. As a high-performance teacher model Initialization parameters during calibration training phase ; Step 4.2: Construct a high-performance teacher model and calibrate the training loss: Initialize the student model as determined in Step 4.

1. Based on this, the total loss during the calibration training phase of the high-performance teacher model is defined as: ; in, This represents the distillation loss of the high-performance teacher model; This represents the distillation weights of the high-performance teacher model. The student model's main output represents the true label cross-entropy loss, while the high-performance teacher model's distillation loss employs a high-temperature soft-label distillation method. , in, This indicates the temperature during the calibration training phase. This indicates the temperature of the high-performance teacher model during the calibration and training phase. The distribution of soft labels below, This indicates the temperature of the student model during the calibration and training phase. The distribution of soft labels below; The refinement stage meets the following requirements: , This indicates the temperature during the basic distillation stage. This indicates the weight corresponding to the distillation loss of the main output layer in step 3. This represents the learning rate of the student model distillation training in step 3, which is used to control the magnitude of parameter updates during the overall distillation training process. Step 4.3: Complete the high-performance teacher model calibration training: Use the high-performance teacher model to calibrate the total loss during the training phase. The student model is trained over a short period of time to obtain the final student model parameters. .

Citation Information

Patent Citations

  • Implicit recommendation method and system based on network self-collaboration

    CN112464104A

  • Training method, image feature extraction method, and image recognition method and device

    CN113326768A