Teacher-student model optimal matching method based on dynamic threshold value

By dynamically adjusting the pruning threshold, the problem that fixed pruning rate in traditional iterative pruning is difficult to adapt to different layers and pruning processes of the model in traditional iterative pruning is solved, and the adaptive removal of deep model redundancy and automatic matching of student models is achieved, which significantly improves the model compression efficiency and performance.

CN119940453APending Publication Date: 2025-05-06GUANGDONG POLYTECHNIC NORMAL UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510127150.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In traditional iterative pruning methods, fixed pruning rate is difficult to adapt to the sensitivity of different layers of the model and the different stages of the pruning process, resulting in the problem of excessive pruning or underpruning.

Method used

Using a teacher-student model best matching method based on dynamic thresholds, the pruning threshold is dynamically adjusted, real-time monitoring and adaptive tuning of depth model redundancy removal process is used, through iterative cycles of pruning and performance evaluation.

Benefits of technology

The excessive pruning and underpruning problems caused by traditional fixed pruning rate methods are avoided, and multiple student models that adapt to different performance-volume needs are automatically produced, which significantly reduces the cost of manual parameter adjustment and enables the teacher-student model to be better matched in knowledge distillation or deployment applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940453A_ABST
    Figure CN119940453A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, and discloses a teacher-student model optimal matching method based on a dynamic threshold value, which comprises the following steps: carrying out iterative pruning on a pre-trained teacher model to obtain an intermediate pruning model; after each iteration pruning, evaluating the performance of the middle pruning model; dynamically adjusting the pruning threshold value of the next iteration pruning based on the performance feedback of the middle pruning model; repeating the first three steps until a preset pruning termination condition is met, and obtaining at least one student model; and selecting a student model matched with the teacher model from the at least one student model. Real-time monitoring and self-adaptive adjustment and optimization of a depth model redundancy removal process are realized through a dynamic evaluation and threshold adjustment mechanism; the problems of excessive pruning and insufficient pruning caused by a traditional fixed pruning rate method are avoided; and a plurality of student models adapting to different performances are produced, so that the manual parameter adjustment cost is remarkably reduced, and the teacher-student models are better matched in knowledge distillation or deployment application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to a teacher-student model optimal matching method based on dynamic threshold. Background Art

[0002] In recent years, deep learning technology has achieved great success in image recognition, natural language processing and other fields. However, high performance is often accompanied by a large model size and computational complexity, which poses a challenge to the deployment of models on resource-constrained devices. To solve this problem, model compression technology has emerged, among which pruning is an important model compression method. Pruning removes unimportant connections or neurons in the model, reduces the number of model parameters and computation, and thus achieves model lightweighting.

[0003] Among the numerous pruning methods, iterative pruning has attracted widespread attention because it can gradually and finely prune the model. Traditional iterative pruning methods usually use a fixed pruning rate, that is, a fixed proportion of weights or channels are removed in each round of pruning. Although this method is simple and easy to implement, it has obvious shortcomings.

[0004] First, a fixed pruning rate is difficult to adapt to the sensitivity of different layers of the model. Different layers of the model have different sensitivities to pruning. Shallow networks usually extract more general features and are more sensitive to pruning, while deep networks extract more specific features and are more tolerant to pruning. Using a fixed pruning rate may lead to over-pruning of sensitive layers, seriously damaging model performance, while under-pruning of insensitive layers leads to poor compression effects. The fundamental reason is that the redundancy and importance of features learned by different layers of the model are different, and a fixed pruning rate cannot distinguish them.

[0005] Secondly, a fixed pruning rate is difficult to adapt to different stages of the pruning process. In the early stages of pruning, there may be more redundancy in the model, which can withstand a higher pruning rate. As pruning progresses, the model gradually approaches its performance limit. At this time, if the same pruning rate is still used, it is easy to cause over-pruning, resulting in a sharp drop in performance. Conversely, if a lower pruning rate is set at the beginning, the compression efficiency will be limited. This is because as pruning progresses, the redundant information in the model gradually decreases, and continuing to use the same pruning rate will touch the bottom line of model performance.

[0006] In order to solve the problems caused by fixed pruning rates, existing technologies have also made some explorations. For example, some methods try to set different fixed pruning rates for different layers, but this still requires manual experience or a large number of experiments to determine the appropriate pruning rate, which increases the design cost and uncertainty. Other methods try to introduce some heuristic rules to adjust the pruning rate during the pruning process, but these rules are often based on experience, lack theoretical support, and are difficult to guarantee universality in various models and tasks.

[0007] Therefore, how to dynamically adjust the pruning rate according to the actual situation of the model to avoid over-pruning or under-pruning has become a key issue that needs to be urgently solved in the field of iterative pruning. Summary of the invention

[0008] The present invention provides a teacher-student model optimal matching method based on dynamic thresholds. Through dynamic evaluation and threshold adjustment mechanisms, real-time monitoring and adaptive tuning of the deep model redundancy removal process are achieved; over-pruning and under-pruning problems caused by traditional fixed pruning rate methods are avoided; multiple student models that meet different performance-volume requirements are automatically generated, which significantly reduces the cost of manual parameter adjustment and enables better matching of teacher-student models in knowledge distillation or deployment applications.

[0009] The present invention provides a teacher-student model optimal matching method based on a dynamic threshold, comprising the following steps:

[0010] S1: Iteratively prune the pre-trained teacher model to obtain an intermediate pruned model;

[0011] S2: After each iterative pruning, the performance of the intermediate pruned model is evaluated;

[0012] S3: Dynamically adjust the pruning threshold for the next iterative pruning based on the performance feedback of the intermediate pruning model;

[0013] S4: Repeat S1 to S3 until the preset pruning termination condition is met and at least one student model is obtained;

[0014] S5: From at least one student model, select a student model that matches the teacher model.

[0015] Preferably, in step S1, iterative pruning includes performing structured pruning on the convolution channels of the teacher model, and removing unimportant channels by evaluating the importance of each channel to obtain an intermediate pruned model.

[0016] Preferably, the importance is evaluated based on the L1 norm, and the L1 norm of each channel in the teacher model is globally calculated and pruned in order from small to large.

[0017] Preferably, in step S2, after completing one pruning, a performance measurement or performance test is performed on the intermediate pruned model by running the intermediate pruned model on a validation set or a test set, and recording its performance on the main task indicators. The main task indicators specifically include classification accuracy, recall rate and loss function value.

[0018] Preferably, in step S3, the dynamic adjustment logic of the pruning threshold is:

[0019] If the model performance drops beyond the preset threshold, the pruning threshold is lowered to reduce the amount of pruning;

[0020] If the model performance decreases slightly, increase the pruning threshold and the amount of pruning.

[0021] Preferably, the reducing or increasing the pruning amount adopts a minimum pruning method;

[0022] The minimum pruning method includes: pruning one channel of a layer each time, and judging whether it exceeds the performance degradation threshold after pruning; if not, continue pruning, otherwise end pruning.

[0023] Preferably, each time a channel of a certain layer is pruned, after pruning, it is determined whether the performance degradation threshold is exceeded; if not, pruning is continued, otherwise pruning is terminated specifically as follows:

[0024] Given a neural network with L convolutional layers, A = (C1, C2, ...C L ) is the original network, where C1 is the number of channels in the first layer; before iterative pruning, determine the target pruning rate R and the acceptable performance loss threshold accLossHold of each layer in each round of pruning, and the current model performance Aacc;

[0025] (1) Prune the current layer, and only delete one filter with the lowest L1-norm each time to achieve minimum pruning;

[0026] (2) Test the model performance Bacc and obtain the loss performance accLoss, specifically: accLoss = Bacc - Aacc;

[0027] (3) If accLoss exceeds the loss threshold accLossHold, perform the next layer pruning and update: Aacc = Bacc, otherwise repeat (1) and (2);

[0028] (4) When each layer is pruned, determine whether the target pruning rate is met. If not, repeat (1), (2), and (3).

[0029] (5) When the pruning rate meets the target, the pruning is completed.

[0030] Preferably, in step S4, the preset pruning termination condition includes:

[0031] The overall pruning rate of the model reaches the preset value;

[0032] The main task indicators dropped to the threshold accepted by users;

[0033] The given number of iterations has been reached.

[0034] Preferably, the overall pruning rate includes a channel pruning rate or a parameter pruning rate.

[0035] Preferably, in step S5, the criteria for the matching student model include: a balance between the highest precision and the lowest number of parameters, the highest accuracy under certain resource constraints, and the minimum model size under a fixed accuracy requirement.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] The present invention discloses a teacher-student model optimal matching method based on dynamic thresholds. Through dynamic evaluation and threshold adjustment mechanisms, real-time monitoring and adaptive tuning of the deep model redundancy removal process are achieved; over-pruning and under-pruning problems caused by traditional fixed pruning rate methods are avoided; multiple student models that meet different performance-volume requirements are automatically generated, which significantly reduces the cost of manual parameter adjustment and enables better matching of teacher-student models in knowledge distillation or deployment applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a flowchart of a method for optimal matching of teacher-student models based on dynamic thresholds provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] like Figure 1 As shown, the present application provides a teacher-student model optimal matching method based on a dynamic threshold, comprising the following steps:

[0041] S1: Iteratively prune the pre-trained teacher model to obtain an intermediate pruned model;

[0042] S2: After each iterative pruning, the performance of the intermediate pruned model is evaluated;

[0043] S3: Dynamically adjust the pruning threshold for the next iterative pruning based on the performance feedback of the intermediate pruning model;

[0044] S4: Repeat S1 to S3 until the preset pruning termination condition is met and at least one student model is obtained;

[0045] S5: From at least one student model, select a student model that matches the teacher model.

[0046] In the above scheme, S1 iteratively prunes the teacher model, which is a cyclic process used to gradually remove redundant connections or channels in the model to achieve the purpose of model compression; after each iterative pruning, the model will generate an intermediate pruned model, which is a model instance in the current pruned state; S2 introduces a performance evaluation link. After each iterative pruning, the current intermediate pruned model will be performance tested, usually by running the model on a validation set or a test set, and recording its performance on key task indicators, such as classification accuracy, recall rate or loss function value, in order to quantify the impact of the pruning operation on model performance; S3 adaptively adjusts the pruning threshold for the next iterative pruning based on the performance feedback information obtained in S2; the pruning threshold determines the number or importance of channels or connections to be removed in the next iterative pruning; the logic of dynamic adjustment is that if the performance of the model deteriorates too quickly after the current pruning operation, the pruning threshold is lowered to slow down the pruning speed; In other words, if the performance degradation is small, the pruning threshold can be increased to speed up the pruning speed; this dynamic adjustment mechanism enables the pruning process to adapt more finely to the sensitivity of the model to pruning, avoiding the problems of over-pruning or under-pruning that may occur in the traditional fixed pruning rate method; S4 discloses the cyclic process of iterative pruning; S1 to S3 will be repeatedly executed until the preset pruning termination condition is met; the pruning termination condition can be a preset target pruning rate, the model performance drops to an acceptable threshold boundary, or reaches a preset upper limit on the number of iterations; through cyclic iterations, the method can generate a series of student models with different pruning degrees and performance performance; S5 selects a student model that matches the teacher model from at least one student model generated by S4. The matching standard can be determined according to the actual application requirements, such as seeking a balance between accuracy and model size, or pursuing the highest accuracy under resource constraints, or pursuing the smallest model size under certain accuracy requirements.

[0047] The invention of this application aims to provide a method that can adaptively optimize the iterative pruning process, overcome the limitations of fixed pruning rates in traditional iterative pruning methods, improve the efficiency and performance of model compression, and ultimately achieve automatic generation of a student model that best matches the teacher model, thereby reducing the cost of manually designing the student model and improving the knowledge distillation effect; the dynamic threshold adjustment mechanism enables the pruning process to be adaptively adjusted according to model performance feedback, avoiding the problem of over-pruning or under-pruning that may be caused by a fixed pruning rate; through the cycle of iterative pruning and performance evaluation, the pruning process can be finely controlled to achieve a higher model compression rate while ensuring model performance; a series of student models are automatically generated, and the best matching student model is selected from them, reducing the cost and uncertainty of manually designing the student model and improving the efficiency and performance of knowledge distillation; and finally obtaining a student model that best matches the teacher model, which can achieve a better balance between model size, computational complexity and performance, and is more suitable for deployment on resource-constrained devices.

[0048] Preferably, in step S1, iterative pruning includes performing structured pruning on the convolution channels of the teacher model, and removing unimportant channels by evaluating the importance of each channel to obtain an intermediate pruned model.

[0049] In the above scheme, the specific method of iterative pruning is to perform structured pruning on the convolution channels of the teacher model. The difference between structured pruning and unstructured pruning is that structured pruning removes the entire structural unit in the model, such as convolution channels or neurons, rather than a single weight connection; the convolution channel is the basic structural unit in the convolutional neural network. Pruning the convolution channel can directly reduce the number of model parameters and the amount of calculation, and maintain the regularity of the model structure, which is more conducive to hardware acceleration and deployment; structured pruning is achieved by evaluating the importance of each channel and removing unimportant channels. Channel importance evaluation is a key step in structured pruning, which determines which channels should be removed. By evaluating the importance of the channel, the channels that have less impact on the model performance can be removed first, thereby achieving model compression while maintaining the model performance as much as possible. After removing unimportant channels, the intermediate pruned model can be obtained.

[0050] Structured pruning can effectively reduce the number of model parameters and computational complexity, while maintaining the regularity of the model structure, which is more conducive to model deployment and hardware acceleration. Pruning based on channel importance evaluation can preferentially remove channels that have less impact on model performance, thereby achieving more effective model compression while ensuring model performance. Compared with unstructured pruning, structured pruning is usually easier to implement and deploy, and is more friendly to hardware accelerators.

[0051] Preferably, the importance is evaluated based on the L1 norm, and the L1 norm of each channel in the teacher model is globally calculated and pruned in order from small to large.

[0052] In the above scheme, the evaluation of channel importance is based on the L1 norm, which is a common indicator for measuring vector sparsity. In model pruning, the L1 norm of a channel is usually used as an indicator for measuring the importance of the channel; the smaller the L1 norm, the less important the channel is generally considered to be; global calculation refers to calculating the L1 norm of the channel in all convolutional layers of the entire teacher model, rather than just calculating it in a single layer; pruning in order from small to large means that in each iterative pruning, the channel with the smallest L1 norm is preferentially removed. This pruning strategy aims to remove the least important channels globally, thereby compressing the model more effectively.

[0053] Evaluating channel importance based on L1 norm is computationally simple, efficient and easy to implement. Calculating channel L1 norm globally can select the least important channels in the global scope for pruning from the perspective of the entire model, thereby improving the effectiveness of pruning. Pruning in ascending order of L1 norm can preferentially remove the least important channels, thereby achieving a higher compression rate while ensuring model performance.

[0054] Preferably, in step S2, after completing one pruning, a performance measurement or performance test is performed on the intermediate pruned model by running the intermediate pruned model on a validation set or a test set, and recording its performance on the main task indicators. The main task indicators specifically include classification accuracy, recall rate and loss function value.

[0055] In the above scheme, the performance evaluation is performed after completing one pruning operation, that is, after each iterative pruning operation, the performance of the current intermediate pruning model is immediately evaluated to obtain timely performance feedback; the specific method of performance evaluation is to run the intermediate pruning model on the validation set or test set and record its performance on the main task indicators; the validation set and test set are data sets used to evaluate the generalization performance of the model. By running the model on these data sets, the performance of the model in practical applications can be truly reflected. The main task indicator refers to the performance evaluation indicator selected according to the specific task type. For example, in image classification tasks, it can be classification accuracy; in target detection tasks, it can be recall rate; in various tasks, the loss function value can also be used as one of the performance indicators. The purpose of recording these indicators is to quantify the impact of pruning operations on model performance and provide a basis for subsequent dynamic threshold adjustment.

[0056] By performing performance evaluation immediately after completing a pruning operation, the impact of the pruning operation on the model performance can be timely obtained, providing timely feedback information for dynamic threshold adjustment. By running the model on the validation set or test set, the performance of the model in actual applications can be truly reflected, ensuring the accuracy of performance evaluation. By recording the main task indicators, the changes in model performance can be quantified, providing a clear basis for dynamic threshold adjustment.

[0057] Preferably, in step S3, the dynamic adjustment logic of the pruning threshold is:

[0058] If the model performance drops beyond the preset threshold, the pruning threshold is lowered to reduce the amount of pruning;

[0059] If the model performance decreases slightly, increase the pruning threshold and the amount of pruning.

[0060] In the above scheme, the adjustment of pruning threshold is based on the degradation of model performance; the specific adjustment logic is divided into two cases:

[0061] If the model performance drops by more than the preset threshold: The preset threshold refers to the maximum acceptable drop in model performance. If the model performance drop evaluated in S2 exceeds this threshold, it indicates that the current pruning operation may be too aggressive and has caused significant damage to the model performance. At this time, it is necessary to lower the pruning threshold to reduce the amount of pruning in the next iteration, thereby slowing down the pruning speed and protecting the model performance.

[0062] If the model performance degradation is small: If the model performance degradation is less than or equal to the preset threshold, it indicates that the current pruning operation has little impact on the model performance and the model may still have redundancy; at this time, the pruning threshold can be increased to increase the pruning amount of the next iterative pruning, thereby speeding up the pruning speed and improving the model compression efficiency; this dynamic adjustment logic forms a negative feedback mechanism that can adaptively adjust the pruning intensity according to the model's sensitivity to pruning to avoid over-pruning or under-pruning.

[0063] Dynamically adjusting the pruning threshold based on the degradation of model performance can make the pruning process more intelligent and avoid the limitations brought by fixed thresholds; lowering the pruning threshold when the performance degradation exceeds the threshold can prevent excessive pruning and protect model performance; raising the pruning threshold when the performance degradation is small can speed up pruning and improve model compression efficiency; the negative feedback mechanism enables the pruning process to adaptively approach the model performance limit, achieving a higher compression rate while ensuring model performance.

[0064] Preferably, the reducing or increasing the pruning amount adopts a minimum pruning method;

[0065] The minimum pruning method includes: pruning one channel of a layer each time, and judging whether it exceeds the performance degradation threshold after pruning; if not, continue pruning, otherwise end pruning.

[0066] In the above scheme, the minimum pruning method is a refined pruning strategy. Its core idea is to prune only one channel of a certain layer each time, and immediately evaluate the performance feedback, and decide whether to continue pruning based on the performance feedback; in each iterative pruning, only the channel with the lowest importance in the current layer is selected for removal. This minimum pruning operation can minimize the potential impact of each pruning on the model performance; after pruning a channel each time, a performance evaluation is immediately performed to determine whether the model performance degradation exceeds the preset performance degradation threshold; if the performance degradation does not exceed the threshold, it means that the current model can still withstand further pruning and the next channel can be pruned; if the performance degradation exceeds the threshold, it means that the current model is close to the performance limit, and continued pruning may cause a sharp drop in performance. At this time, the pruning of the current layer should be terminated, and the next layer should be pruned, or the current round of iterative pruning should be terminated; the minimum pruning method combined with the dynamic threshold adjustment mechanism can achieve more refined pruning control, and maximize the model's pruning space while ensuring model performance.

[0067] The minimum pruning method is adopted to achieve more refined pruning control and improve the accuracy and efficiency of pruning. The minimum pruning strategy of pruning one channel at a time can minimize the potential impact of each pruning on the model performance and achieve a smoother performance degradation curve. The performance degradation threshold is immediately determined after pruning, and performance feedback can be obtained in time. The pruning behavior is dynamically adjusted according to the feedback to achieve more sensitive threshold adjustment. If the performance does not exceed the threshold, pruning continues, otherwise pruning is terminated. It can explore the prunable space of the model more finely and maximize the compression rate while ensuring model performance. The minimum pruning method combined with the dynamic threshold adjustment mechanism can achieve a smarter and more efficient iterative pruning process.

[0068] Preferably, each time a channel of a certain layer is pruned, after pruning, it is determined whether the performance degradation threshold is exceeded; if not, pruning is continued, otherwise pruning is terminated specifically as follows:

[0069] Given a neural network with L convolutional layers, A = (C1, C2, ...C L ) is the original network, where C1 is the number of channels in the first layer; before iterative pruning, determine the target pruning rate R and the acceptable performance loss threshold accLossHold of each layer in each round of pruning, and the current model performance Aacc;

[0070] (1) Prune the current layer, and only delete one filter with the lowest L1-norm each time to achieve minimum pruning;

[0071] (2) Test the model performance Bacc and obtain the loss performance accLoss, specifically: accLoss = Bacc - Aacc;

[0072] (3) If accLoss exceeds the loss threshold accLossHold, perform the next layer pruning and update: Aacc = Bacc, otherwise repeat (1) and (2);

[0073] (4) When each layer is pruned, determine whether the target pruning rate is met. If not, repeat (1), (2), and (3).

[0074] (5) When the pruning rate meets the target, the pruning is completed.

[0075] In the above scheme, in the current convolutional layer, the L1 norm of all filters is calculated, and the filter with the smallest L1 norm is selected for pruning to achieve minimum pruning; after pruning, the performance Bacc of the model is tested, and the performance loss accLoss is calculated, that is, the difference between the performance Bacc after pruning and the performance Aacc before pruning; it is determined whether the performance loss accLoss exceeds the preset performance loss threshold accLossHold. If it exceeds, the pruning of the current layer is stopped, and the next layer is pruned, and Aacc is updated to Bacc as the benchmark performance of the next layer pruning; otherwise, if the performance loss does not exceed the threshold, steps (1) and (2) are continued in the current layer, that is, the next channel is pruned; when all convolutional layers have completed a round of pruning, it is determined whether the overall pruning rate of the current model reaches the target pruning rate R; if not, steps (1), (2) and (3) are repeated to perform the next round of iterative pruning; when the overall pruning rate of the model reaches the target pruning rate R, the iterative pruning process ends.

[0076] Preferably, in step S4, the preset pruning termination condition includes:

[0077] The overall pruning rate of the model reaches the preset value;

[0078] The main task indicators dropped to the threshold accepted by users;

[0079] The given number of iterations has been reached.

[0080] In the above scheme, when the overall pruning rate of the model (such as channel pruning rate or parameter pruning rate) reaches a preset target value, the iterative pruning is stopped; this condition is applicable to application scenarios that need to compress the model to a specific size; when the performance of the model on the main task indicators (such as classification accuracy) drops to the minimum level acceptable to the user, the iterative pruning is stopped; this condition is applicable to application scenarios that have minimum requirements for model performance; when the number of iterative pruning rounds reaches a preset upper limit, the iterative pruning is stopped; this condition can be used as a safety mechanism to prevent the iterative pruning process from looping infinitely, or to limit the pruning time when resources are limited; these pruning termination conditions can be used alone or in combination, and users can choose the appropriate termination condition according to actual application needs.

[0081] A variety of optional pruning termination conditions make the method of the present invention more flexible and adaptable to different application scenarios and user needs; the termination condition based on the pruning rate can control the compression degree of the model and is suitable for scenarios with strict requirements on the model size; the termination condition based on the performance indicator can ensure the performance bottom line of the model and is suitable for scenarios with minimum requirements on model performance; the termination condition based on the number of iterations can limit the pruning time and resource consumption and is suitable for scenarios with limited resources.

[0082] Preferably, the overall pruning rate includes a channel pruning rate or a parameter pruning rate.

[0083] In the above scheme, the channel pruning ratio refers to the ratio of the number of removed channels to the total number of original channels. The channel pruning ratio more directly reflects the compression degree of structured pruning; the parameter pruning ratio refers to the ratio of the number of removed parameters to the total number of original parameters. The parameter pruning ratio more comprehensively reflects the degree of reduction in the number of model parameters, including the parameter reduction brought about by channel pruning and other types of pruning (such as weight pruning).

[0084] Preferably, in step S5, the criteria for the matching student model include: a balance between the highest precision and the lowest number of parameters, the highest accuracy under certain resource constraints, and the minimum model size under a fixed accuracy requirement.

[0085] In the above scheme, the balance between the highest accuracy and the lowest number of parameters refers to seeking the best balance between accuracy and model size, and selecting the student model with the smallest number of model parameters under the premise of ensuring high accuracy. This standard is suitable for application scenarios that need to balance accuracy and model size; the highest accuracy under certain resource constraints refers to selecting the student model that can achieve the highest accuracy under given resource constraints (such as computing resources and memory resources). This standard is suitable for application scenarios where resource-constrained devices are deployed and it is necessary to pursue the best performance under limited resources; the minimum model size under fixed accuracy requirements refers to selecting the student model with the smallest model size under the premise of meeting the preset accuracy requirements. This standard is suitable for application scenarios that have minimum requirements for model accuracy and hope to reduce the model size as much as possible. These matching criteria can be selected according to actual application needs, or used in combination to find the student model that best meets the application needs.

[0086] A variety of optional matching criteria make the method of the present invention more flexible and adaptable to different application scenarios and optimization goals; the standard of balancing accuracy and parameter quantity can find the best compromise between accuracy and model size; the highest accuracy standard under resource constraints can maximize model performance under resource-constrained conditions; the minimum model size standard under fixed accuracy requirements can minimize the model size while meeting accuracy requirements; users can choose appropriate matching criteria according to actual application requirements, find the student model that best meets application requirements, and improve the flexibility and efficiency of model deployment.

[0087] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for optimal matching of teacher-student models based on dynamic threshold, characterized in that: The following steps are involved: S1: Iteratively prune the pre-trained teacher model to obtain an intermediate pruned model; S2: After each iterative pruning, the performance of the intermediate pruned model is evaluated; S3: Dynamically adjust the pruning threshold for the next iterative pruning based on the performance feedback of the intermediate pruning model; S4: Repeat S1 to S3 until the preset pruning termination condition is met and at least one student model is obtained; S5: From at least one student model, select a student model that matches the teacher model.

2. According to claim 1, a teacher-student model optimal matching method based on dynamic threshold is characterized in that: In step S1, iterative pruning includes performing structured pruning on the convolution channels of the teacher model, and removing unimportant channels by evaluating the importance of each channel to obtain an intermediate pruned model.

3. The optimal matching method of teacher-student model based on dynamic threshold according to claim 2, characterized in that: The importance is evaluated based on the L1 norm, and the L1 norm of each channel in the teacher model is globally calculated and pruned in order from small to large.

4. The optimal matching method of teacher-student model based on dynamic threshold according to claim 3, characterized in that: In step S2, after completing one pruning, the intermediate pruned model is performance measured or tested by running the intermediate pruned model on a validation set or a test set, and recording its performance on the main task indicators. The main task indicators specifically include classification accuracy, recall rate, and loss function value.

5. The method for optimal matching of teacher-student models based on dynamic threshold according to claim 4, characterized in that: In step S3, the dynamic adjustment logic of the pruning threshold is: If the model performance drops beyond the preset threshold, the pruning threshold is lowered to reduce the amount of pruning; If the model performance decreases slightly, increase the pruning threshold and the amount of pruning.

6. The optimal matching method of teacher-student model based on dynamic threshold according to claim 5, characterized in that: The reduction of pruning amount or increase of pruning amount adopts the minimum pruning method; The minimum pruning method includes: pruning one channel of a layer each time, and judging whether it exceeds the performance degradation threshold after pruning; if not, continue pruning, otherwise end pruning.

7. The method for optimal matching of teacher-student models based on dynamic threshold according to claim 6, characterized in that: Each time a channel of a certain layer is pruned, it is determined whether the performance degradation threshold is exceeded after pruning; if not, pruning continues, otherwise pruning ends specifically as follows: Given a neural network with L convolutional layers, A = (C1, C2, ...C L ) is the original network, where C1 is the number of channels in the first layer; before iterative pruning, determine the target pruning rate R and the acceptable performance loss threshold accLossHold of each layer in each round of pruning, and the current model performance Aacc; (1) Prune the current layer, and only delete one filter with the lowest L1-norm each time to achieve minimum pruning; (2) Test the model performance Bacc and obtain the loss performance accLoss, specifically: accLoss = Bacc - Aacc; (3) If accLoss exceeds the loss threshold accLossHold, perform the next layer pruning and update: Aacc = Bacc, otherwise repeat (1) and (2); (4) When each layer is pruned, determine whether the target pruning rate is met. If not, repeat (1), (2), and (3). (5) When the pruning rate meets the target, the pruning is completed.

8. The method for optimal matching of teacher-student models based on dynamic threshold according to claim 7, characterized in that: In step S4, the preset pruning termination condition includes: The overall pruning rate of the model reaches the preset value; The main task indicators dropped to the threshold accepted by users; The given number of iterations has been reached.

9. The method for optimal matching of teacher-student models based on dynamic threshold according to claim 8, characterized in that: The overall pruning rate includes a channel pruning rate or a parameter pruning rate.

10. The method for optimal matching of teacher-student models based on dynamic threshold according to claim 9, characterized in that: In step S5, the criteria for the matching student model include: a balance between the highest precision and the lowest number of parameters, the highest accuracy under certain resource constraints, and the minimum model size under a fixed accuracy requirement.

Citation Information

Patent Citations

  • Deep learning model acquisition method and device, electronic equipment and medium

    CN117236375A

  • Reverse channel pruning compression method and device based on multilevel knowledge distillation

    CN117744737A

  • Neural network model cascade compression method

    CN118312772A

  • Model compression method and system, electronic equipment and storage medium

    CN118798294A

  • Model optimization system and method based on deep learning

    CN119337949A