Efficient large model training optimization method based on dynamic resource allocation and knowledge distillation

Through dynamic resource allocation and knowledge distillation methods, the problems of differences in the structure of teacher models and student models and inaccurate knowledge expression are solved, and the stability of knowledge learning and training efficiency are improved.

CN119849594BActive Publication Date: 2025-06-06山东亚微软件股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510336242.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-06
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

In the prior art, there are large differences in the architecture between the teacher model and the student model, and the expression of knowledge is not accurate enough during the distillation process, resulting in unstable knowledge learned.

Method used

By building a real-time resource game controller, the Nash equilibrium algorithm is used to dynamically allocate computing resources, combine cross-layer consistency KL divergence and timing smooth KL divergence, dynamically adjust the distillation loss weight, and comprehensively weigh the resource utilization and distillation effect through the Pareto multi-objective optimization function.

Benefits of technology

It achieves the optimal balance between training convergence speed, computing resource utilization rate and distillation effect while ensuring the quality of knowledge learning, and improves the training efficiency and the adaptability of large models under limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119849594B_ABST
    Figure CN119849594B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of machine learning technology, and in particular, to an efficient large model training optimization method based on dynamic resource allocation and knowledge distillation. It comprises the following steps: identifying the knowledge bottleneck layer of the student model, obtaining the KL divergence by calculating the difference in attention distribution between the teacher and student models on the knowledge bottleneck layer, and using the KL divergence as the distillation loss term, then obtaining the resource utilization rate during the training process, combining the resource utilization rate with the weight of the distillation loss term, constructing a Pareto multi-objective optimization function, and using a proximal strategy optimization algorithm to train a resource allocation agent; when the contribution of a module of the teacher model to the student model is lower than a threshold, the module is frozen and its forward calculation is stopped. The method ensures that the model achieves an optimal balance between the training convergence speed, computing resource utilization, and distillation effect, which not only improves the training efficiency, but also enhances the adaptability of the large model under limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular to an efficient large model training optimization method based on dynamic resource allocation and knowledge distillation. Background Art

[0002] In the field of deep learning, with the continuous increase in model size, the development of technologies such as the Transformer architecture and Mixture of Experts (MoE) has made it easy for the number of parameters of trained large-scale models to exceed the trillion level. However, although these large-scale models have shown excellent performance in various tasks, they also bring significant challenges: on the one hand, the training and reasoning of large models require a lot of computing resources and time, which is impractical for resource-limited environments (such as mobile devices or edge computing); on the other hand, when deploying these large and complex models to actual application scenarios, they often face limitations in storage space, memory usage, and real-time performance.

[0003] In order to overcome these problems, knowledge distillation is widely used as an effective model compression technology. It achieves model lightweighting by transferring the knowledge of a large and complex teacher model to a smaller student model while retaining the performance of the original model as much as possible. However, when traditional knowledge distillation methods transfer knowledge between the teacher model and the student model, the learned knowledge is unstable due to the large architectural differences between the teacher model and the student model, and the imprecise expression of knowledge during the distillation process. Therefore, an efficient large model training optimization method based on dynamic resource allocation and knowledge distillation is designed. Summary of the invention

[0004] The purpose of the present invention is to provide an efficient large model training optimization method based on dynamic resource allocation and knowledge distillation, so as to solve the problem that the architectural differences between the teacher model and the student model proposed in the above background technology are large, and the expression of knowledge is not accurate enough in the distillation process, resulting in unstable learned knowledge.

[0005] To achieve the above object, the present invention aims to provide an efficient large model training optimization method based on dynamic resource allocation and knowledge distillation, comprising the following steps:

[0006] S1. Build a real-time resource game controller, dynamically allocate computing resources to the teacher model and the student model through the Nash equilibrium algorithm, and determine the resource allocation ratio based on the real-time training benefit ratio of the two, wherein the teacher model and the student model both contain a composite network structure of deep convolutional layers, dilated convolutional layers and fully connected layers;

[0007] S2, based on the dual criteria of gradient variance continuously lower than the historical mean threshold or layer output feature L2 norm continuously lower than the preset threshold, the attention distribution of the corresponding layer of the teacher model is transferred to the knowledge bottleneck layer, and the cross-layer knowledge transfer loss is calculated by KL divergence;

[0008] S3, build a dynamic Pareto multi-objective optimization function, integrate resource utilization indicators, distillation loss weight coefficients and remaining training time constraints, and drive the proximal strategy optimization algorithm to generate resource allocation strategies;

[0009] S4. Dynamically evaluate the contribution of each module of the teacher model to the student model, freeze the modules whose contribution is lower than the threshold, and simulate the output feature distribution of the frozen modules through a lightweight generative adversarial network.

[0010] As a further improvement of this technical solution, in S1: the benefit ratios of the teacher model and the student model are respectively: ;

[0011] ;

[0012] When the teacher model has a higher benefit ratio than the student model, more than 50% of computing resources will be allocated to it. Otherwise, the computing resources will be tilted towards the student model. When the two are equal, they will be allocated in a balanced ratio.

[0013] As a further improvement of the technical solution, the specific implementation of the dual criterion in S2 includes:

[0014] The gradient variance threshold is set to 30%-50% of the historical sliding window mean;

[0015] The L2 norm threshold of the layer output feature is set to 20%-40% of the preset feature norm baseline value.

[0016] As a further improvement of the technical solution, the method for generating the cross-layer knowledge transfer loss in S2 includes:

[0017] Extract the attention matrix of the teacher model and the student model at the knowledge bottleneck layer and ;

[0018] Calculate the weighted KL divergence loss value, where the weight coefficient is positively correlated with the importance score of the feature channel;

[0019] The KL divergence loss is dynamically added to the total loss function of the student model and fed back to the real-time resource game controller in step S1.

[0020] As a further improvement of the technical solution, the dynamic Pareto multi-objective optimization function in S3 includes:

[0021] Resource efficiency optimization item: build composite indicators based on GPU memory occupancy, CUDA core utilization and data throughput;

[0022] Knowledge transfer optimization item: adjust the distillation loss weight according to the exponential decay of training progress, and the decay rate is inversely correlated with the remaining training time;

[0023] Stability constraints: impose cross-layer attention distribution consistency loss and KL divergence change rate constraints for adjacent training batches.

[0024] As a further improvement of the technical solution, the implementation of the proximal strategy optimization algorithm in S3 includes:

[0025] Define the state space vector: including the current resource allocation ratio, model loss value, gradient statistics, hardware monitoring indicators and training stage identifier;

[0026] Define the action space vector: including the adjustment of the teacher / student model resource allocation ratio, the adjustment of the distillation loss weight, and the activation flag of the adversarial generation network;

[0027] Design a composite reward function: Integrate the Pareto optimization objective function value, resource saving reward, and knowledge transfer stability reward.

[0028] As a further improvement of the technical solution, the teacher model module contribution evaluation method in S4 includes:

[0029] Calculate the marginal contribution of each module to the reduction of the student model loss, and use the back-propagation path integral method to quantify the module influence;

[0030] Set a dynamic adjustment threshold. When the module contribution is lower than the threshold for N consecutive training cycles, the freezing operation is triggered, where N increases with the training progress.

[0031] As a further improvement of the technical solution, the feature simulation process of the lightweight generative adversarial network includes:

[0032] The generator network adopts a deep separable convolution structure, and the number of parameters is controlled at 15%-25% of the original module;

[0033] The discriminator network introduces a dynamic attention mechanism to adjust the feature similarity evaluation weights according to the current training stage of the student model;

[0034] An alternating training strategy is adopted to simultaneously optimize the generator and discriminator parameters during the freezing of module activations.

[0035] As a further improvement of this technical solution, an exception handling mechanism is also included:

[0036] When it is detected that the gradient amplitude of the student model exceeds the safety threshold, the resource allocation ratio is forcibly reset to teacher model: student model = 7:3;

[0037] During the model convergence plateau period, the distillation loss weight of the knowledge bottleneck layer is automatically increased to 1.5-2 times the baseline value.

[0038] As a further improvement of this technical solution, the implementation of the efficient large model training optimization method in a distributed training system includes:

[0039] Cross-node resource coordination: Synchronize resource allocation strategy tables of multiple computing nodes through distributed Nash equilibrium algorithm;

[0040] Layered distillation mechanism: implement inter-layer attention migration within nodes and implement model output soft target distillation across nodes

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] In this efficient large model training optimization method based on dynamic resource allocation and knowledge distillation, the cross-layer consistent KL divergence and temporal smoothing KL divergence are introduced to dynamically adjust the distillation loss weight, so that the student model can avoid learning unstable knowledge while ensuring the quality of knowledge learning. In addition, combined with Pareto multi-objective optimization, a comprehensive trade-off is made between task loss, distillation loss, computing resource utilization and data distribution stability to ensure that the model achieves the optimal balance between training convergence speed, computing resource utilization and distillation effect. This method not only improves training efficiency, but also enhances the adaptability of large models under limited resources, providing a new optimization strategy for efficient training of large-scale intelligent systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 The figure is a flow chart of the overall method of the present invention. DETAILED DESCRIPTION

[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0045] See also Figure 1 As shown, an efficient large model training optimization method based on dynamic resource allocation and knowledge distillation is provided, including the following steps:

[0046] S1. Build a real-time resource game controller, dynamically allocate computing resources to the teacher model and the student model through the Nash equilibrium algorithm, and determine the real-time training benefit ratio allocation of each model based on the real-time training benefit ratio of the teacher model and the student model; wherein both the teacher model and the student model include deep convolutional layers, dilated convolutional layers, and fully connected layers;

[0047] Nash equilibrium is a concept in game theory that describes a stable strategy combination, that is, when no participant can obtain greater benefits by unilaterally changing the strategy, the system reaches an equilibrium state. In our computing resource allocation problem:

[0048] Two game participants: teacher model (T) and student model (S)

[0049] Strategy:

[0050] Computing resource allocation (resources allocated to the teacher model);

[0051] Computing resource allocation (resources allocated to the student model);

[0052] Satisfy the total computing resource constraints: ;

[0053] income:

[0054] The benefit of the teacher model depends on the impact of computing resources on its training benefits;

[0055] The benefit of the student model depends on the impact of computing resources on its training benefits;

[0056] The goal of Nash equilibrium is to achieve a stable resource allocation state by making it impossible for either party to obtain better benefits by adjusting resource allocation alone.

[0057] Furthermore, deep convolutional layers are mainly used to automatically extract features of input data (such as images, audio, etc.). Through a series of convolution operations, it can capture local dependencies in space or time in the input data, and as the number of layers increases, it can gradually identify more abstract and complex patterns and features. This is crucial for the teacher model and the student model, because they need to perform subsequent learning and prediction based on these features;

[0058] The dilated convolution layer (also called the hole convolution layer) expands the receptive field by inserting "holes" between the convolution kernel elements without increasing the number of parameters or reducing the resolution. This enables the model to capture a wider range of information without significantly increasing the computational cost, and is particularly suitable for tasks that require consideration of broader contextual information. For both teacher and student models, the dilated convolution layer can help them better understand global information, thereby improving the expressiveness of the model.

[0059] The fully connected layer is usually located deeper in the network and is used to convert the distributed feature representations learned by the previous layers into specific outputs, such as category probabilities in classification tasks. Each neuron is connected to all neurons in the previous layer to achieve comprehensive processing of all features. In the teacher model and the student model, the fully connected layer acts as a bridge from low-level features to the final decision results, responsible for integrating information from the convolutional layer and making the final judgment or prediction.

[0060] In large model training, computing resources are usually limited, and knowledge distillation involves the collaborative training of the teacher model and the student model. In order to ensure maximum resource utilization, the Nash equilibrium algorithm is used to dynamically allocate computing resources between the teacher model and the student model. The resource allocation ratio is determined based on the real-time training benefit ratio, so that the teacher model and the student model can still achieve optimal learning efficiency when computing resources are limited. This allocation strategy will affect the computing power of the teacher model in the subsequent distillation process, thereby affecting the effect of knowledge distillation.

[0061] The training benefit ratio is the ratio of the product of the rate of decrease of the loss function and the gradient amplitude (that is, the contribution of the current computing resources to the training effect), as follows:

[0062] ;

[0063] ;

[0064] In the formula, is the training benefit ratio of the teacher model; is the training benefit ratio of the student model; is the loss change of the teacher model; is the game coefficient of the teacher model; is the game coefficient of the student model; is the loss change of the student model; is the L2 norm of the gradient magnitude of the teacher model; is the L2 norm of the gradient magnitude of the student model.

[0065] Based on the real-time training benefit ratio of the teacher model and the student model, the real-time training benefit ratio distribution of each model is determined as follows: When , the computing resources allocated to the teacher model are greater than those allocated to the student model; this indicates that the teacher model currently contributes more to the overall training and should be allocated more computing resources; on the contrary, when When , the computing resources allocated to the student model are greater than those allocated to the teacher model, which means that the student model needs more computing resources to better learn the knowledge of the teacher model. , the student model and the teacher model evenly distribute computing resources.

[0066] Initial training allocation The above resources are transferred to the teacher model to accelerate convergence. When the student model gradient variance is lower than the preset threshold, the resource transfer mechanism is started, and the resource allocation ratio is adjusted to the teacher model and the student model according to the linear attenuation strategy. ; Freeze the teacher model parameters in the later stages of training, retaining only The computing resources are used to fine-tune the teacher model output and pass the game coefficient , Optimize resource allocation formula;

[0067] S2, calculate the gradient variance and L2 norm of each layer of the student model to mark whether the layer is a knowledge bottleneck layer, introduce the attention distribution of the corresponding layer in the teacher model to the knowledge bottleneck layer marked in the student model, and obtain the KL divergence by calculating the difference in attention distribution between the teacher and student models on the knowledge bottleneck layer. The KL divergence is used as the distillation loss term and fed back to the resource game controller of S1; the output logits distillation loss is used for non-knowledge non-bottleneck layers to reduce unnecessary calculations;

[0068] In S2, the gradient is the signal used to update the model parameters during the optimization process, and the variance of the gradient can reflect the update stability of different layers during the model training process. A layer with a large gradient variance may indicate that the learning process of this layer is unstable or has difficulties. The L2 norm is usually used to measure the size of the gradient, reflecting the update amplitude of each layer during the training process. A larger L2 norm indicates that the parameter update of this layer is larger, which may help learning; while a smaller L2 norm may indicate that the layer contributes less to the overall optimization of the model;

[0069] By calculating the gradient variance and L2 norm, we can identify which layers show greater instability or smaller update amplitude during training. These layers are the "knowledge bottleneck layers" of the student model. Learning these layers is often the most difficult part and requires more guidance and optimization.

[0070] The knowledge bottleneck layer is based on the following conditions: if the gradient variance is continuously lower than the threshold of the gradient variance or the L2 norm is continuously lower than the threshold of the L2 norm, it is marked as a knowledge bottleneck layer; specifically: the gradient variance is lower than the threshold for 5 consecutive batches Or the L2 norm is below the threshold for 5 consecutive batches , it is determined to be the knowledge bottleneck layer;

[0071] The attention distribution matrix reflects the degree of attention paid by the teacher model to the input data at this layer. It can help the student model better understand which features are important and learn how to process the data.

[0072] KL divergence is used to measure the difference in attention distribution between the teacher model and the student model on the corresponding layer. The smaller the KL divergence, the closer the student model is to the learning method of the teacher model.

[0073] This KL divergence will be fed back to the resource game controller to adjust the allocation of computing resources. If the KL divergence is large (i.e. the student model is learning slowly), it may mean that the bottleneck layer needs more computing resources to accelerate learning. Therefore, the resource game controller will dynamically adjust the allocation ratio of computing resources based on this information to ensure that the student model can better learn the knowledge of the teacher model.

[0074] In S2, the specific steps of obtaining the KL divergence and using the KL divergence as the distillation loss term and feeding it back to the resource game controller of S1 are as follows;

[0075] Get the KL divergence:

[0076] ;

[0077] In the formula, is the KL divergence; is the total number of input features in the current layer. When calculating the KL divergence, the normalization term It can be used to ensure that the calculation process is not affected by the number of features, thereby avoiding excessive impact on KL divergence when the number of features is large; For the characteristics The weighting factor of , which indicates the importance of the feature to the learning process, can usually be obtained through some heuristic rules or through training; For the teacher model The attention distribution matrix on the features, which represents the attention distribution matrix of the teacher model on the input features at this layer The degree of "attention" of the teacher model is usually represented by the activation value or output of the layer. The attention distribution is a vector in which each element corresponds to a certain part of the input feature. This distribution reflects how the teacher model processes different parts of the input data; For the student model The attention distribution on the features also indicates the attention of the student model on the input features at this layer. The degree of “attention”;

[0078] The student model is usually smaller than the teacher model. Directly copying knowledge will lead to insufficient learning. Therefore, it is necessary to find the knowledge bottleneck layer, that is, the layer that contributes more to the training loss, and enhance its knowledge distillation. This can accurately enhance the learning of important layers, rather than applying knowledge distillation globally and indiscriminately to reduce the amount of calculation. The optimized distillation strategy will affect the way subsequent data is transmitted, making knowledge transfer more efficient.

[0079] in,

[0080] ;

[0081] In the formula, For the The L2 norm of the feature gradient; For the The variance of the feature gradient; For the The L2 norm of the feature gradient; For the The variance of the feature gradient; ; ;

[0082] Distillation loss term:

[0083] ;

[0084] In the formula, is the total distillation loss; For the The distillation loss weight of the layer; is the total number of distillation layers; ;

[0085] Feeding the total distillation loss back into the resource game controller, the training benefit ratio becomes:

[0086] ;

[0087] ;

[0088] In the formula, is the training benefit ratio of the optimized teacher model; is the weight coefficient of distillation loss in the training benefit ratio; is the training benefit ratio of the optimized student model;

[0089] S3. Obtain resource utilization during training, combine resource utilization with the weight of distillation loss term, construct Pareto multi-objective optimization function, and use proximal strategy optimization algorithm to train resource allocation agent based on Pareto multi-objective optimization function. S3 step further dynamically adjusts computing resources to optimize computing resource utilization and avoid inefficient training caused by fixed resource allocation method. The resource allocation strategy optimized in this step will directly affect the scale of teacher model calculation and the subsequent freezing mechanism.

[0090] The resource utilization in S3 includes the utilization of the teacher model, the utilization of the student model, and the utilization of other components during the training process;

[0091] ;

[0092] In the formula, for resource utilization; is the weight coefficient of GPU memory occupancy; is the GPU memory occupancy rate; is the weight coefficient for calculating the unit utilization; To calculate the unit utilization;

[0093] ;

[0094] In the formula, The video memory occupied by the teacher model; The video memory occupied by the student model; The video memory occupied by other intermediate variables; is the total GPU memory;

[0095] ;

[0096] In the formula, is the computation time of the teacher model; is the computation time of the student model; The calculation time for other operations; is the total computation time;

[0097] In S3, resource utilization is combined with the weight of the distillation loss term to construct a Pareto multi-objective optimization function. The specific steps are as follows:

[0098] Calculate the distillation loss weight:

[0099] ;

[0100] In the formula, is the weight coefficient of the distillation loss in the objective optimization function; is the upper limit of resource utilization; for resource utilization; Loss for the mission; Loss for the initial mission; The upper limit of training time; is the total training time; ; In the formula, For the The duration of the training round; is the total number of training rounds;

[0101] This formula dynamically adjusts the distillation loss weight , ensuring that distillation is strengthened when resources are sufficient and reducing the impact of distillation on calculation when resources are tight; when task loss When is larger, the weight of the distillation loss increases to enhance the guidance of the teacher model to the student model; when the training time approaches the upper limit, the weight of the distillation loss decreases to ensure that the training can be completed on time;

[0102] Construct a Pareto multi-objective optimization function:

[0103] ;

[0104] The constraints are:

[0105] ;

[0106] ;

[0107] In the formula, It is a Pareto multi-objective optimization function; is the resource utilization bonus coefficient; Allocate proportions for resources; is the resource limit;

[0108] In the process of constructing the Pareto multi-objective optimization function, the representation space of different layers of deep distillation has cross-layer inconsistency and drastic fluctuations in data distribution, which makes it difficult for the student model to correctly learn the knowledge of the teacher model. Therefore, in the process of constructing the Pareto multi-objective optimization function, the data distribution stability constraint is considered for optimization. The specific steps after optimization are as follows:

[0109] Calculate the cross-layer consistency KL divergence:

[0110] ;

[0111] In the formula, The cross-layer consistency KL divergence measures the difference in attention distribution between the teacher model and the student model at different layers, and is used to evaluate the consistency between model layers during the knowledge distillation process; For the The KL divergence of the layer; this divergence measures the consistency of the student model's learning of the teacher model's knowledge at different layers. If the KL divergence of some layers is too high, it means that the knowledge transfer of this layer is unstable. By increasing the optimization objective function The penalty term can guide the training to reduce the distribution difference across layers and improve the knowledge distillation effect;

[0112] Calculate the time series smoothing KL divergence:

[0113] ;

[0114] In the formula, It is the time series smoothing KL divergence, which measures the change in attention distribution between the current batch and the previous batch; is the KL divergence of the current batch; is the KL divergence of the previous batch; this divergence measures the stability of data distribution over time during the knowledge distillation process. If the KL divergence between different batches changes too much, it means that the data distribution is unstable, which may make it difficult for training to converge. By introducing The penalty term can reduce the drastic changes in data distribution between batches and make knowledge distillation smoother and more stable;

[0115] Calculate the distillation loss weight:

[0116] ;

[0117] In the formula, is the weight coefficient of the distillation loss after optimization in the objective optimization function; A hyperparameter to control the KL divergence of cross-layer consistency; It is a hyperparameter to control the KL divergence of time series smoothing; when the data distribution fluctuates greatly, the exponential term tends to 0, reducing the impact of distillation loss and preventing the student model from learning unstable knowledge.

[0118] Construct the optimized Pareto multi-objective optimization function:

[0119] ;

[0120] In the formula, is the optimized Pareto multi-objective optimization function; Hyperparameters that influence the optimization objective to control cross-layer consistency KL divergence; Hyperparameters that control the influence of the KL divergence optimization objective on time series smoothing;

[0121] The constraints are: ;

[0122] .

[0123] This optimization method combines data distribution stability and computing resource utilization to construct an adaptive optimization mechanism. By introducing cross-layer consistency KL divergence and temporal smoothing KL divergence, the distillation loss weight is dynamically adjusted so that the student model can avoid learning unstable knowledge while ensuring the quality of knowledge learning. In addition, combined with Pareto multi-objective optimization, a comprehensive trade-off is made between task loss, distillation loss, computing resource utilization and data distribution stability to ensure that the model achieves the optimal balance between training convergence speed, computing resource utilization and distillation effect. This method not only improves training efficiency, but also enhances the adaptability of large models under limited resources, providing a new optimization strategy for efficient training of large-scale intelligent systems.

[0124] In S3, the resource allocation agent is trained using a proximal strategy optimization algorithm based on the Pareto multi-objective optimization function. The specific steps are as follows:

[0125] S31. Initialize strategy parameters ; Initialize old strategy parameters; Initialize clipping parameters , ; Initialize resource utilization reward coefficient , The initialization step provides a starting point for subsequent policy updates. and resource utilization bonus coefficient The initialization ensures that the policy optimization can be carried out within a stable range, while avoiding training instability caused by excessive policy updates. The initialization process provides a good foundation for the learning and training of the agent by setting reasonable initial values;

[0126] S32. Execute the current strategy in the environment and collect trajectory data:

[0127] ;

[0128] In the formula, is the current state; For the current action, including resource allocation ratio and distillation weights ; is the current reward function, ; The core purpose of this stage is to collect environmental feedback and record trajectory data by actually executing the current strategy to form an experience replay pool for strategy optimization. Trajectory data includes state, action, reward, and next state, which provides actual data support for calculating the advantage function and objective function of the strategy. Through continuous sampling, the agent can gradually adjust the strategy to optimize the resource allocation and distillation loss weight adjustment;

[0129] S33. Calculate the objective function and update the policy parameters using the gradient descent method:

[0130] Calculate the objective function:

[0131] ;

[0132] In the formula, is the objective function; is the time step expectations; Take action under the current strategy probability; Take the same action as the old strategy probability; For the old strategy; For the current strategy; is the advantage function, which measures the improvement of the current strategy over the old strategy; Used to constrain the scope of strategy changes and ensure stable updates; the decision variables of the agent include teacher-student resource allocation ratio, distillation loss weight adjustment, etc., to ensure that the training process converges to the optimal solution; To trim the scope, prevent the strategy from changing too much, and maintain stable training;

[0133] Update policy parameters by gradient descent :

[0134] ;

[0135] In the formula, is the learning rate;

[0136] Objective Function By considering the probability ratio of the current strategy to the old strategy and the guidance of the advantage function, the strategy can be optimized in each iteration. We ensure that the policy update is not too drastic, avoid the risk of over-optimization, and maintain the stability and reliability of training. Using the gradient descent method to update the policy parameters can make the agent gradually approach the optimal solution, effectively optimizing the weights in the resource allocation and distillation process;

[0137] Calculating the advantage function in the objective function The specific steps are as follows:

[0138] Calculating cumulative rewards :

[0139] ;

[0140] In the formula, To represent the time step Start, go to the future time step End of cumulative bonus; ; For Reward At time step The discount factor decreases as the number of time steps increases; For the time step The immediate reward obtained from the environment. The reward here is obtained through a reward function, which usually represents the "performance" of the model under a certain behavior. The reward function is directly mapped to the optimization objective function Negative value of ; This formula calculates the time step To the end time step The cumulative reward is calculated by the discount factor The rewards at each time step are weighted so that rewards farther away from the current time step have less influence;

[0141] Calculate state value function :

[0142] ;

[0143] In the formula, To indicate the status The expected cumulative reward that can be obtained according to the current strategy reflects the average performance in a given state and estimates the long-term reward that the strategy may bring from this state; is the expected action, which represents the weighted average of all possible future state sequences. Usually, the expectation is calculated based on the probability distribution under the current strategy; the state value function It is used to evaluate the long-term performance of following the current strategy starting from a certain state. It assigns a value to each state, representing the total reward that can be expected from that state;

[0144] Calculate the advantage function :

[0145] ;

[0146] In the formula, is the advantage function, which means from the current state The advantage function is the additional reward for taking a specific action relative to the average behavior performed according to the strategy in the current state. In other words, the advantage function represents the average performance of an action relative to the strategy; the advantage function Calculate the difference between the actual reward for an action and the average reward expected from the policy in the current state. , it means that the return of taking this action is better than the average level; if , it means that the return of this action is below average;

[0147] S34. Copy the current strategy parameters to the old strategy parameters: ; By copying the current policy parameters to the old policy parameters, the agent can maintain a good policy comparison basis in the next update, making the calculation of the probability ratio and advantage function in the objective function more accurate, ensuring the continuity and stability of the training process;

[0148] S35. Repeat steps S32 to S34 until the strategy converges; by repeated sampling, calculating the objective function and updating the strategy, the agent can continuously make adjustments in the feedback and gradually optimize the resource allocation strategy and distillation loss weight. This process enables the agent to learn how to optimize model training under limited resources, improve resource utilization and training effects, and finally converge to an optimal resource allocation strategy and distillation strength. The repeated training process helps the agent to adaptively adjust its strategy until convergence is achieved.

[0149] S4. During the training of the resource allocation agent (freezing some modules before the teacher model fully guides the student model may lead to a decline in learning effect. Low-contribution modules should be gradually frozen after the training is stable), calculate the contribution value of each module of the teacher model to the performance of the student model. When the contribution of a module of the teacher model to the student model is lower than the threshold, freeze the module and stop its forward calculation (freeze the module with low contribution). For the frozen module in the teacher model, use a lightweight generative adversarial network to simulate its output distribution (to ensure that the student model can still learn key knowledge, while reducing the amount of calculation and improving training efficiency);

[0150] The specific steps of S4 are as follows:

[0151] S41. In the later training process, the contribution of different modules of the teacher model (such as feature extraction modules or attention mechanism modules of each layer) to the performance of the student model is evaluated in real time. By calculating the distillation loss contribution of the teacher model in a specific module, the influence of each module on the learning process of the student model is determined. This process is usually quantified by comparing the difference between the output of the teacher model and the output of the student model. For example, the KL divergence is used to calculate the contribution of each module to the error of the student model.

[0152] S42. Set a contribution threshold in advance. When the contribution value of a module of the teacher model to the student model is lower than the contribution threshold, it is determined that the module has no effect on the student model learning, and the parameters of the module are frozen, and the forward calculation of the module in training is stopped. This can be done through dynamic threshold adjustment, and continuous optimization based on training progress or other real-time indicators to ensure the rationality of the freezing decision. The frozen module will no longer participate in the gradient calculation, thereby reducing the computational burden and improving training efficiency. Through module-level freezing, unnecessary waste of computing resources is effectively reduced, the training speed is improved, and it is ensured that the student model can focus on learning more valuable knowledge;

[0153] S43. For the frozen modules, a lightweight generative adversarial network is used to simulate the output distribution of the frozen modules, wherein GAN is composed of a generator and a discriminator, the output of the frozen modules simulated by the adversarial network is generated by the generator, and the discriminator determines the similarity between the output of the frozen modules simulated by the adversarial network generated by the generator and the output of the real modules; in this way, the student model can continue to obtain information from the frozen modules of the teacher model. Using a lightweight generative adversarial network to replace the actual calculation of the frozen modules can not only ensure the transmission of information, but also effectively reduce the computational burden. This method avoids the negative impact of the frozen modules on the training results while maintaining the effect of knowledge distillation;

[0154] S44. Regularly evaluate the impact of frozen modules on the performance of the student model, and dynamically adjust the threshold or freezing strategy based on the evaluation results; if it is found that the performance of the student model is greatly affected after freezing a module, or the freezing strategy causes a decline in training results, unfreeze the module in a timely manner and resume its learning process. Through a regular performance feedback mechanism, the freezing strategy can be flexibly adjusted to ensure the adaptability and continuous optimization of model training;

[0155] This step further reduces the amount of computation of the teacher model and improves training efficiency. Since the generative adversarial network can learn the output distribution of the frozen module, the student model can still learn complete knowledge. This mechanism is combined with the previous sub-graph activation mechanism, so that the teacher model can flexibly adjust the calculation range during training to improve efficiency.

[0156] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and descriptions are only preferred examples of the present invention, and are not intended to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements all fall within the scope of the present invention to be protected.

Claims

1. An efficient large model training optimization method based on dynamic resource allocation and knowledge distillation, characterized in that: The steps include: S1. Build a real-time resource game controller, dynamically allocate computing resources to the teacher model and the student model through the Nash equilibrium algorithm, and determine the resource allocation ratio based on the real-time training benefit ratio of the two, wherein the teacher model and the student model both contain a composite network structure of deep convolutional layers, dilated convolutional layers and fully connected layers; S2, based on the dual criteria of gradient variance continuously lower than the historical mean threshold or layer output feature L2 norm continuously lower than the preset threshold, the attention distribution of the corresponding layer of the teacher model is transferred to the knowledge bottleneck layer, and the cross-layer knowledge transfer loss is calculated by KL divergence; In S1, the benefit ratios of the teacher model and the student model are: In the formula, is the training benefit ratio of the teacher model; is the training benefit ratio of the student model; is the loss change of the teacher model; is the game coefficient of the teacher model; is the game coefficient of the student model; is the loss change of the student model; is the L2 norm of the gradient magnitude of the teacher model; is the L2 norm of the gradient magnitude of the student model; When the teacher model has a higher benefit ratio than the student model, more than 50% of computing resources will be allocated to it. Otherwise, the student model will be allocated more computing resources. When the two are equal, they will be allocated in a balanced ratio. The method for generating the cross-layer knowledge transfer loss in S2 includes: Extract the attention matrix of the teacher model and the student model at the knowledge bottleneck layer and ; Calculate the weighted KL divergence loss value, where the weight coefficient is positively correlated with the importance score of the feature channel; The KL divergence loss is dynamically added to the total loss function of the student model and fed back to the real-time resource game controller of step S1; S3, build a dynamic Pareto multi-objective optimization function, integrate resource utilization indicators, distillation loss weight coefficients and remaining training time constraints, and drive the proximal strategy optimization algorithm to generate resource allocation strategies; S4. Dynamically evaluate the contribution of each module of the teacher model to the student model, freeze the modules whose contribution is lower than the threshold, and simulate the output feature distribution of the frozen modules through a lightweight generative adversarial network.

2. The efficient large model training optimization method based on dynamic resource allocation and knowledge distillation according to claim 1 is characterized in that: The specific implementation of the dual criteria in S2 includes: The gradient variance threshold is set to 30%-50% of the historical sliding window mean; The L2 norm threshold of the layer output feature is set to 20%-40% of the preset feature norm baseline value.

3. The efficient large model training optimization method based on dynamic resource allocation and knowledge distillation according to claim 2 is characterized in that: The dynamic Pareto multi-objective optimization function in S3 includes: Resource efficiency optimization item: build composite indicators based on GPU memory occupancy, CUDA core utilization and data throughput; Knowledge transfer optimization item: adjust the distillation loss weight according to the exponential decay of training progress, and the decay rate is inversely correlated with the remaining training time; Stability constraints: impose cross-layer attention distribution consistency loss and KL divergence change rate constraints for adjacent training batches.

4. The efficient large model training optimization method based on dynamic resource allocation and knowledge distillation according to claim 3 is characterized in that: The implementation of the proximal strategy optimization algorithm in S3 includes: Define the state space vector: including the current resource allocation ratio, model loss value, gradient statistics, hardware monitoring indicators and training stage identifier; Define the action space vector: including the adjustment of the teacher / student model resource allocation ratio, the adjustment of the distillation loss weight, and the activation flag of the adversarial generation network; Design a composite reward function: Integrate the Pareto optimization objective function value, resource saving reward, and knowledge transfer stability reward.

5. The efficient large model training optimization method based on dynamic resource allocation and knowledge distillation according to claim 4 is characterized in that: The teacher model module contribution evaluation method in S4 includes: Calculate the marginal contribution of each module to the reduction of the student model loss, and use the back-propagation path integral method to quantify the module influence; Set a dynamic adjustment threshold. When the module contribution is lower than the threshold for N consecutive training cycles, the freezing operation is triggered, where N increases with the training progress.

6. The efficient large model training optimization method based on dynamic resource allocation and knowledge distillation according to claim 5 is characterized in that: The feature simulation process of the lightweight generative adversarial network includes: The generator network adopts a deep separable convolution structure, and the number of parameters is controlled at 15%-25% of the original module; The discriminator network introduces a dynamic attention mechanism to adjust the feature similarity evaluation weights according to the current training stage of the student model; An alternating training strategy is adopted to simultaneously optimize the generator and discriminator parameters during the freezing of module activations.

7. The efficient large model training optimization method based on dynamic resource allocation and knowledge distillation according to claim 6 is characterized in that: It also includes exception handling mechanisms: When it is detected that the gradient amplitude of the student model exceeds the safety threshold, the resource allocation ratio is forcibly reset to teacher model: student model = 7:3; During the model convergence plateau period, the distillation loss weight of the knowledge bottleneck layer is automatically increased to 1.5-2 times the baseline value.

8. The efficient large model training optimization method based on dynamic resource allocation and knowledge distillation according to claim 7 is characterized in that: The implementation of efficient large model training optimization methods in distributed training systems includes: Cross-node resource coordination: Synchronize resource allocation strategy tables of multiple computing nodes through distributed Nash equilibrium algorithm; Layered distillation mechanism: implement inter-layer attention transfer within nodes and implement model output soft target distillation across nodes.

Citation Information

Patent Citations

  • CPU resource allocation optimization method in knowledge distillation period based on edge calculation

    CN118626249A

  • Model compression method, system and equipment based on dynamic adaptive distillation

    CN119337946A