Large model incremental training method and system based on dynamic sparsification

Through dynamic sparseness and multi-dimensional evaluation of parameter importance, the large model structure is adaptively adjusted and incremental training management is carried out, which solves the problems of inefficient and resource consumption of large models in the existing technology, and achieves good protection of efficient training, low resource consumption and model performance.

CN119669714BActive Publication Date: 2025-05-09HANGZHOU MICRO MACRO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510186056.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-09
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The existing large-model training optimization methods have problems such as low training efficiency, high computing resource consumption and large model performance losses, especially in incremental training scenarios, which are prone to catastrophic forgetting.

Method used

The incremental training method of large-model based on dynamic sparse is adopted to evaluate parameter importance in multiple dimensions, adaptive structural adjustment, build an empirical playback buffer pool with adjustable capacity, and propose multi-objective loss functions, and incremental training and management of large-model models, and use hybrid precision quantization strategies for model compression optimization.

Benefits of technology

It significantly improves the speed of large-scale model training, reduces computing and storage overhead, ensures the effect of compressed models, eliminates catastrophic forgetting problems, and optimizes the actual usage cost of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669714B_ABST
    Figure CN119669714B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of model training technology, and specifically relates to a large model incremental training method and system based on dynamic sparsification. The method includes: S1, using a multi-dimensional evaluation method to evaluate the importance of parameters of a large model to reflect the contribution of parameters to the performance of the large model; S2, using an adaptive structural adjustment method to dynamically sparsify the structure of the large model; S3, by constructing an experience replay buffer pool with adjustable capacity and proposing a multi-objective loss function, incremental training management of the large model; S4, using a mixed precision quantization strategy to perform model compression optimization on the large model. The present invention has the characteristics of being able to dynamically adjust the network structure by real-time evaluation of the importance of model parameters, thereby achieving efficient training while ensuring model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of model training, and specifically relates to a large model incremental training method and system based on dynamic sparsification. Background Art

[0002] With the rapid development of Large Language Models (LLMs), their applications in various fields are becoming more and more extensive. However, the training and fine-tuning of large models face the problem of huge computing resource consumption. Taking GPT-3 as an example, it takes thousands of GPUs and millions of dollars to fully train a model with 175 billion parameters. Even model fine-tuning often requires a large investment of computing resources.

[0003] At present, the common large model training optimization methods mainly include:

[0004] 1. Full Fine-tuning: Update all parameters of the model, which consumes a lot of computing resources and is prone to overfitting;

[0005] 2. LoRA (Low-Rank Adaptation): Reduces trainable parameters through low-rank matrix decomposition, but the fixed structure may limit the model's expressiveness;

[0006] 3. P-tuning: Only train the embedding of continuous prompt words, although the number of parameters is small, the effect is limited;

[0007] 4. Model pruning: The model is pruned statically using preset rules, which lacks flexibility.

[0008] However, the above methods have the following problems:

[0009] 1. The parameter update strategy is fixed and cannot be dynamically adjusted according to the task points;

[0010] 2. The model compression ratio needs to be preset manually and lacks adaptive capabilities;

[0011] 3. Catastrophic forgetting is prone to occur in incremental training scenarios;

[0012] 4. The performance of the model after compression is significantly lost.

[0013] Therefore, a training method is needed that can adaptively adjust the model structure and dynamically balance computing resources and model performance. Summary of the invention

[0014] The present invention aims to overcome the problems of low training efficiency, high consumption of computing resources and large model performance loss in the existing large model training optimization methods in the prior art, and provides a large model incremental training method and system based on dynamic sparsification, which can dynamically adjust the network structure by real-time evaluation of the importance of model parameters, thereby achieving efficient training while ensuring model performance.

[0015] In order to achieve the above-mentioned object of the invention, the present invention adopts the following technical solutions:

[0016] The large model incremental training method based on dynamic sparsification includes the following steps:

[0017] S1, uses a multi-dimensional evaluation method to evaluate the importance of parameters of the large model and reflect the contribution of parameters to the performance of the large model;

[0018] S2, uses an adaptive structural adjustment method to dynamically sparsify the large model structure;

[0019] S3, manages incremental training of large models by building an experience replay buffer pool with adjustable capacity and proposing a multi-objective loss function;

[0020] S4, adopts a mixed precision quantization strategy to perform model compression optimization on large models.

[0021] Preferably, step S1 comprises the following steps:

[0022] S11, evaluates parameter sensitivity by calculating the influence of parameters on the loss function; specifically, for the loss function In the parameters The sensitivity calculation at is approximated by the second-order Taylor expansion:

[0023] ;

[0024] in, Indicated in the parameter The loss function value at Representation parameters Small changes in is the loss function with respect to the parameter The first derivative of ; is the loss function with respect to the parameter The second derivative of

[0025] Defining parameters Sensitivity score for:

[0026] ;

[0027] in, Represents the absolute value of the first-order derivative, reflecting the direct impact of parameter changes on the loss function; It represents the absolute value of the second-order derivative, reflecting the impact of parameter changes on the stability of the model; is a balance factor, which is used to adjust the relative importance of the first-order term and the second-order term, and its value range is [0,1];

[0028] Use mini-batch data to further estimate the sensitivity score :

[0029] ;

[0030] Where B represents a mini-batch dataset; |B| represents the batch size; git and hit represent the first-order and second-order derivatives calculated on the batch V, respectively;

[0031] S12, the final sensitivity score of the parameter is smoothed by exponential sliding average:

[0032] ;

[0033] in, Indicates Step sensitivity score; is the smoothing coefficient, the value range is [0,1]; Indicates Step sensitivity score;

[0034] S13, the final sensitivity score of the parameter is smoothed by exponential sliding average, and then combined with the evaluation indicators of other dimensions to perform multi-dimensional weighted combination calculation; specifically, the parameter The comprehensive importance score Ii is calculated as follows:

[0035] ;

[0036] in, is the parameter sensitivity score;

[0037] is the activation value of the parameter affecting the score, and the calculation formula is: ;in, is the dataset size, For parameters In the sample The activation value on ;

[0038] Ti is the task-specific score, calculated according to the specific task type:

[0039] For classification tasks: ;

[0040] in, Representation parameters With category labels The mutual information between them is used to measure the parameters The degree of information contribution to the classification result; the greater the mutual information, the more significant the contribution of the parameter to the classification task;

[0041] For the build task: ;

[0042] in, Represents KL divergence, which is used to measure parameter The output distribution P is similar to the parameter-free The difference between the output distribution Q when P and Q are respectively The output probability distribution when ; the larger the KL divergence, the greater the parameter The more significant the impact on the generated results;

[0043] is the weight coefficient of each dimension, and satisfies ;

[0044] S14, after obtaining the comprehensive importance score Ii, the parameter classification is performed using the quantile-based adaptive threshold method, and the classification is as follows:

[0045] when , then the corresponding parameters are important parameters; among them, is the upper quartile of the score distribution, i.e. the 75% quantile;

[0046] when , then the corresponding parameter is a secondary parameter; is the lower quartile of the score distribution, i.e. the 25% quantile;

[0047] when , then the corresponding parameter is a non-critical parameter;

[0048] Introducing a dynamic adjustment mechanism:

[0049] ;

[0050] in, is the threshold at the current moment; is the historical threshold at time t-1; is the target threshold calculated based on current performance and resource constraints; is the smoothing factor, and its value range is [0,1.

[0051] Preferably, step S2 comprises the following steps:

[0052] S21, organize the large model parameters into equal-sized computation blocks , the size of each block is k×k; the side length k of each computational block is determined by the following formula:

[0053] ;

[0054] Where M is the total number of model parameters; N is the number of target calculation blocks; kmax is the maximum allowed block size; Indicates rounding up operation;

[0055] S22, for each calculation block The importance of:

[0056] ;

[0057] in, To calculate the number of parameters in the block; To calculate the parameters within the block The overall importance score of Represents the importance score of the i-th computational block.

[0058] Preferably, step S2 further comprises the following steps:

[0059] S23, design a dynamic sparsity calculation method to maintain the local continuity of the large model network structure; the specific calculation process is as follows:

[0060] S231, calculate the basic sparsity sb according to resource constraints:

[0061] sb=1-(Ravail / Rreq);

[0062] Among them, Ravail is the current available computing resources; Rreq is the resources required for complete training of the large model; the value range of sb is [0,1];

[0063] S232, considering the performance requirements of large models, adjust the basic sparsity:

[0064] ;

[0065] in, is the performance change; is the baseline performance level; is the performance sensitivity coefficient, with a value range of [0,1]; sp is the sparsity after considering the performance constraint;

[0066] S233, in order to maintain the local continuity of the network structure, the inter-layer correlation calculation is introduced:

[0067] ;

[0068] in, is the current layer index; For the layer The set of directly connected layers; For Layer The topological distance with layer j; is the distance attenuation coefficient; Presentation Layer The degree of connection with adjacent layers;

[0069] S234, based on the inter-layer correlation, calculate the actual sparsity of each layer and adjust the layer sparsity:

[0070] ;

[0071] in, For Layer The final sparsity of is the continuity protection coefficient, with a value range of [0,1];

[0072] S235, during the training process, the sparsity is dynamically updated through a sliding window:

[0073] ;

[0074] in, is the level sparsity at time t; is the level sparsity at the previous moment; is the target sparsity currently calculated; is the smoothing factor, with a value range of [0,1];

[0075] S236, to ensure the rationality of sparsification, set the following constraints for constraint condition checking:

[0076] Global constraints: ;

[0077] Local constraints: and ;

[0078] in, Indicates The sparsity of the next layer adjacent to the layer, Indicates The sparsity of the previous layer adjacent to the layer; L is the total number of network layers; smax is the maximum allowed sparsity; is the sparsity difference threshold of adjacent layers;

[0079] S24, based on the calculated block importance and target sparsity, performs structure pruning:

[0080] S241, sorting the calculation blocks in descending order according to the importance scores;

[0081] S242, before retention important blocks; m represents the total number of original computing blocks, s represents the target sparsity, and m(1-s) represents the number of important computing blocks that need to be retained; Indicates a round-down operation;

[0082] S243, performing sparse processing on the remaining blocks.

[0083] Preferably, step S2 further comprises the following steps:

[0084] S25, design a progressive weight inheritance mechanism to ensure the smoothness and performance stability of large model structure adjustment, which specifically includes the following steps:

[0085] S251, weight importance evaluation first evaluates the importance score of each weight :

[0086] ;

[0087] in, is the gradient of the loss function with respect to the weight W; represents norm operation; is a small constant to prevent division by zero; Reflects the contribution of weights to the performance of large models;

[0088] S252, calculate inheritance coefficient based on weight importance :

[0089] ;

[0090] in, is the sigmoid activation function; is the smoothing coefficient, with a value range of [0,1], which is used to control the impact of importance on inheritance strength; As the basic inheritance rate, ensure the minimum inheritance ratio; The value range of is [0,1];

[0091] S253, adopts different inheritance strategies for weights at different levels:

[0092] ;

[0093] in, is the current layer index; L is the total number of network layers; f(x) is the layer modulation function: ; is the level influence factor, with a value range of [0,1], which is used to control the inheritance strength of weights at different levels;

[0094] S254, weight update is performed using interpolation:

[0095] ;

[0096] in, is the updated weight; is the original weight; is the target weight; is the calculated inheritance coefficient;

[0097] S255, introduces momentum term to smooth weights and perform momentum adjustment:

[0098] ;

[0099] in, is the momentum value at the current moment; represents the momentum value at time t-1; is the momentum coefficient, with a value range of [0,1]; is the final weight value;

[0100] S256, dynamically adjusts inherited parameters based on performance changes during training:

[0101] ;

[0102] in, is the performance change; and These are all initial parameter values; and All are adjustment coefficients; and All are adjusted parameters;

[0103] S257, set the safety threshold of weight change for stability protection:

[0104] Maximum change: ;

[0105] Relative change constraint: ;

[0106] in, and r are preset thresholds;

[0107] S258, regularly evaluate the performance of the large model and perform structural recovery when necessary, including the following process:

[0108] Calculate the performance degradation ;

[0109] Among them, P0 and P1 represent the initial performance and current performance of the model respectively;

[0110] when Exceeding performance threshold The recovery mechanism is triggered when

[0111] Gradually restore the pruned important computing blocks;

[0112] Re-evaluate and adjust sparsity;

[0113] The structure is continuously updated during the training process. The specific process is as follows:

[0114] Re-evaluate the importance of computational blocks every T training steps;

[0115] Adjust the large model structure according to the new importance score;

[0116] Update weight inheritance coefficient;

[0117] Optimize computing resource allocation.

[0118] Preferably, step S3 comprises the following steps:

[0119] S31, design the following multi-objective loss function to simultaneously optimize the performance of the new task and maintain the original capabilities of the model:

[0120] ;

[0121] in, is the total loss function; , , is the weight coefficient of each loss term, and ; for mission-specific losses; is the knowledge distillation loss; To maintain losses for the structure;

[0122] The specific definitions of each loss function are as follows:

[0123] Mission Specific Losses Use corresponding loss functions for different types of tasks:

[0124] For classification tasks: ;

[0125] CrossEntropy represents the cross entropy loss function, which is used to measure the difference between the predicted probability distribution and the true label distribution;

[0126] y_pred represents the output probability distribution predicted by the model, and y_true represents the distribution of the true label;

[0127] For the build task: ;

[0128] NLL stands for Negative Log-Likelihood, which is used to measure the difference between the model-generated sequence and the target sequence;

[0129] For regression tasks: ;

[0130] MSE stands for Mean Squared Error, which is used to calculate the average square difference between the predicted value and the true value;

[0131] Knowledge Distillation Loss Used to maintain the performance of the model on the original task:

[0132] ;

[0133] in, is the output probability distribution of the teacher model (original model); is the output probability distribution of the student model (current model); T is the temperature parameter used to soften the probability distribution, T>1; is the balance factor, with a value range of [0,1]; KL is the Kullback-Leibler divergence;

[0134] Structure retention loss Used to maintain key structural features of the model:

[0135] ;

[0136] in, , , is the weight coefficient, and ;

[0137] Each sub-item includes feature similarity loss , Attention Consistency Loss and gradient similarity loss .

[0138] S32, during the training process, the loss weights are dynamically adjusted in the following way:

[0139] ;

[0140] in, Represents the weight vector of each loss item at time t. The weight value is normalized to the interval (0,1) through the softmax function, and the sum of all weights is ensured to be 1;

[0141] The calculation formula for weight update is:

[0142] ;

[0143] in, Represents the logarithmic value of the weight at time t; Represents the logarithmic value of the weight at time t-1; Represents the learning rate, which is used to control the step size of weight update, and its value range is (0,1]; Represents the gradient of the loss function calculated on the validation set with respect to the weight;

[0144] Buffer pool initialization builds a dynamically sized buffer pool ERB:

[0145] ;

[0146] in, are memory units for different tasks; kt is the number of current tasks; the capacity of each memory unit can be adjusted dynamically;

[0147] S33, calculate the corresponding importance score for each sample x :

[0148] ;

[0149] in, is the loss value of the sample; G(x) is the gradient magnitude score; D(x) is the diversity score based on the distance between the sample and the stored samples; , , is the weight coefficient, and ;

[0150] S34, dynamically adjust the capacity of each memory unit based on task importance :

[0151] ;

[0152] Among them, C is the upper limit of total capacity; Assign temperature parameters to capacity to control the hardness or softness of the assignment; is an indicator of the importance of the task;

[0153] S35, using stratified sampling method to select storage samples:

[0154] ;

[0155] in, Represents the probability of selecting sample x; Select temperature parameters for the samples to control the randomness of the selection;

[0156] S36, Design a progressive update strategy:

[0157] Formulate new sample entry rules, as follows:

[0158] When the buffer pool is not full, the sample is directly deposited;

[0159] When the buffer pool is full, samples are replaced with the following probabilities:

[0160] ;

[0161] in, represents the probability of replacing the stored old sample x_old with the new sample x_new; I(x_new) represents the importance score of the new sample; I(x_old) represents the importance score of the old sample; is the sigmoid function;

[0162] Establish a regular cleanup mechanism as follows:

[0163] Calculate sample redundancy ;

[0164] Among them, R(x) represents the average similarity between sample x and the samples stored in the memory unit; sim(x, ) represents the difference between sample x and stored sample The similarity between them; n represents the number of samples involved in the calculation; when R(x) exceeds the preset threshold, it indicates that the redundancy between the corresponding sample and the stored sample is high, and it is considered to be removed to maintain sample diversity;

[0165] S37, adopts stratified sampling method to train sampling strategy, including task-level sampling and sample-level sampling:

[0166] Task-level sampling: ;

[0167] Sample-level sampling: ;

[0168] in, is the task importance weight; is the sampling temperature parameter;

[0169] S38, design an adaptive learning rate adjustment strategy based on parameter importance, as follows:

[0170] For parameters , the corresponding learning rate It is calculated as follows:

[0171] ;

[0172] in, is the basic learning rate; For parameters Importance score; is the learning rate modulation coefficient, the value range is [0,1]; is a monotonically decreasing function.

[0173] Preferably, in step S4, the mixed precision quantization strategy specifically refers to:

[0174] Different parameters are given different numerical precisions, depending on their importance:

[0175] Important parameters maintain FP16 precision to ensure accuracy;

[0176] Secondary parameters are quantized using INT8 to save storage space;

[0177] Non-critical parameters are further reduced to INT4 precision.

[0178] The present invention also provides a large model incremental training system based on dynamic sparsification, including:

[0179] The parameter evaluation module is used to evaluate the importance of the parameters of the large model using a multi-dimensional evaluation method to reflect the contribution of the parameters to the performance of the large model;

[0180] Dynamic sparsification control module, used to dynamically sparsify the large model structure using an adaptive structural adjustment method;

[0181] The incremental training management module is used to manage the incremental training of large models by building an experience replay buffer pool with adjustable capacity and proposing a multi-objective loss function;

[0182] The optimization module is used to use a mixed precision quantization strategy to perform model compression optimization on large models.

[0183] Compared with the prior art, the present invention has the following beneficial effects: (1) In terms of training efficiency, the present invention significantly improves the speed of large model training; through dynamic sparsification and parameter importance evaluation mechanism, the calculation amount of non-essential parameters is reduced, so that the training time is shortened by more than 60%. At the same time, due to the adoption of a sparsification scheme based on computing blocks, the computing density is improved, the GPU utilization rate is increased by 35%, and the overall training throughput is increased by 2.5 times; (2) In terms of resource consumption, the present invention greatly reduces the computing and storage overhead; through mixed precision quantization and dynamic structure adjustment, the GPU video memory usage in the model training process is reduced by 50%, and the storage space requirement is reduced by 55%; especially in the multi-card training scenario, through the optimized parameter update strategy, the communication overhead is significantly reduced, making the distributed training more efficient. The rate is increased by 40%; (3) In terms of model performance, the present invention well guarantees the effect of the compressed model; through a carefully designed parameter importance evaluation mechanism and a progressive weight inheritance strategy, the performance loss of the core task is controlled within 3% when the model size is reduced by 60%; at the same time, through an effective incremental training mechanism, the adaptability of new tasks is significantly improved, and the catastrophic forgetting problem is basically eliminated; (4) In terms of deployment and application, the present invention optimizes the actual use cost of the model; through adaptive model compression technology, the requirements for hardware devices are reduced, so that the model can run on more popular hardware platforms; actual measurements show that the compressed model reduces the inference delay by 45%, and the memory usage during runtime is reduced by 50%, which greatly reduces the cost of model deployment and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS

[0184] Figure 1 A flowchart of the large model incremental training method based on dynamic sparsification in the present invention in practical application;

[0185] Figure 2 This is a percentage improvement effect diagram obtained by using the large model incremental training method based on dynamic sparsification in the present invention. DETAILED DESCRIPTION

[0186] In order to more clearly illustrate the embodiments of the present invention, the specific implementation methods of the present invention will be described below with reference to the accompanying drawings. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings and other implementation methods can be obtained based on these accompanying drawings without creative work.

[0187] The present invention provides a large model incremental training method based on dynamic sparsification, comprising the following steps:

[0188] 1. Use a multi-dimensional evaluation method to evaluate the importance of parameters of large models and reflect the contribution of parameters to the performance of large models;

[0189] 2. Adopt adaptive structural adjustment method to dynamically control the sparseness of large model structure;

[0190] 3. Manage incremental training of large models by building an experience replay buffer pool with adjustable capacity and proposing a multi-objective loss function;

[0191] 4. Use mixed precision quantization strategy to perform model compression optimization on large models.

[0192] The specific calculation process of each step is as described in the summary of the invention, and will not be elaborated in detail in the embodiment.

[0193] In addition, the present invention also provides a large model incremental training system based on dynamic sparsification, including:

[0194] The parameter evaluation module is used to evaluate the importance of the parameters of the large model using a multi-dimensional evaluation method to reflect the contribution of the parameters to the performance of the large model;

[0195] Dynamic sparsification control module, used to dynamically sparsify the large model structure using an adaptive structural adjustment method;

[0196] The incremental training management module is used to manage the incremental training of large models by building an experience replay buffer pool with adjustable capacity and proposing a multi-objective loss function;

[0197] The optimization module is used to use a mixed precision quantization strategy to perform model compression optimization on large models.

[0198] Based on the technical solution of the present invention, the implementation process of the present invention in practical application is illustrated by the following case scenario, and the specific application implementation plan is as follows:

[0199] like Figure 1 As shown, the method of the present invention takes the incremental training of the GPT-3 175B large-scale language model as an example for explanation during training. The entire training process is divided into three stages: the initial deployment stage, the incremental update stage, and the performance protection stage.

[0200] 1. Initial deployment phase

[0201] (1) Basic model preparation: GPT-3 175B is selected as the base model, batch_size=32, learning_rate=1e-5 is set, and the parameter importance evaluation module is initialized.

[0202] (2) Parameter importance evaluation: For the input data batch, the sensitivity of each parameter is calculated. The actual evaluation results show:

[0203] The sensitivity of the Transformer attention layer parameters is about 0.82;

[0204] The sensitivity of FFN layer parameters is about 0.65;

[0205] The Embedding layer parameter sensitivity is about 0.43;

[0206] (3) Initial structure tailoring, based on the sensitivity score, the parameters are divided into three categories:

[0207] Important parameters (sensitivity > 0.7): maintain FP16 accuracy;

[0208] Minor parameters (0.3≤sensitivity≤0.7): reduced to INT8 precision;

[0209] Non-critical parameters (sensitivity < 0.3): Reduce to INT4 precision or prune;

[0210] The actual compression effect shows that the model size is reduced from 350GB to 140GB, the computing resource requirements are reduced by 55%, and the inference delay is reduced from 200ms to 110ms.

[0211] 2. Incremental update phase

[0212] Take the newly added medical field data training as an example:

[0213] (1) Data access processing: collect 1 million medical conversation data, standardize the data format, and calculate the data feature distribution.

[0214] (2) Dynamic evaluation and update: Calculate the impact of the newly added medical data batch and perform layered updates based on the impact scores:

[0215] High impact layer (impact > 0.6): full update;

[0216] Medium impact layer (0.3<influence≤0.6): sparse update;

[0217] Low impact layer (impact ≤ 0.3): keep parameters frozen;

[0218] (3) Incremental training management. The training configuration uses batch_size=16, learning_rate=5e-6, and max_steps=10000. During the training process, the performance changes are checked every 100 steps and the update strategy is adjusted dynamically.

[0219] 3. Performance protection stage

[0220] (1) Performance monitoring: a comprehensive evaluation is performed every 1000 steps, including:

[0221] Calculate the accuracy of medical question answers;

[0222] Test general task performance;

[0223] Check resource usage;

[0224] (2) Structural recovery check: When the performance drops by more than 10%, the recovery mechanism is triggered, including:

[0225] Calculate the contribution of each layer to performance;

[0226] Recover accuracy layer by layer in order of importance;

[0227] Stop when performance recovers to more than 90% of baseline;

[0228] The actual recovery effect shows that the accuracy of medical question-and-answering has recovered from 82% to 93%, the performance of general tasks has remained above 95% of the baseline, and the additional resource overhead has increased by about 15%.

[0229] (3) Long-term stability protection, regular maintenance operations, including:

[0230] Conduct a full assessment once a week;

[0231] Dynamically adjust the sparsity threshold;

[0232] Update parameter importance scores;

[0233] After 6 months of actual operation verification, the system has shown good stability, including:

[0234] No noticeable performance degradation;

[0235] Dynamic sparsity is adaptively adjusted between 35%-65%;

[0236] Resource utilization increased by 40%;

[0237] The actual application data of this embodiment shows that compared with the traditional fixed-ratio pre-trained model pruning method, the present invention has achieved significant improvements in training efficiency, storage space, inference latency, etc., including:

[0238] Training efficiency increased by 2.5 times;

[0239] Save 60% of storage space;

[0240] Inference latency reduced by 45%;

[0241] The model performance loss is controlled within 3%.

[0242] These data fully demonstrate the practical value and technical advantages of the present invention in the large-model incremental training scenario.

[0243] like Figure 2 As shown, through a large number of experimental verifications, the method of the present invention has obvious advantages in the following aspects:

[0244] First, in terms of training efficiency, the speed of large model training has been significantly improved. Through dynamic sparsification and parameter importance evaluation mechanisms, the amount of calculation of non-essential parameters has been reduced, shortening the training time by more than 60%. At the same time, due to the use of a sparsification solution based on computing blocks, the computing density has been improved, the GPU utilization rate has been increased by 35%, and the overall training throughput has been increased by 2.5 times.

[0245] Secondly, in terms of resource consumption, the present invention significantly reduces computing and storage overhead. Through mixed precision quantization and dynamic structural adjustment, the GPU memory usage of the model training process is reduced by 50%, and the storage space requirement is reduced by 55%. Especially in the multi-card training scenario, the communication overhead is significantly reduced through the optimized parameter update strategy, which improves the efficiency of distributed training by 40%.

[0246] Third, in terms of model performance, the present invention well ensures the effect of the compressed model. Through a carefully designed parameter importance evaluation mechanism and a progressive weight inheritance strategy, the performance loss of the core task is controlled within 3% when the model size is reduced by 60%. At the same time, through an effective incremental training mechanism, the adaptability of new tasks has been significantly improved, and the catastrophic forgetting problem has been basically eliminated.

[0247] Finally, in terms of deployment and application, the present invention optimizes the actual use cost of the model. Through adaptive model compression technology, the requirements for hardware devices are reduced, allowing the model to run on more popular hardware platforms. Actual measurements show that the compressed model reduces inference latency by 45% and memory usage during runtime by 50%, which greatly reduces the cost of model deployment and maintenance.

[0248] The innovative features of the present invention are as follows:

[0249] 1. In terms of the dynamic evaluation method of parameter importance, a multi-dimensional evaluation mechanism is proposed, including a parameter sensitivity calculation method based on the second-order Taylor expansion, an influence analysis method based on the activation value distribution, and a task-specific evaluation index system. This evaluation system can fully and accurately reflect the contribution of parameters to model performance, providing a reliable basis for subsequent structural adjustments.

[0250] 2. In terms of the implementation mechanism of adaptive sparsification, the present invention designs a unique dynamic structure adjustment scheme, including an adaptive sparsity calculation method based on resource constraints, a local sparsification strategy based on computing blocks, and a progressive weight inheritance mechanism. These technologies ensure that the model structure can be dynamically adjusted according to actual needs, while improving the stability of the training process.

[0251] 3. In terms of anti-forgetting strategies for incremental training, this paper proposes innovative solutions, including an experience replay mechanism based on adjustable capacity, a differentiated parameter update strategy, and a loss function design for multi-objective optimization. These technologies effectively solve the problem of catastrophic forgetting in incremental learning, allowing the model to continue to learn new knowledge while maintaining its original capabilities.

[0252] 4. In terms of performance protection methods for model compression, the present invention has developed a complete set of technical systems, including a mixed precision quantization scheme based on parameter importance, a multi-level structural optimization technology, and a performance recovery mechanism. These technologies optimize the model size and computational efficiency while maximizing the model performance.

[0253] The above description is only a detailed description of the preferred embodiments and principles of the present invention. For ordinary technicians in this field, according to the ideas provided by the present invention, there will be changes in the specific implementation methods, and these changes should also be regarded as the protection scope of the present invention.

Claims

1. A large model incremental training method based on dynamic sparsification, characterized in that: It includes the following steps; S1. Adopt a multi-dimensional evaluation method to evaluate the parameter importance of the large model, reflecting the contribution of parameters to the performance of the large model; S2. Adopt an adaptive structure adjustment method to perform dynamic sparsity control on the large model structure; S3. Through constructing a capacity-adjustable experience replay buffer pool and proposing a multi-objective loss function, perform incremental training management on the large model; the incremental training management is based on 1 million formatted medical dialogue data; S4. Adopt a mixed-precision quantization strategy to optimize the model compression of the large model; Step S1 includes the following steps: S11. Evaluate the parameter sensitivity by calculating the influence degree of the parameter on the loss function; specifically, for the sensitivity calculation of the loss function L(θ) at the parameter θi, use the second-order Taylor expansion approximation: L(θi+Δθi)≈L(θi)+g i Δθi+(1 / 2)h i (Δθi) 2 ; Among them, L(θi) represents the loss function value at parameter θi, and Δθi represents the small change of parameter θi; is the first-order derivative of the loss function with respect to the parameter θi; is the second-order derivative of the loss function with respect to the parameter θi; Define the sensitivity score Si of the parameter θi as: Si=|g i |+λ|h i |; Among them, |g i | represents the absolute value of the first-order derivative, reflecting the direct impact of parameter changes on the loss function; |h i | represents the absolute value of the second-order derivative, reflecting the impact of parameter changes on model stability; λ is a balance factor, which is used to adjust the relative importance of the first-order term and the second-order term, and its value range is [0,1]; Use mini-batch data to further estimate the sensitivity score Si: Si=(1 / |B|)∑ V (|git|+λ|hit|); where B represents the mini-batch data set; |B| represents the batch size; git and hit respectively represent the first-order and second-order derivatives calculated on the batch V; S12. The final sensitivity score of the parameter is smoothed by exponential moving average: Si_p = αSi_p+(1-α)Si_(p-1); where Si_p represents the sensitivity score at the p-th step; α is the smoothing coefficient, and its value range is [0,1]; Si_(p-1) represents the sensitivity score at the (p-1)-th step; After the final sensitivity score of the parameter is smoothed by exponential moving average, combined with the evaluation indicators of other dimensions, perform multi-dimensional weighted combination calculation; specifically, the comprehensive importance score Ii of the parameter θi is calculated as follows: Ii = w1Si+w2Ai+w3Ti; where Si is the parameter sensitivity score; Ai is the activation value influence score of the parameter, and the calculation formula is: Ai=(1 / |D|)∑ x |a iu |; where |D| is the size of the data set, a iu is the activation value of parameter θi on sample x; Ti is the task-specific score, calculated according to the specific task type: For the classification task: Ti = MI(θi,Y); where MI(θi,Y) represents the mutual information between the parameter θi and the class label Y, used to measure the information contribution degree of the parameter θi to the classification result; the larger the mutual information, the more significant the contribution of the parameter to the classification task; For the generation task: Ti = KL(P||Q); where KL(P||Q) represents the KL divergence, used to measure the difference degree between the output distribution P with the parameter θi and the output distribution Q without the parameter θi; P and Q are the output probability distributions with and without the parameter θi respectively; the larger the KL divergence, the more significant the influence of the parameter θi on the generation result; w1, w2, w3 are the weight coefficients of each dimension, and satisfy w1+w2+w3 = 1; S14. After obtaining the comprehensive importance score Ii, use a quantile-based adaptive threshold method to classify the parameters, and the classification is as follows: When Ii>Q3, the corresponding parameter is an important parameter; where Q3 is the upper quartile of the score distribution, that is, the 75% quantile; When Q1≤Ii≤Q3, the corresponding parameter is a secondary parameter; Q1 is the lower quartile of the score distribution, that is, the 25% quantile; When Ii<Q1, the corresponding parameter is a non-critical parameter; Introduce a dynamic adjustment mechanism: τ_t=βτ_(t-1)+(1-β)τ*; Among them, τ_t is the threshold at the current moment; τ_(t-1) is the historical threshold at time t-1; τ* is the target threshold calculated based on the current performance and resource constraints; β is the smoothing factor, and its value range is [0,1].

2. The large model incremental training method based on dynamic sparsification according to claim 1 is characterized in that: Step S2 includes the following steps: S21, divide the large model parameters into equal-sized computational blocks {B1, B2, ..., B i }, the size of each block is k×k; the side length k of each computational block is determined by the following formula: Where M is the total number of model parameters; N is the number of target calculation blocks; kmax is the maximum allowed block size; Indicates rounding up operation; S22, for each calculation block B i The importance of: I(B i )=(1 / |B i |)∑ k I(θ k ); Among them, |B i | is the number of parameters in the calculation block; I(θ k ) is the parameter θ in the calculation block k The comprehensive importance score of i ) represents the importance score of the i-th computational block.

3. The large model incremental training method based on dynamic sparsification according to claim 2 is characterized in that: Step S2 also includes the following steps: S23, design a dynamic sparsity calculation method to maintain the local continuity of the large model network structure; the specific calculation process is as follows: S231, calculate the basic sparsity sb according to resource constraints: sb = 1 - (Ravail / Rreq); Among them, Ravail is the current available computing resources; Rreq is the resources required for complete training of the large model; the value range of sb is [0,1]; S232, considering the performance requirements of large models, adjust the basic sparsity: sp=sb*(1-λp*ΔP / P0) Among them, ΔP is the performance change; P0 is the baseline performance level; λp is the performance sensitivity coefficient, ranging from [0,1]; sp is the sparsity after considering the performance constraint; S233, in order to maintain the local continuity of the network structure, the inter-layer correlation calculation is introduced: C(l)=(1 / |N(l)|)∑ j∈N(l) exp(-d(l,j) / σ); Where l is the current layer index; N(l) is the set of layers directly connected to layer l; d(l,j) is the topological distance between layer l and layer j; σ is the distance decay coefficient; C(l) represents the degree of association between layer l and adjacent layers; S234, based on the inter-layer correlation, calculate the actual sparsity of each layer and adjust the layer sparsity: s(l)=sp*(1-μ*C(l)); Where s(l) is the final sparsity of layer l; μ is the continuity protection coefficient, ranging from [0,1]; S235, during the training process, the sparsity is dynamically updated through a sliding window: s_t(l)=βs_(t-1)(l)+(1-β)s*(l); Where s_t(l) is the level sparsity at time t; s_(t-1)(l) is the level sparsity at the previous time; s*(l) is the target sparsity calculated currently; β is the smoothing factor, and its value range is [0,1]; S236, to ensure the rationality of sparsification, set the following constraints for constraint condition checking: Global constraint: (1 / L)∑ l s(l)≤smax; Local constraints: |s(l)-s(l+1)|≤δs and |s(l)-s(l-1)|≤δs; Where s(l+1) represents the sparsity of the next layer adjacent to the lth layer, s(l-1) represents the sparsity of the previous layer adjacent to the lth layer; L is the total number of network layers; smax is the maximum allowed sparsity; δs is the sparsity difference threshold of adjacent layers; S24, based on the calculated block importance and target sparsity, performs structure pruning: S241, sorting the calculation blocks in descending order according to the importance scores; S242, before retention important blocks; m represents the total number of original computing blocks, s represents the target sparsity, and m(1-s) represents the number of important computing blocks that need to be retained; Indicates a round-down operation; S243, performing sparse processing on the remaining blocks.

4. The large model incremental training method based on dynamic sparsification according to claim 3 is characterized in that: Step S2 also includes the following steps: S25, design a progressive weight inheritance mechanism to ensure the smoothness and performance stability of large model structure adjustment, which specifically includes the following steps: S251, weight importance evaluation first evaluates the importance score E(W) of each weight: in, is the gradient of the loss function with respect to the weight W; ||·|| represents the norm operation; ε is a small constant to prevent division by zero; E(W) reflects the contribution of the weight to the performance of the large model; S252, based on the weight importance, calculate the inheritance coefficient γ: γ=σ(αE(W)+β h ) Among them, σ(·) is the sigmoid activation function; α is the smoothing coefficient, which ranges from [0,1] and is used to control the influence of importance on inheritance strength; β h is the basic inheritance rate to ensure the minimum inheritance ratio; the value range of γ is [0,1]; S253, adopts different inheritance strategies for weights at different levels: γ_l=γ*f(l / L) Where l is the current layer index; L is the total number of network layers; f(x) is the layer modulation function: f(x) = 1-λL|x-0.5|; λL is the layer influence factor, ranging from [0,1], which is used to control the inheritance strength of weights at different layers; S254, weight update is performed using interpolation: W_new=γ_l*W_old+(1-γ_l)*W_target; Among them, W_new is the updated weight; W_old is the original weight; W_target is the target weight; γ_l is the calculated inheritance coefficient; S255, introduces momentum term to smooth weights and perform momentum adjustment: m_t=U*m_(t-1)+(1-U)*(W_new-W_old)W_final =W_old+m_t Among them, m_t is the momentum value at the current moment; m_(t-1) represents the momentum value at time t-1; U is the momentum coefficient, ranging from [0,1]; W_final is the final weight value; S256, dynamically adjusts inherited parameters based on performance changes during training: α_t=α_0exp(-ρ*ΔP)β_t=β_0*(1+ξ*ΔP) Among them, ΔP is the performance change; α_0 and β_0 are both initial parameter values; ρ and ξ are both adjustment coefficients; α_t and β_t are both adjusted parameters; S257, set the safety threshold of weight change for stability protection: Maximum change: ||W_new-W_old||≤δmax; Relative change constraint: ||W_new-W_old|| / ||W_old||≤r; Among them, δmax and r are preset thresholds; S258, regularly evaluate the performance of the large model and perform structural recovery, which specifically includes the following processes: The computing performance decrease is δp = (P0-P1) / P0; Among them, P0 and P1 represent the initial performance and current performance of the model respectively; When δp exceeds the performance threshold δp_max, the recovery mechanism is triggered; Gradually restore the pruned important computing blocks; Re-evaluate and adjust sparsity; The structure is continuously updated during the training process. The specific process is as follows: Re-evaluate the importance of computational blocks every T training steps; Adjust the large model structure according to the new importance score; Update weight inheritance coefficient; Optimize computing resource allocation.

5. The large model incremental training method based on dynamic sparsification according to claim 4 is characterized in that: Step S3 includes the following steps: S31, design the following multi-objective loss function to simultaneously optimize the performance of the new task and maintain the original capabilities of the model: L total =λ1L task +λ2L distill +λ3L_struct; Among them, L total is the total loss function; λ1, λ2, λ3 are the weight coefficients of each loss term, and ∑λ i =1;L task is the task-specific loss; L_distill is the knowledge distillation loss; L_struct is the structure preservation loss; S32, during the training process, the loss weights are dynamically adjusted in the following way: λL_t=softmax(vL_t); Among them, λL_t represents the weight vector of each loss item at time t. The weight value is normalized to the interval (0,1) through the softmax function, and the sum of all weights is ensured to be 1; The calculation formula for weight update is: Where vL_t represents the logarithmic value of the weight at time t; vL_(t-1) represents the logarithmic value of the weight at time t-1; η represents the learning rate, which is used to control the step size of the weight update and has a value range of (0,1]; Represents the gradient of the loss function calculated on the validation set with respect to the weight; Buffer pool initialization builds a dynamically sized buffer pool ERB: ERB={M1,M2,...,M kt }; Among them, M kt are memory units for different tasks; kt is the number of current tasks; the capacity of each memory unit can be adjusted dynamically; S33, calculate the corresponding importance score I(x) for each sample x: I(x)=α1L(x)+α2G(x)+α3D(x); Where L(x) is the loss value of the sample; G(x) is the gradient magnitude score; D(x) is the diversity score based on the distance between the sample and the stored samples; α1, α2, α3 are weight coefficients, and ∑α i =1; S34, dynamically adjust the capacity Ci of each memory unit based on the task importance: Among them, C is the upper limit of total capacity; β c is the temperature parameter for capacity allocation, controlling the hardness or softness of allocation; Ij is the importance index of the task; S35, using stratified sampling method to select storage samples: P(select|x)=softmax(I(x) / τ s ); Among them, P(select|x) represents the probability of selecting sample x; τ s Select temperature parameters for the samples to control the randomness of the selection; S36, Design a progressive update strategy: Formulate new sample entry rules, as follows: When the buffer pool is not full, the sample is directly deposited; When the buffer pool is full, samples are replaced with the following probabilities: P(replace|x_old,x_new)=σ(I(x_new)-I(x_old)); Where P(replace|x_old,x_new) represents the probability of replacing the stored old sample x_old with the new sample x_new; I(x_new) represents the importance score of the new sample; I(x_old) represents the importance score of the old sample; σ(·) is the sigmoid function; Establish a regular cleanup mechanism as follows: Calculate sample redundancy R(x) = (1 / n)∑ i sim(x,x i ); Among them, R(x) represents the average similarity between sample x and the samples stored in the memory unit; sim(x,x i ) represents the difference between sample x and stored sample x i The similarity between them; n represents the number of samples involved in the calculation; when R(x) exceeds the preset threshold, it indicates that the redundancy between the corresponding sample and the stored sample is high, and it is considered to be removed to maintain sample diversity; S37, adopts stratified sampling method to train sampling strategy, including task-level sampling and sample-level sampling: S38, design an adaptive learning rate adjustment strategy based on parameter importance, as follows: For the parameter θ i , the corresponding learning rate η i It is calculated as follows: or i =η_base*f(κ*Ip); Among them, η_base is the basic learning rate; Ip is the parameter θ i The importance score of ; κ is the learning rate modulation coefficient, which ranges from [0,1]; f(·) is a monotonically decreasing function.

6. The large model incremental training method based on dynamic sparsification according to claim 5 is characterized in that: In step S4, the mixed precision quantization strategy specifically refers to: Different parameters are given different numerical precisions, depending on their importance: Important parameters maintain FP16 precision to ensure accuracy; Secondary parameters are quantized using INT8 to save storage space; Non-critical parameters are further reduced to INT4 precision.

7. A large model incremental training system based on dynamic sparsification, used to implement the large model incremental training method based on dynamic sparsification according to any one of claims 1 to 6, characterized in that: The large model incremental training system based on dynamic sparsification includes: The parameter evaluation module is used to evaluate the importance of the parameters of the large model using a multi-dimensional evaluation method to reflect the contribution of the parameters to the performance of the large model; Dynamic sparsification control module, used to dynamically sparsify the large model structure using an adaptive structural adjustment method; The incremental training management module is used to manage the incremental training of large models by building an experience replay buffer pool with adjustable capacity and proposing a multi-objective loss function; The optimization module is used to use a mixed precision quantization strategy to perform model compression optimization on large models.

Citation Information

Patent Citations

  • Three-dimensional target detection method and system based on cross-modal decoupling knowledge transfer

    CN118799665A

  • Method and system for training a neural network model using gradual knowledge distillation

    US20230222326A1