A model pruning method, device and medium
By iterative training and importance score evaluation of large language models and dynamically allocating pruning ratios, the calculation and storage problems of large language models when deploying in resource-constrained scenarios are solved, and the applicability of the model to edge devices and real-time systems is improved.
Patent Information
- Application Number
- CN202510671470.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-23
AI Technical Summary
When existing large language models are deployed in resource-constrained scenarios, the computational overhead and storage costs are high, and the global pruning method may lead to reduced model accuracy.
By obtaining text data, iterative training of the target big model, determine the importance scores of each module in different tasks, and dynamically allocate the pruning ratio according to the scores, and give priority to modules with low pruning importance, and retain modules with high importance.
While ensuring model accuracy, it reduces computing and storage resource overhead and expands the deployment scale of large models on edge devices and real-time systems.
Smart Images

Figure CN120181171B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deep learning technology, and in particular to a model pruning method, device, and medium. Background Art
[0002] With the rapid development of large language models (LLMs), they are widely used in various task scenarios, such as text generation, semantic understanding, and multi-task reasoning. However, the high computational overhead, storage costs, and energy consumption of current large language models limit their large-scale deployment in resource-constrained scenarios, such as edge devices and real-time systems.
[0003] To address these technical issues, large language models can currently be pruned using global pruning methods, reducing the number of model parameters and computational complexity, thereby lowering storage and energy requirements. Specifically, model pruning is achieved by removing model weights or structures based on the overall model weight distribution. However, this approach often ignores the functional heterogeneity of different modules and may over-prune key modules, resulting in reduced model accuracy.
[0004] Therefore, how to ensure the computational accuracy of large language models while reducing the overhead of model computing and storage resources and expanding the deployment scale in resource-constrained scenarios such as edge devices and real-time systems is an urgent problem that needs to be solved by technical personnel in this field. Summary of the Invention
[0005] In view of this, one aspect of the present application provides a model pruning method, the method comprising:
[0006] Acquire text data; wherein the text data includes data of multiple different tasks;
[0007] Iteratively training a pre-built target large model using the text data to obtain a training result;
[0008] Determining the importance scores of each module in the target large model in the multiple different tasks based on the training results;
[0009] According to the importance score, the pruning ratio of each module is determined; and based on the pruning ratio, each module is iteratively pruned until a preset iteration condition is reached; when the importance score is smaller, the pruning ratio is larger, indicating that the corresponding module has lower importance in the multiple different tasks.
[0010] Optionally, the text data includes first text data of a self-supervised learning task and second text data of multiple downstream tasks; the training results include self-supervised training results corresponding to the first text data and downstream task training results corresponding to the second text data; and determining, based on the training results, the importance scores of each module in the target large model in the multiple different tasks, includes:
[0011] Determining, according to the self-supervised training results, a self-learning task score of each module in the self-supervised learning task;
[0012] Determining, based on the downstream task training results, a multi-task comprehensive score of each module in the multiple downstream tasks;
[0013] The importance score is determined based on the self-learning task score and the multi-task comprehensive score; when the self-learning task score is larger, the importance score is larger; when the multi-task comprehensive score is larger, the importance score is larger.
[0014] Optionally, determining the self-learning task score of each module in the self-supervised learning task based on the self-supervised training result includes:
[0015] Obtaining the weight matrix of each module;
[0016] Determining a self-supervised loss function value of the target large model based on the self-supervised training result;
[0017] Determining the gradient of the self-supervised loss function value with respect to the weight matrix;
[0018] The self-learning task score is determined based on the gradient and the weight matrix; when the absolute value of the gradient and the weight matrix are larger, the self-learning task score is larger, indicating that the corresponding module is more important in the self-supervised learning task.
[0019] Optionally, determining, based on the downstream task training results, the multi-task comprehensive scores of each module in the multiple downstream tasks, includes:
[0020] Obtaining a weight matrix for each module and a task weight for representing the importance of each downstream task;
[0021] Determining downstream loss function values corresponding to different downstream tasks according to the downstream task training results;
[0022] Determining the gradient of the downstream loss function value with respect to the weight matrix;
[0023] Determining, according to the gradient and the weight matrix, a single-task score of each module under different downstream tasks; when the absolute value of the gradient and the weight matrix are larger, the single-task score is larger, indicating that the corresponding module is more important in the corresponding downstream task;
[0024] The task weights and the single-task scores are weighted and summed to obtain the multi-task comprehensive score.
[0025] Optionally, determining the pruning ratio of each module according to the importance score includes:
[0026] sorting the importance scores in ascending order;
[0027] Prune the modules corresponding to the first proportion at the front of the sorting result at a first pruning ratio; prune the modules corresponding to the second proportion in the middle of the sorting result at a second pruning ratio; and prune the modules corresponding to the third proportion at the end of the sorting result at a third pruning ratio;
[0028] The sum of the first proportion, the second proportion, and the third proportion is equal to 1, and there is no intersection between them; the first pruning ratio is greater than the second pruning ratio; and the second pruning ratio is greater than the third pruning ratio.
[0029] Optionally, determining the pruning ratio of each module according to the importance score includes:
[0030] Get the preset basic pruning ratio and target adjustment range;
[0031] Sorting the importance scores in ascending order to obtain a ranking value of each module in the sorting result;
[0032] The pruning ratio is determined according to the basic pruning ratio, the target adjustment range, and the ranking value; the smaller the ranking value, the smaller the importance score, and the larger the pruning ratio.
[0033] Optionally, the iterative pruning of each module based on the pruning ratio includes:
[0034] Determine a base loss function value of the target large model after the first training; and determine a current loss function value of the target large model;
[0035] Determining a loss function increase amount according to the basic loss function value and the current loss function value;
[0036] Determine whether the increase in the loss function is less than a threshold;
[0037] If less than, return to the step of iteratively training the pre-built target large model using the text data to obtain a training result;
[0038] If it is not less than, and the preset iteration condition is not met, the pruning ratio of the previous pruning is maintained, and the target adjustment amplitude is adjusted.
[0039] Optionally, adjusting the target adjustment range includes:
[0040] Obtain a preset adjustment coefficient and the historical adjustment range corresponding to the last pruning; the adjustment coefficient is greater than 0 and less than 1;
[0041] Determining a current amplitude adjustment amount according to the adjustment coefficient and the loss function increase;
[0042] Determining the target adjustment range according to the historical adjustment range and the current adjustment range; the greater the increase in the loss function, the greater the current adjustment range, and the smaller the target adjustment range;
[0043] Correspondingly, the preset iteration condition is that the target adjustment amplitude is smaller than a preset amplitude.
[0044] Another aspect of the present application provides a model pruning device, the device comprising:
[0045] A text data acquisition module, configured to acquire text data; wherein the text data includes data of multiple different tasks;
[0046] An iterative training module is used to iteratively train a pre-built target large model using the text data to obtain training results;
[0047] An importance score determination module, configured to determine the importance scores of each module in the target large model in the plurality of different tasks based on the training results;
[0048] The model pruning module is used to determine the pruning ratio of each module according to the importance score; and iteratively prune each module based on the pruning ratio until a preset iteration condition is met; when the importance score is smaller, the pruning ratio is larger, indicating that the corresponding module is less important in the multiple different tasks.
[0049] Another aspect of the present application provides a model pruning device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the steps of the model pruning method are implemented.
[0050] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the model pruning method when the program is executed by a processor.
[0051] The present application provides a model pruning method, device, and medium, which have the following beneficial effects: by determining the importance scores of different modules in the model under different tasks, the pruning ratio of each module can be dynamically allocated according to the importance score, ensuring that modules with small importance scores, low importance, and small contribution to the model processing tasks are allocated a larger pruning ratio, while modules with large importance scores and high importance are allocated a smaller pruning ratio, thereby avoiding global pruning of all modules of the large model and ignoring the functional differences between modules in different tasks, which leads to reduced model accuracy. While ensuring the computational accuracy of the large language model, the overhead of model computing resources and storage resources can be reduced through model pruning, which can expand the deployment scale of large models in resource-constrained scenarios such as edge devices and real-time systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 A schematic diagram of a flow chart of a model pruning method provided in an embodiment of the present application;
[0053] Figure 2 A schematic diagram of the structure of a model pruning system provided in an embodiment of the present application;
[0054] Figure 3 A schematic diagram of the principle of a model pruning method provided in an embodiment of the present application;
[0055] Figure 4 A schematic diagram of the interaction principle of a model pruning system provided in an embodiment of the present application;
[0056] Figure 5 A schematic structural diagram of a pruning device of a model provided in an embodiment of the present application;
[0057] Figure 6 A schematic structural diagram of a model pruning device provided in another embodiment of the present application.
[0058] The accompanying drawings are numbered as follows: 50 is a text data acquisition module, 51 is an iterative training module, 52 is an importance score determination module, 53 is a model pruning module, 60 is a memory, 61 is a processor, 62 is a display screen, 63 is an input and output interface, 64 is a communication interface, 65 is a power supply, 66 is a communication bus, 601 is a computer program, 602 is an operating system, and 603 is data. DETAILED DESCRIPTION
[0059] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0060] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0061] Figure 1 A flow chart of a model pruning method provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the method includes:
[0062] S10: Acquire text data; wherein the text data includes data of multiple different tasks;
[0063] S11: Iteratively train the pre-built target large model using text data to obtain training results;
[0064] In a specific embodiment, to avoid global pruning of all modules in a model, that is, pruning all modules in the model at the same rate, thereby ignoring the importance of modules in different tasks, resulting in over-pruning of important modules and reduced module accuracy, the present application first determines the importance of each module in the model before pruning the model to solve this technical problem.
[0065] Specifically, the pre-built target large model is iteratively trained using text data. Based on the iterative training results, the importance of each module in different tasks can be determined. It should be noted that the acquired text data includes data from multiple different tasks. This allows the importance of each module in the target large model to be determined for each task, allowing for real-time adjustments to the pruning ratio of each module.
[0066] In a specific embodiment, the multiple different tasks may include, but are not limited to, masked language modeling tasks, next sentence prediction tasks, mathematical reasoning tasks, text generation tasks, translation tasks, and question-answering tasks. It should be noted that the text data for different tasks can be obtained from open source databases or manually set. This application does not limit the method for obtaining text data. After obtaining the text data, it is preprocessed to improve the quality of the text data.
[0067] Furthermore, the pre-built target large model is iteratively trained using text data to obtain training results. The target large model may include, but is not limited to, the GPT (Generative Pre-trained Transformer) large model series, LaMDA (Language Model for Dialogue Applications), and Tongyi Qianwen.
[0068] In an optional embodiment, the target large model can be a large language model based on the Transformer architecture, and each module in the target large model refers to an independent computing unit with a clear function in the target large model. Correspondingly, each module in the large language model based on the Transformer architecture includes a self-attention module, a feedforward neural network module, a normalization module, and a word embedding module. In another optional embodiment, the text data of different tasks are input into the target large model in units of a specified number of samples for training, so as to more intuitively view the contribution of each module under different tasks.
[0069] S12: Based on the training results, determine the importance scores of each module in the target large model in multiple different tasks;
[0070] It will be appreciated that in specific embodiments, based on the input text data and the training output of the target large model, the model's loss function (i.e., the difference between the input and output) can be determined, thereby measuring the importance of each module in different tasks. Specifically, based on the training results, the importance scores of each module in the target large model for different tasks can be calculated.
[0071] In an optional embodiment, the module gradient can reflect the sensitivity of model parameters to the loss function. During model training, each module parameter will affect the final loss function. By analyzing the gradient information of different module parameters, the importance score of each module in different tasks can be determined.
[0072] It should be noted that, in an optional embodiment, the importance score is used to evaluate the contribution of each module in different tasks, that is, the importance of each module in different tasks. The higher the importance score, the higher the importance of the representation module in the task.
[0073] S13: Determine the pruning ratio of each module based on the importance score; and iteratively prune each module based on the pruning ratio until the preset iteration condition is met; the smaller the importance score, the larger the pruning ratio, indicating that the corresponding module has lower importance in multiple different tasks.
[0074] Furthermore, after determining the importance score of each module in step S12, the pruning ratio for each module can be determined based on the importance score. Specifically, it can be understood that when the importance score is smaller, it indicates that the corresponding module is less important in multiple different tasks. In this case, the model can be assigned a higher pruning ratio, thereby reducing the computing and storage resource utilization of the target large model. Conversely, to avoid reducing the accuracy of the target large model, modules with higher importance scores are assigned a lower pruning ratio, thereby ensuring the computational accuracy of the model.
[0075] In an optional embodiment, when determining the pruning ratios of different modules, a mapping relationship between importance scores and pruning ratios may be pre-constructed, thereby allocating pruning ratios to each module according to the mapping relationship.
[0076] Therefore, the model pruning method provided in the embodiment of the present application determines the importance scores of different modules in the model under different tasks, so as to dynamically allocate the pruning ratio of each module according to the importance score, ensuring that modules with small importance scores, low importance, and low contribution to the model processing tasks are allocated a larger pruning ratio, while modules with large importance scores and high importance are allocated a smaller pruning ratio, avoiding global pruning of all modules of the large model, ignoring the functional differences between modules in different tasks, and resulting in reduced model accuracy. While ensuring the computational accuracy of the large language model, the model pruning reduces the overhead of model computing resources and storage resources, which can expand the deployment scale of large models in resource-constrained scenarios such as edge devices and real-time systems.
[0077] In an optional embodiment, in order to improve the generalization ability of the target large model, its ability to understand language structure, and the training results of the target large model in different tasks, thereby improving the calculation accuracy of the pruning ratio of each module, the iterative training of the target large model is divided into two parts, specifically, including a self-supervised training part and a plurality of downstream task training parts.
[0078] Correspondingly, the text data includes first text data for a self-supervised learning task and second text data for multiple downstream tasks. In an optional embodiment, the first text data includes but is not limited to masked language modeling text data and next sentence prediction text data, and the second text data includes but is not limited to mathematical reasoning text data, text generation data, translation text data, and question-answering text data.
[0079] Self-supervised learning is an unsupervised learning method that uses the structure of the data to construct pseudo-labels, allowing the model to be trained without manually annotated data. In fact, self-supervised learning can be understood as a pre-training phase for the initially constructed target large model, so that the target large model can learn the knowledge in the first text data.
[0080] The downstream task can be understood as the stage of fine-tuning the target large model. The downstream task includes multiple different tasks in order to determine the importance of each module in the target large model in different downstream tasks.
[0081] It is understandable that different training results are obtained under different training tasks. Specifically, the training results include the self-supervised training results corresponding to the first text data and the downstream task training results corresponding to the second text data.
[0082] Figure 2 This is a structural diagram of a model pruning system provided in an embodiment of the present application. For ease of understanding, the following will be combined with Figure 2 For explanation. Figure 2 As shown, in an optional embodiment, the pruning method of the model provided by this application is Figure 2 The model pruning system shown is implemented. Specifically, the model pruning system includes a self-supervised learning unit, a cross-task learning unit and a model pruning unit.
[0083] In a specific embodiment, the self-supervised learning unit is used to perform self-supervised learning training on the target large model using the first text data, and the cross-task learning unit is used to fine-tune the target large model for multiple different downstream tasks using the second text data. Furthermore, the model pruning unit determines the pruning ratio for each module in the target large model based on the training results of the self-supervised learning unit and the cross-task learning unit, and prunes the model.
[0084] On this basis, as an optional embodiment, the importance scores of each module in the target large model in multiple different tasks are determined according to the training results, including:
[0085] According to the self-supervised training results, the self-learning task score of each module in the self-supervised learning task is determined;
[0086] According to the training results of downstream tasks, the multi-task comprehensive scores of each module in multiple downstream tasks are determined;
[0087] The importance score is determined based on the self-learning task score and the multi-task comprehensive score; the greater the self-learning task score, the greater the importance score; and the greater the multi-task comprehensive score, the greater the importance score.
[0088] It is understandable that, whether in self-supervised learning tasks or in multiple downstream tasks, each module in the target large model will show different degrees of contribution. Therefore, in order to more accurately determine the pruning ratio of each module, in a specific embodiment, the importance of each module in the self-supervised learning task and the downstream task is determined.
[0089] Specifically, the self-learning task score of each module in the self-supervised learning task is determined based on the self-supervised training results. Simultaneously, the multi-task comprehensive score of each model in different downstream tasks is determined based on the downstream task training results. It is understood that the higher the self-learning task score, the more important the representation module is in the self-supervised learning task, and the higher the multi-task comprehensive score, the more important the representation module is across multiple downstream tasks.
[0090] In an optional embodiment, when determining the final importance score of each module based on the self-learning task score and the multi-task comprehensive score, a first weight can be assigned to the self-learning task score and a second weight can be assigned to the multi-task comprehensive score, and the sum of the first weight and the second weight is equal to 1. Specifically, the importance score can be calculated according to formula (1):
[0091] (1)
[0092] in, The first The importance score of each module, is the first weight, is the second weight, where , used to balance the self-learning task score and the multi-task comprehensive score, that is, to adjust the importance score of the self-learning task score and the multi-task comprehensive score In a specific embodiment, the weight can be set according to actual business needs, for example, it can be set to 0.5. The first The self-learning task score of each module, The first The multi-task comprehensive score of each module.
[0093] Based on formula (1), in a specific embodiment, when the self-learning task score The larger the The more important a module is in the self-supervised learning task, the greater its corresponding importance score is, and the smaller the allocated pruning ratio is. The larger the The higher the comprehensive importance of a module in multiple downstream tasks, the greater the corresponding importance score and the smaller the allocated pruning ratio.
[0094] Figure 3 A schematic diagram of the principle of a model pruning method provided in an embodiment of the present application is shown as follows: Figure 3 As shown, in a specific embodiment, the first text data is input into the target large model for training. And based on the self-supervised training results, the self-learning task score can be calculated. At the same time, the second text data is input into the target large model for training, and the multi-task comprehensive score can be calculated based on the downstream task training results. . Further, according to the self-learning task score and multi-task comprehensive score , the importance scores of each module in self-supervised tasks and multiple downstream tasks can be calculated .
[0095] Therefore, the model pruning method provided in the embodiment of the present application uses a self-supervised learning unit to evaluate the importance of each model module to the current input first data in real time. At the same time, combined with the multi-downstream task perception of the cross-task learning unit, it realizes the dynamic adjustment of the pruning strategy, improves the adaptability of the target large model to different tasks, improves the accuracy of the pruning ratio distribution of each module, reduces the model calculation and storage resources, and ensures the calculation accuracy of the model.
[0096] In an optional embodiment, the self-learning task score of each module in the self-supervised learning task is determined based on the self-supervised training results, including:
[0097] Get the weight matrix of each module;
[0098] Determine the self-supervised loss function value of the target large model based on the self-supervised training results;
[0099] Determine the gradient of the self-supervised loss function with respect to the weight matrix;
[0100] The self-learning task score is determined based on the gradient and weight matrix. When the absolute value of the gradient and the weight matrix are larger, the self-learning task score is larger, indicating that the corresponding module is more important in the self-supervised learning task.
[0101] In a specific embodiment, the first text data is input into the target large model for pre-training to obtain the self-supervised training results, and the importance scores of each part of the model for model training are evaluated based on the self-supervised training results, that is, the self-learning task score is determined. .
[0102] It's understandable that in large models, gradients reflect the sensitivity of model parameters to the loss function. During model training, the parameters of each module influence the final loss function. By analyzing the gradient information of different module parameters, we can understand the importance of these modules in a specific task.
[0103] Therefore, in an optional embodiment, the self-learning task score is calculated When , the weight matrix of each module is obtained, and the gradient of the self-supervised loss function value to the weight matrix is determined based on the self-supervised training results. In order to calculate the self-learning task score based on the gradient and weight matrix Specifically, it can be calculated according to formula (2):
[0104] (2)
[0105] in, is the number of samples of the first text data input in a single self-supervised learning task, is the self-supervised loss function value, For the The weight matrix of each module, is the self-supervised loss function value Weight matrix The gradient, It means taking the absolute value of all elements in the result matrix and then finding the average value. is the gradient With the weight matrix Perform element-wise product.
[0106] In a specific embodiment, formula (2), that is, the self-learning task score Can measure the The module has a self-supervised loss function value The importance of The weight matrix of each module and the weight matrix The greater the absolute value of the gradient is, the more The more significant the contribution of each module to the loss. At the same time, through batch averaging, the global importance of weights can be analyzed. That is, the overall importance of model weights on the entire first text dataset can be evaluated through batch averaging, rather than just the importance of a single sample.
[0107] Therefore, in the model pruning method provided in the embodiment of the present application, the self-supervised learning unit introduces self-supervised learning tasks in the model training stage, such as masked language modeling and next sentence prediction, etc., by inputting the first text data into the target large model for training, and calculating the self-learning task scores of each module of the model that contributes to the training, thereby evaluating the importance of each part of the model.
[0108] In an optional embodiment, based on the downstream task training results, the multi-task comprehensive scores of each module in multiple downstream tasks are determined respectively, including:
[0109] Obtain the weight matrix of each module and the task weight used to characterize the importance of each downstream task;
[0110] According to the training results of downstream tasks, determine the downstream loss function values corresponding to different downstream tasks;
[0111] Determine the gradient of the downstream loss function value with respect to the weight matrix;
[0112] Based on the gradient and weight matrix, the single-task score of each module under different downstream tasks is determined. When the absolute value of the gradient and the weight matrix are larger, the single-task score is larger, indicating that the corresponding module is more important in the corresponding downstream task.
[0113] The task weights and single-task scores are weighted and summed to obtain the multi-task comprehensive score.
[0114] In order to evaluate the contribution of different modules in the target large model to different downstream tasks, in a specific embodiment, The second text data of different downstream tasks are input into the target large model for training, and the downstream task training results are obtained. Based on the downstream task training results, the multi-task comprehensive score of each module in multiple downstream tasks is calculated. .
[0115] Similarly, in a specific embodiment, it is necessary to obtain the weight matrix of each module At the same time, we also need to obtain task weights to characterize the importance of each downstream task. Understandably, in different practical scenarios, the target large model pays different attention to different tasks. To improve the large model's adaptability to different downstream tasks and further enhance model pruning accuracy, we assign different task weights to different downstream tasks.
[0116] Furthermore, based on the downstream task training results of different downstream tasks, the corresponding downstream loss function value is determined, and then the gradient of the downstream loss function value with respect to the weight matrix is determined. Thus, the single task score of each module in a single downstream task can be calculated based on the gradient and the weight matrix. Specifically, according to formula (3), it is calculated:
[0117] (3)
[0118] in, The first The module in Single-task scores in downstream tasks, For the The downstream loss function value of the downstream task.
[0119] In a specific embodiment, according to formula (3), in the In the downstream tasks, if The weight matrix of each module and the weight matrix The greater the absolute value of the gradient is, the more Modules for The more significant the contribution of the downstream loss function value of the downstream task is, the higher the single task score is. The bigger.
[0120] Furthermore, the task weight and single task score are weighted and summed to obtain the task comprehensive score , for details, see formula (4):
[0121] (4)
[0122] in, for The maximum value of the downstream tasks, For the The task weights of downstream tasks, where In a specific embodiment, if you prefer to retain the model effects of several downstream tasks, you can set a higher task weight for the corresponding downstream tasks. If each downstream task is equally important, then the task weight Both are set to 1.
[0123] Therefore, in the model pruning method provided in the embodiment of the present application, the cross-task learning unit simultaneously introduces multiple downstream tasks in the model training stage, inputs the second text data of different tasks into the model for training, calculates the single-task score contributed by each module of the model to the training of each downstream task, and calculates the multi-task comprehensive score of each module of the model for all tasks, thereby evaluating the importance of each module in multiple downstream tasks.
[0124] In an optional embodiment, determining the pruning ratio of each module according to the importance score includes:
[0125] Sort the importance scores in ascending order;
[0126] The modules corresponding to the first proportion of the sorted results are pruned at a first pruning ratio; the modules corresponding to the second proportion in the middle of the sorted results are pruned at a second pruning ratio; the modules corresponding to the third proportion at the end of the sorted results are pruned at a third pruning ratio;
[0127] The sum of the first proportion, the second proportion, and the third proportion is equal to 1, and there is no intersection between them; the first pruning ratio is greater than the second pruning ratio; and the second pruning ratio is greater than the third pruning ratio.
[0128] In an optional embodiment, the target large model is divided into multiple local modules, and different pruning operations are performed on each independent module. Specifically, modules with low importance are pruned first, that is, modules with small importance scores are pruned first.
[0129] Specifically, the importance score of each module Perform ascending sorting, that is, sorting from small to large, that is, sorting based on importance from low to high. In an optional embodiment, pruning ratios are allocated to each module based on a pre-set segmented pruning strategy. Specifically, in an optional embodiment, the modules corresponding to the first proportion (e.g., top 20%) of the sorting results are pruned at a first pruning ratio (e.g., the first pruning ratio is 25%), the modules corresponding to the second proportion in the middle of the sorting results (e.g., 60% of the middle part of the sorting results) are pruned at a second pruning ratio (e.g., the second pruning ratio is 15%), and the modules corresponding to the third proportion (e.g., bottom 20%) at the end of the sorting results are pruned at a third pruning ratio (e.g., the third pruning ratio is 5%).
[0130] It should be noted that, in the specific embodiment, the first proportion, the second proportion and the third proportion have no overlap, and the sum of the first proportion, the second proportion and the third proportion is equal to 1, ensuring that all modules can be allocated pruning proportions.
[0131] In an optional embodiment, a piecewise function strategy for pruning ratios is set. Specifically, a first pruning ratio is within a first pruning ratio range (e.g., 15%-25%), a second pruning ratio is within a second pruning ratio range (e.g., 5%-15%), and a third pruning ratio is within a third pruning ratio range (e.g., 0-5%). The first pruning ratio range, the second pruning ratio range, and the third pruning ratio range do not intersect, and any pruning ratio within the first pruning ratio range is greater than any pruning ratio within the second pruning ratio range, and any pruning ratio within the second pruning ratio range is greater than any pruning ratio within the third pruning ratio range.
[0132] That is to say, according to the ranking results, the modules with the first proportion can choose the specific first pruning ratio within the first pruning ratio range. Similarly, the modules with the second and third proportions can choose the specific first pruning ratio within the second and third pruning ratio ranges respectively. It should be noted that no matter whether it is the modules in the first proportion or the modules in the second and third proportions, the importance score should be met when choosing the specific pruning ratio. The higher the value, the larger the corresponding pruning ratio.
[0133] like Figure 3 As shown, in a specific embodiment, the importance score The larger the value, the smaller the allocated pruning ratio, and the more conservative the corresponding pruning strategy. The smaller it is, the larger the allocated pruning ratio is, and the more aggressive the corresponding pruning strategy is.
[0134] Of course, in an optional embodiment, the importance score is pre-established. The mapping relationship between the pruning ratio and the importance score is obtained Finally, the corresponding pruning ratio is selected according to the mapping relationship. Furthermore, the performance of the target large model is verified through the test set data, and the pruning ratio and mapping relationship of each module are adjusted according to the verification results.
[0135] Therefore, in the model pruning method provided in the embodiment of the present application, the model pruning unit combines the self-supervised learning unit and the cross-task learning unit to calculate the importance score of each module, and sorts all modules in the model according to the importance score from low to high, generates a pruning priority list, dynamically allocates the pruning ratio of each module according to the module pruning priority, and adopts a piecewise function strategy to perform modular pruning within each module.
[0136] In order to more accurately determine the pruning ratio of each module, as an optional embodiment, determining the pruning ratio of each module includes:
[0137] Get the preset basic pruning ratio and target adjustment range;
[0138] Sort the importance scores in ascending order to obtain the ranking value of each module in the ranking results;
[0139] The pruning ratio is determined based on the basic pruning ratio, target adjustment range, and ranking value. The smaller the ranking value, the smaller the importance score, and the larger the pruning ratio.
[0140] In a specific embodiment, the basic pruning ratio of the module is pre-set, and the target adjustment range of the pruning ratio can be adjusted. Sort in ascending order, and determine the ranking value of each module in the ranking result after sorting, where the ranking value range is 1 to .
[0141] Furthermore, the pruning ratio is determined according to the basic pruning ratio, the target adjustment range and the ranking value. Specifically, it can be calculated according to formula (5):
[0142] (5)
[0143] in, The first The pruning ratio of each module, is the basic pruning ratio. In an optional embodiment, the basic pruning ratio Can be set to 5%. For the The ranking value of each module in the ranking result, among which the module with a ranking value of 1 is the highest priority pruning module. is the total number of all modules in the target large model, that is, The maximum value. is the target adjustment amplitude. In an optional embodiment, the target adjustment amplitude It can be set to 10%, or it can be set according to actual business needs.
[0144] According to formula (5), in a specific embodiment, when the ranking value is smaller, the importance score is The smaller it is, the lower the importance of the representation module in the task, the higher the pruning ratio priority, and a higher pruning ratio can be allocated for pruning.
[0145] Therefore, the model pruning method provided in the embodiment of the present application can calculate the precise pruning ratio according to the actual importance of each module, accurately prune each module, ensure the accuracy of the model, and reduce the usage of model calculation and storage resources.
[0146] In an optional embodiment, iteratively pruning each module based on the pruning ratio includes:
[0147] Determine the base loss function value of the target large model after the first training; and determine the current loss function value of the target large model;
[0148] Determine the loss function increase amount based on the basic loss function value and the current loss function value;
[0149] Determine whether the increase in the loss function is less than the threshold;
[0150] If it is less than, return to the step of iteratively training the pre-built target large model through the text data to obtain the training results;
[0151] If it is not less than and the preset iteration condition is not met, the pruning ratio of the previous pruning is maintained and the target adjustment amplitude is adjusted.
[0152] In a specific embodiment, in order to further improve the accuracy of model pruning, the pruning ratio of each module is dynamically adjusted. Specifically, during the pruning process, the performance indicators of the model on the validation set are detected, and the pruning strategy is dynamically adjusted according to the changes in the performance indicators to ensure the stability of the model performance during the pruning process. Figure 2 As shown, the model pruning system also includes a feedback adjustment unit for dynamically adjusting the pruning ratio through a verification data set.
[0153] Therefore, in an optional embodiment, the text data also includes a verification text data set, such as Figure 3 As shown, in a specific embodiment, after determining the pruning ratio of each module based on the method of the above embodiment, the verification text dataset is input into the target large model to test the current performance indicators of the target large model. Specifically, the current loss function value of the target large model and the basic loss function value after the target large model is first trained using the first text data and the second text data can be calculated according to formula (6):
[0154] (6)
[0155] in, is the basic loss function value, is the current loss function value.
[0156] Furthermore, the increase in loss function is calculated according to formula (7):
[0157] (7)
[0158] in, is the amount of loss function increase. In a specific embodiment, the current loss function value And the basic loss function value obtained after the first training The smaller the difference, the corresponding loss function increases The smaller.
[0159] In an optional embodiment, as Figure 3 As shown, if the loss function increases If the value is less than the threshold, indicating that the current loss change is within an acceptable range, the pruning iteration can be continued, that is, the pre-built target large model is iteratively trained through text data to iteratively prune each module.
[0160] In another optional embodiment, if the loss function increases by If the loss function is not less than the threshold, the adjustment mechanism of the target adjustment range is triggered. Specifically, in an optional embodiment, when the loss function increases by When it is not less than the threshold, the target adjustment range can be reduced, thereby reducing the pruning ratio of each module.
[0161] Therefore, the feedback adjustment unit monitors the model's performance on the validation set during pruning and dynamically adjusts the pruning strategy based on performance changes to ensure model performance stability during pruning. If the model's performance on the validation set does not exceed the acceptable threshold, pruning continues. If it does, the model is rolled back to its state before the pruning attempt and the pruning strategy is adjusted until further pruning yields negligible benefits, at which point pruning is terminated.
[0162] Based on the above embodiment, as an optional embodiment, Figure 3 As shown, the loss function increases If it is not less than the threshold, it is first determined whether the preset iteration condition is met, that is, whether the target adjustment range is less than the preset range (for example, the preset range is 0.5%).
[0163] In an optional embodiment, if the target adjustment range is not less than the preset range, the pruning ratio of the last pruning is maintained and the target adjustment range is adjusted. Specifically, in an optional embodiment, adjusting the target adjustment range includes:
[0164] Get the preset adjustment coefficient and the historical adjustment range corresponding to the last pruning; the adjustment coefficient is greater than 0 and less than 1;
[0165] Determine the current amplitude adjustment amount based on the adjustment coefficient and the increase in the loss function;
[0166] The target adjustment range is determined based on the historical adjustment range and the current adjustment range. The greater the increase in the loss function, the greater the current adjustment range and the smaller the target adjustment range.
[0167] In a specific embodiment, the target adjustment range is calculated according to formula (8):
[0168] (8)
[0169] in, Adjust the amplitude for the target. is the historical adjustment range corresponding to the last pruning, is an adjustment coefficient. In a specific embodiment, the adjustment coefficient is greater than 0 and less than 1. For example, it can be set to 0.5, indicating that the pruning rate decreases by 0.5% every time the loss exceeds the threshold by 1%. is the current amplitude adjustment amount.
[0170] In another optional embodiment, as Figure 3 As shown, if the target adjustment range is smaller than the preset range, for example, , indicating that the benefit of the current pruning can be ignored, that is, the preset iteration condition is reached, and the model pruning can be ended at this time.
[0171] Therefore, the self-supervised learning unit introduces self-supervised tasks during the model inference phase to evaluate the weights of various model components. The cross-task learning unit introduces multiple downstream tasks during the model training phase to evaluate the weights of various model components in different tasks. The model pruning unit prunes the model based on the evaluation results of the self-supervised learning unit and the cross-task learning unit. The feedback adjustment unit monitors model performance during the pruning process and dynamically adjusts the pruning strategy based on performance changes to ensure the stability of model performance during the pruning process. The various modules collaborate with each other to achieve dynamic pruning of large language models based on the feedback adjustment mechanism.
[0172] Figure 4 This is a schematic diagram of the interaction principle of a model pruning system provided in an embodiment of the present application. In order to facilitate those skilled in the art to better understand the pruning method of the model provided in this application, the following will be combined with Figure 4 Further explanation.
[0173] like Figure 4 As shown in the figure, the model pruning system includes a self-supervised learning unit, a cross-task learning unit, a model pruning unit, and a feedback adjustment unit. The module pruning unit includes an importance score determination unit and a pruning ratio determination unit. The importance score determination unit is used to calculate the importance score of each module in the target large model for different tasks. The pruning ratio determination unit is used to determine the pruning ratio of each module based on the importance score and prune each module based on the pruning ratio.
[0174] In a specific embodiment, Figure 4 As shown, the data input layer of the model pruning system obtains text data, specifically including first text data used by the self-supervised learning unit for self-supervised learning tasks and second text data used by the cross-task learning unit for training multiple downstream tasks. In an optional embodiment, the data input layer reads and preprocesses the text data to obtain high-quality first and second text data.
[0175] Furthermore, the self-supervised learning unit performs a self-supervised learning task on the target large model using the first text data and, based on the self-supervised training results, calculates the self-learning task score of each module in the self-supervised learning task. Simultaneously, the cross-task learning unit fine-tunes the target large model for different downstream tasks using the second text data and, based on the downstream task training results, determines the multi-task composite score of each module in multiple downstream tasks.
[0176] Furthermore, the importance score in the module pruning unit is used to calculate the importance score of each module according to the self-learning task score and the multi-task comprehensive score, and the current pruning ratio is determined by the pruning ratio determination module, and the model is pruned according to the pruning ratio.
[0177] During the specific pruning process, after the model pruning unit prunes the target large model, it provides the pruned large model to the feedback adjustment unit. The feedback adjustment unit monitors the performance indicators of the model in real time, determines the pruning adjustment strategy based on the performance monitoring results, and returns the pruning adjustment strategy to the model pruning unit so that the model pruning unit can dynamically adjust the pruning strategy.
[0178] In this way, a closed-loop pruning is formed for the model. Specifically, by quantifying the importance score of each module in the large model for self-supervised learning tasks and various different tasks, a piecewise function strategy is adopted to perform differentiated pruning based on the importance score of each module in the large model. During the pruning process, the performance fluctuation of the model on the validation set is monitored in real time, and the pruning ratio is optimized through the gradient-weight joint formula, so as to achieve dynamic compression of the model parameters while ensuring the model effect.
[0179] In the above embodiments, the model pruning method is described in detail. The present application also provides a corresponding embodiment of a model pruning device.
[0180] Figure 5 A schematic diagram of the structure of a pruning device of a model provided in an embodiment of the present application is shown as follows: Figure 5 As shown, the device includes:
[0181] The text data acquisition module 50 is used to acquire text data; wherein the text data includes data of multiple different tasks;
[0182] Iterative training module 51 is used to iteratively train the pre-built target large model using text data to obtain training results;
[0183] Importance score determination module 52, used to determine the importance score of each module in the target large model in multiple different tasks based on the training results;
[0184] The model pruning module 53 is used to determine the pruning ratio of each module according to the importance score; and iteratively prune each module based on the pruning ratio until the preset iteration condition is met; when the importance score is smaller, the pruning ratio is larger, indicating that the corresponding module is less important in multiple different tasks.
[0185] In addition, the pruning device of the model provided in the embodiment of the present application also includes:
[0186] A self-learning task score determination module is used to determine the self-learning task score of each module in the self-supervised learning task according to the self-supervised training results;
[0187] A multi-task comprehensive score determination module is used to determine the multi-task comprehensive scores of each module in multiple downstream tasks based on the downstream task training results;
[0188] The first determination module is used to determine the importance score according to the self-learning task score and the multi-task comprehensive score; when the self-learning task score is larger, the importance score is larger; when the multi-task comprehensive score is larger, the importance score is larger.
[0189] The weight matrix acquisition module is used to obtain the weight matrix of each module;
[0190] A self-supervised loss function value determination module is used to determine the self-supervised loss function value of the target large model based on the self-supervised training results;
[0191] Gradient determination module, used to determine the gradient of the self-supervised loss function value with respect to the weight matrix;
[0192] The second determination module is used to determine the self-learning task score based on the gradient and the weight matrix. When the absolute value of the gradient and the weight matrix are larger, the self-learning task score is larger, indicating that the corresponding module is more important in the self-supervised learning task.
[0193] The weight matrix acquisition module is also used to obtain the weight matrix of each module and the task weight used to characterize the importance of each downstream task;
[0194] The downstream loss function value determination module is used to determine the downstream loss function values corresponding to different downstream tasks based on the downstream task training results;
[0195] The gradient determination module is also used to determine the gradient of the downstream loss function value with respect to the weight matrix;
[0196] The single-task score determination module is used to determine the single-task score of each module under different downstream tasks based on the gradient and weight matrix. When the absolute value of the gradient and the weight matrix are larger, the single-task score is larger, indicating that the corresponding module is more important in the corresponding downstream task.
[0197] The third determination module is used to perform weighted summation of task weights and single task scores to obtain a multi-task comprehensive score.
[0198] Sorting module, used to sort the importance scores in ascending order;
[0199] The module pruning module is used to prune the modules corresponding to the first proportion at the front of the sorting result at a first pruning ratio; prune the modules corresponding to the second proportion in the middle of the sorting result at a second pruning ratio; and prune the modules corresponding to the third proportion at the end of the sorting result at a third pruning ratio; wherein the sum of the first proportion, the second proportion, and the third proportion is equal to 1, and there is no intersection between them; the first pruning ratio is greater than the second pruning ratio; and the second pruning ratio is greater than the third pruning ratio.
[0200] The pruning parameter acquisition module is used to obtain the preset basic pruning ratio and target adjustment range;
[0201] The ranking value acquisition module is used to sort the importance scores in ascending order and obtain the ranking value of each module in the ranking results;
[0202] The pruning ratio determination module is used to determine the pruning ratio based on the basic pruning ratio, the target adjustment range and the ranking value; the smaller the ranking value, the smaller the importance score, and the larger the pruning ratio.
[0203] A loss value determination module is used to determine the basic loss function value of the target large model after the first training; and to determine the current loss function value of the target large model;
[0204] A loss function increase amount determination module is used to determine the loss function increase amount based on the basic loss function value and the current loss function value;
[0205] The judgment module is used to determine whether the increase in the loss function is less than the threshold; if it is less than, the iterative training module is called; if it is not less than, and the preset iteration conditions are not met, the pruning ratio of the previous pruning is maintained, and the target adjustment range is adjusted.
[0206] The pruning parameter acquisition module is also used to obtain the preset adjustment coefficient and the historical adjustment range corresponding to the previous pruning; the adjustment coefficient is greater than 0 and less than 1;
[0207] A current amplitude adjustment amount determination module is used to determine the current amplitude adjustment amount according to the adjustment coefficient and the loss function increase;
[0208] The target adjustment amplitude determination module is used to determine the target adjustment amplitude based on the historical adjustment amplitude and the current amplitude adjustment amount; the greater the increase in the loss function, the greater the current amplitude adjustment amount, and the smaller the target adjustment amplitude; correspondingly, the preset iteration condition is that the target adjustment amplitude is less than the preset amplitude.
[0209] Figure 6 A schematic structural diagram of a model pruning device provided in another embodiment of the present application is shown in FIG. Figure 6 As shown, the pruning device of the model includes: a memory 60 for storing a computer program;
[0210] The processor 61 is configured to implement the steps of the model pruning method mentioned in the above embodiment when executing a computer program.
[0211] The pruning device of the model provided in this embodiment may include but is not limited to a laptop computer or a desktop computer.
[0212] Among them, the processor 61 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 61 can be implemented in at least one hardware form of a digital signal processor (DSP), a field programmable gate array (FPGA), and a programmable logic array (PLA). The processor 61 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 61 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 61 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.
[0213] The memory 60 may include one or more computer-readable storage media, which may be non-transitory. The memory 60 may also include high-speed random access memory, and non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 60 is used to store at least the following computer program 601, wherein, after the computer program is loaded and executed by the processor 61, it can implement the relevant steps of the model pruning method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 60 may also include an operating system 602 and data 603, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 602 may include Windows, Unix, Linux, etc. The data 603 may include but is not limited to the relevant data involved in the model pruning method, etc.
[0214] In some embodiments, the model pruning device may further include a display screen 62 , an input / output interface 63 , a communication interface 64 , a power supply 65 , and a communication bus 66 .
[0215] Those skilled in the art will understand that Figure 6 The structure shown in the figure does not constitute a limitation on the pruning device of the model, and may include more or fewer components than shown in the figure.
[0216] The model pruning device provided in an embodiment of the present application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the model pruning method in the above embodiment.
[0217] It should be noted that although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that these operations be performed in the particular order shown or performed sequentially, or that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products.
Claims
1. A model pruning method, characterized in that: The method comprises: Acquire text data; wherein the text data includes data of multiple different tasks; Iteratively training a pre-built target large model using the text data to obtain training results; the text data includes first text data for a self-supervised learning task and second text data for multiple downstream tasks; the training results include self-supervised training results corresponding to the first text data and downstream task training results corresponding to the second text data; Determining the importance scores of each module in the target large model in the multiple different tasks based on the training results; Determining a pruning ratio for each module according to the importance score; and iteratively pruning each module based on the pruning ratio until a preset iteration condition is met; when the importance score is smaller, the pruning ratio is larger, indicating that the corresponding module is less important in the multiple different tasks; Determining the importance scores of each module in the target large model in the multiple different tasks based on the training results includes: Determining, according to the self-supervised training results, a self-learning task score of each module in the self-supervised learning task; Determining, based on the downstream task training results, a multi-task comprehensive score of each module in the multiple downstream tasks; The importance score is determined based on the self-learning task score and the multi-task comprehensive score; when the self-learning task score is larger, the importance score is larger; when the multi-task comprehensive score is larger, the importance score is larger.
2. The pruning method of the model according to claim 1, characterized in that Determining the self-learning task score of each module in the self-supervised learning task according to the self-supervised training result includes: Obtaining the weight matrix of each module; Determining a self-supervised loss function value of the target large model based on the self-supervised training result; Determining the gradient of the self-supervised loss function value with respect to the weight matrix; The self-learning task score is determined based on the gradient and the weight matrix; when the absolute value of the gradient and the weight matrix are larger, the self-learning task score is larger, indicating that the corresponding module is more important in the self-supervised learning task.
3. The pruning method of the model according to claim 1, characterized in that Determining the multi-task comprehensive scores of each module in the multiple downstream tasks based on the downstream task training results includes: Obtaining a weight matrix for each module and a task weight for representing the importance of each downstream task; Determining downstream loss function values corresponding to different downstream tasks according to the downstream task training results; Determining the gradient of the downstream loss function value with respect to the weight matrix; Determining, according to the gradient and the weight matrix, a single-task score of each module under different downstream tasks; when the absolute value of the gradient and the weight matrix are larger, the single-task score is larger, indicating that the corresponding module is more important in the corresponding downstream task; The task weights and the single-task scores are weighted and summed to obtain the multi-task comprehensive score.
4. The pruning method of the model according to claim 1, characterized in that Determining the pruning ratio of each module according to the importance score includes: sorting the importance scores in ascending order; Prune the modules corresponding to the first proportion at the front of the sorting result at a first pruning ratio; prune the modules corresponding to the second proportion in the middle of the sorting result at a second pruning ratio; and prune the modules corresponding to the third proportion at the end of the sorting result at a third pruning ratio; The sum of the first proportion, the second proportion, and the third proportion is equal to 1, and there is no intersection between them; the first pruning ratio is greater than the second pruning ratio; and the second pruning ratio is greater than the third pruning ratio.
5. The pruning method of the model according to claim 1, characterized in that: Determining the pruning ratio of each module according to the importance score includes: Get the preset basic pruning ratio and target adjustment range; Sorting the importance scores in ascending order to obtain a ranking value of each module in the sorting result; The pruning ratio is determined according to the basic pruning ratio, the target adjustment range, and the ranking value; the smaller the ranking value, the smaller the importance score, and the larger the pruning ratio.
6. The pruning method of the model according to claim 5, characterized in that: The iterative pruning of each module based on the pruning ratio includes: Determine a base loss function value of the target large model after the first training; and determine a current loss function value of the target large model; Determining a loss function increase amount according to the basic loss function value and the current loss function value; Determine whether the increase in the loss function is less than a threshold; If less than, return to the step of iteratively training the pre-built target large model using the text data to obtain a training result; If it is not less than, and the preset iteration condition is not met, the pruning ratio of the previous pruning is maintained, and the target adjustment amplitude is adjusted.
7. The model pruning method according to claim 6, characterized in that: The adjusting the target adjustment range includes: Obtain a preset adjustment coefficient and the historical adjustment range corresponding to the last pruning; the adjustment coefficient is greater than 0 and less than 1; Determining a current amplitude adjustment amount according to the adjustment coefficient and the loss function increase; Determining the target adjustment range according to the historical adjustment range and the current adjustment range; the greater the increase in the loss function, the greater the current adjustment range, and the smaller the target adjustment range; Correspondingly, the preset iteration condition is that the target adjustment amplitude is smaller than a preset amplitude.
8. A model pruning device, characterized in that: The device comprises: A text data acquisition module, configured to acquire text data; wherein the text data includes data of multiple different tasks; an iterative training module, configured to iteratively train a pre-built target large model using the text data to obtain training results; the text data includes first text data for a self-supervised learning task and second text data for multiple downstream tasks; the training results include self-supervised training results corresponding to the first text data and downstream task training results corresponding to the second text data; An importance score determination module, configured to determine the importance scores of each module in the target large model in the plurality of different tasks based on the training results; a model pruning module, configured to determine a pruning ratio for each module according to the importance score; and iteratively prune each module based on the pruning ratio until a preset iteration condition is met; the smaller the importance score, the larger the pruning ratio, indicating that the corresponding module has lower importance in the multiple different tasks; A self-learning task score determination module, configured to determine the self-learning task score of each module in the self-supervised learning task according to the self-supervised training result; a multi-task comprehensive score determination module, configured to determine, based on the downstream task training results, the multi-task comprehensive scores of each module in the multiple downstream tasks; The first determination module is used to determine the importance score based on the self-learning task score and the multi-task comprehensive score; when the self-learning task score is larger, the importance score is larger; when the multi-task comprehensive score is larger, the importance score is larger.
9. A model pruning device, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps of the pruning method of the model described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the pruning method of the model described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Model pruning method and device, equipment, storage medium and program product
CN116523025A
Model pruning method, device and equipment and storage medium
CN118690815A