Large-scale pre-training model rapid compression method and system

By splitting the large-scale pre-trained model into multiple independent modules and serially compressing, the problem of high resource consumption during model compression is solved, and efficient compression and excellent model performance under limited resource conditions are achieved.

CN120218140APending Publication Date: 2025-06-27SUZHOU INST FOR ADVANCED STUDY USTC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510327982.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Large-scale pre-trained models consume high computing power, storage and time during the compression process, and are difficult to effectively compress in scenarios where computing power or video memory is limited.

Method used

By splitting the large-scale pre-trained model into multiple series modules according to the cascade structure of the skeleton, and splitting each module into multiple basic units, and compressing independently and serially, each module is independently compressed in the subsequent process, reducing the resource requirements in the overall compression process.

Benefits of technology

It significantly reduces the computing power, storage and time overhead of large-scale pre-trained models during compression, so that they can still be compressed efficiently under limited resource conditions, and improves the model performance under high sparseness and low precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218140A_ABST
    Figure CN120218140A_ABST
Patent Text Reader

Abstract

The invention discloses a large-scale pre-training model rapid compression method and system. The method comprises the steps of obtaining an initial weight, a preset sparseness, preset precision, calibration data and standard data of a trained large-scale pre-training model; splitting the pruned large-scale pre-training model into a plurality of modules connected in series according to a cascade structure of a skeleton, and splitting each module into a plurality of basic units; respectively compressing each basic unit in a first module in a skeleton of a large-scale pre-training model; updating the weight in a second module in the large-scale pre-training model skeleton; and executing the same compression operation in the second module in the skeleton, and executing the same weight updating operation in the third module in the skeleton until the last module in the skeleton is completely executed, thereby obtaining a compressed model. According to the method, the problems of too high computing power consumption, too high storage and too long time in the compression process of a large-scale pre-training model are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of model compression in artificial intelligence, and particularly relates to a method and system for quickly compressing large-scale pre-trained models. Background Art

[0002] Since the proposal of the Transformer architecture and Vision Transformer, significant progress has been made in both vision and language models, making full use of the basic structure of the Transformer. In language modeling, model series such as OPT and Llama utilize the Transformer architecture to build large language models and achieve excellent performance. Similarly, important vision models such as the Segment anything model (SAM) also adopt Vision Transformer to build powerful image representations and perform well in segmentation, classification, and various downstream applications. The scales of these models are extremely large. For example, the largest OPT model contains 175 billion parameters, and the latest Llama v3 model has as many as 405 billion parameters. The largest SAM model has 636 million parameters. Although the number of parameters is significantly smaller, its deployment in certain scenarios is still challenging. To address the deployment problem of large-scale models, model compression techniques are usually adopted to reduce their scales to meet hardware limitations.

[0003] The large scale of large-scale pre-trained models results in huge costs for training or fine-tuning such models, making it difficult to deploy them in some scenarios with limited computing power and devices. Currently, the problem of difficult deployment of large-scale pre-trained models in such scenarios can be alleviated through model compression methods.

[0004] In model compression techniques, it mainly includes pruning, quantization, knowledge distillation, and low-rank decomposition: Pruning: reducing the model scale by removing unimportant weights or neurons; Quantization: converting the model weights from commonly used high-precision data types to low-precision data types to reduce the storage space of the model; Knowledge Distillation: re-creating a small-scale model and using the large-scale model to guide the training process of the small-scale model to transfer knowledge from the large-scale model to the small-scale model, enabling the small-scale model to obtain a performance close to that of the large-scale model to replace the large-scale model; Low-rank Decomposition: decomposing the weight matrix in the large-scale model into the product of multiple small matrices to reduce the number of model parameters.

[0005] The above method can effectively alleviate the problem of difficult deployment of large-scale pre-trained models, and pruning is a commonly used compression algorithm. Traditional pruning techniques can be divided into two categories: iterative pruning and one-shot pruning. Iterative methods are based on the "Lottery Ticket Hypothesis" and gradually find a sparse sub-network with performance similar to the original model through repeated pruning and fine-tuning. Although it can improve the pruning effect, it requires a large amount of computing resources for multiple rounds of fine-tuning. Especially when the model is large, the memory requirements and pruning time increase significantly. In addition, some of the latest pruning methods further reduce the memory consumption during fine-tuning by combining Low-Rank Adaptation technology, but still need to optimize the global model to a certain extent. However, the pruning algorithm itself also has high requirements for computing power and video memory. Therefore, there are certain difficulties in compressing large-scale pre-trained models in scenarios with limited computing power or video memory. In contrast, one-shot methods do not require global fine-tuning, significantly reducing the demand for memory and computing resources during compression. However, the performance of the one-shot pruning is not good in the case of high sparsity (50%-90%). In addition, most of the existing one-shot pruning methods compensate for errors for a single layer and do not fully consider the error propagation between layers during the pruning process, which may lead to error accumulation at high sparsity, thus affecting the performance of the overall model.

[0006] Pruning is an important model compression technique, usually divided into multiple pruning and one-shot compression methods according to the pruning frequency. Traditional iterative pruning methods are based on the "Lottery Ticket Hypothesis" and gradually identify a sparse sub-network through the process of gradually pruning and fine-tuning, making its performance equivalent to the original model, thus reducing the number of model parameters. Classic methods such as magnitude pruning and recent methods such as the LLM-Pruner method, which use block pruning combined with LoRA fine-tuning, all follow this hypothesis. However, for large models, the resource consumption of global fine-tuning is very large, requiring a large amount of GPU memory. Although the LLM-Pruner reduces the memory requirements by using LoRA, it still needs to store intermediate results and gradients. And methods such as LLM Surgeon, although avoiding global fine-tuning, still require a large amount of memory and time due to the use of second-order derivatives.

[0007] In contrast, one-shot pruning methods have received more attention by minimizing memory usage during fine-tuning. SparseGPT avoids global fine-tuning and accelerates the pruning process by compensating for accuracy within each basic layer. Wanda further simplifies this process, completely eliminating the compensation step and achieving performance equivalent to SparseGPT using a unique saliency scoring without fine-tuning. Summary of the Invention

[0008] In view of the deficiencies of the prior art, the present invention proposes a method and system for quickly compressing a large-scale pre-trained model to solve the problem of excessive consumption of computing power, storage, and time during the compression of a large-scale pre-trained model.

[0009] To achieve the foregoing invention object, the present invention adopts the following solutions: One aspect of the present invention provides a method for quickly compressing a large-scale pre-trained model, including the following steps: S1. Obtain the initial weights, preset sparsity, preset precision, calibration data, and standard data of the trained large-scale pre-trained model; S2. Split the pruned large-scale pre-trained model into multiple cascaded modules according to the cascaded structure of the backbone, and split each module into multiple basic units, where each module is independently and serially compressed in the subsequent process; S3. In the first module in the backbone of the large-scale pre-trained model, compress each of the basic units to obtain a compressed single module; S4. Update the weights in the second module in the backbone of the large-scale pre-trained model; S5. Perform the same operation as in step S3 in the second module in the backbone, and perform the same operation as in step S4 in the third module in the backbone until step S3 is completed in the last module in the backbone to obtain a compressed model.

[0010] Another aspect of the present invention provides a system for quickly compressing a large-scale pre-trained model, including: A data acquisition module for obtaining the initial weights, preset sparsity, preset precision, calibration data, and standard data of the trained large-scale pre-trained model; A module splitting module for splitting the pruned large-scale pre-trained model into multiple cascaded modules according to the cascaded structure of the backbone, and splitting each module into multiple basic units, where each module is independently and serially compressed in the subsequent process; A compression module for compressing each of the basic units in the first module in the backbone of the large-scale pre-trained model to obtain a compressed single module; A weight update module for updating the weights in the second module in the backbone of the large-scale pre-trained model; An iteration module for performing the same operation as the compression module in the second module in the backbone, and performing the same operation as the weight update module in the third module in the backbone until the last module in the backbone is completed to obtain a compressed model.

[0011] Compared with the prior art, the present invention has at least the following advantages: (1) Compared with the iterative compression method in the background art, the compression method proposed by the present invention significantly reduces the computing power, storage, and time overhead of large-scale pre-trained models during the entire compression process, making it possible to compress large-scale pre-trained models on 24G video memory consumer-grade graphics cards such as NVIDIA RTX 4090; (2) Compared with the one-time fast compression method (such as SparseGPT and Wanda in the background art), without increasing the video memory consumption, the performance of the model at high sparsity and low precision is significantly improved; (3) The present invention tested the language model on the Wikitext dataset. Taking the 70% sparsity of OPT-6.7B as an example, the perplexity (the smaller the value, the better) of the model compressed by the present invention is 17.82, which is significantly lower than 20.48 of SparseGPT and 159.2 of Wanda; taking the 2:4 sparsity and 3-Bit quantization of OPT-6.7B as an example, the perplexity of the model compressed by the present invention is 18.38, which is significantly lower than 20.07 of SparseGPT. The present invention also tested the vision model on the COCO dataset. Taking the 90% sparsity of SAM-H as an example, the mean intersection over union (the higher the value, the better) of the model compressed by the present invention is 57.60, which is significantly higher than 18.15 of SparseGPT and 1.20 of Wanda; taking the 2:4 sparsity and 3-Bit quantization of SAM-H as an example, the mean intersection over union of the model compressed by the present invention is 68.75, which is significantly higher than 66.39 of SparseGPT. In terms of video memory consumption, taking the compression of the OPT-6.7B model as an example, the highest video memory occupied by the present invention during the compression process is 7.8G, which is less than 8.9G of SparseGPT and only 0.1G higher than 7.7G of Wanda in terms of video memory occupancy. In terms of time overhead, taking the compression of the OPT-6.7B model as an example, the time consumed by the present invention during the compression process is 1 hour, which is twice that of 0.5 hours of SparseGPT and within an acceptable range. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0013] Figure 1 It is a schematic flowchart of a method for quickly compressing a large-scale pre-trained model provided by a typical embodiment of the present invention; Figure 2Schematic diagram of the innovation point of the compression method provided by a typical embodiment of the present invention; Figure 3 Schematic diagram of the implementation details of the compression method provided by a typical embodiment of the present invention. Detailed implementation manners

[0014] To make the objectives, technical solutions and advantages of the present invention clearer, the following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings. Examples of these preferred implementation manners are illustrated in the accompanying drawings. The implementation manners of the present invention shown in the drawings and described according to the drawings are merely exemplary, and the present invention is not limited to these implementation manners.

[0015] To alleviate the difficulty of pruning large-scale pre-trained models in scenarios with limited computing power and storage resources, the objective of the present invention is to provide a method and system for quickly compressing large-scale pre-trained models to achieve model pruning in scenarios with limited computing power and video memory conditions.

[0016] One aspect of the present invention provides a method for quickly compressing a large-scale pre-trained model, including the following steps: S1. Obtain the initial weights, preset sparsity, preset accuracy, calibration data, and standard data of the trained large-scale pre-trained model; S2. Split the pruned large-scale pre-trained model into multiple cascaded modules according to the cascade structure of the backbone, and split each module into multiple basic units, where each module is independently and serially compressed in the subsequent process; S3. In the first module in the backbone of the large-scale pre-trained model, compress each of the basic units to obtain a compressed single module; S4. Update the weights in the second module in the backbone of the large-scale pre-trained model; S5. Perform the same operation as in step S3 in the second module in the backbone, and perform the same operation as in step S4 in the third module in the backbone until step S3 is completed for the last module in the backbone to obtain a compressed model.

[0017] In one embodiment, step S1 specifically includes: S11. Obtain a data set according to different application scenarios of the large-scale pre-trained model, obtain the original full-precision weights of the relevant model as the initial weights before compression, select the preset sparsity and preset accuracy, and rearrange the sparsity; S12. Randomly sample a first preset number of samples in the data set as calibration data, and the calibration data is updated along with the forward propagation; Before preprocessing the calibration data and forward-inferring it to the backbone structure in the large-scale pre-trained model, a calibration data consisting entirely of intermediate activations is formed. S14. Copy a calibration data set as the standard data, and the standard data is used for weight error compensation.

[0018] In one embodiment, the step S2 specifically includes: S21. Divide the backbone structure in the large-scale pre-trained model into multiple modules with similar or identical structures but different weight values in sequence according to its tandem structure. S22. Split a single module into multiple compression units, and the compression units include linear layers, convolutional layers and their variants.

[0019] In one embodiment, the step S3 operates on the t-th module, starting from t = 0, and further includes: S31. Forward-infer the standard data to the t-th module. S32. In the t-th module, respectively construct compression constraint conditions for each of the compression units according to the corresponding weights and calibration data, where the constraint conditions are used to solve the importance scores of each weight in each of the compression units; the compression method here uses the SparseGPT method described in the background art to compress each of the compression units; the weights corresponding to the compression units are all the weight parameters included in the compression units. For example, if a compression unit contains a linear layer, then the corresponding weights refer to the parameter matrix in the linear layer.

[0020] S33. Sort the weights in each compression unit from large to small according to the importance scores, and select the least important weights to form a pruning 0-1 mask, where 0 corresponds to the weights not to be pruned, and 1 corresponds to the weights to be pruned. S34. According to the compression constraint conditions and the pruning 0-1 mask, solve the weight change amount of each compression unit after weight compression relative to before compression. S35. Add the weight change amount to the weight matrix of each compression unit to obtain the compressed t-th module. S36. Forward-infer the calibration data set to the compressed t-th module.

[0021] In one embodiment, the step S4 specifically includes: Step S41. Forward-infer the standard data to the (t + 1)-th module. Step S42: Use the calibration data as training data, the standard data as labels, and the mean squared error as the loss function to train the (t + 1)-th module until convergence, obtaining the (t + 1)-th module with updated weights. The expression for the mean squared error is: ; where is the calibration data of the i-th layer, is the standard data; Step S43: Forward-infer the calibration data set to the (t + 1)-th module.

[0022] In the above steps S31, S36, S41, and S43, for the forward inference of all standard data and calibration data, if a large number of intermediate results are saved in the video memory, it will cause a large video memory occupation. Therefore, the present invention uses block intermediate storage, that is, all intermediate results are stored in the system memory. When using data for forward inference, the data is read into the video memory one by one. After the previous data is processed by the GPU, the next data will be read into the video memory; after a single data is processed by the GPU, it is returned to its original position in the system memory, thus reducing the video memory overhead of the entire forward inference.

[0023] Another aspect of the present invention provides a large-scale pre-trained model fast compression system, including: A data acquisition module, configured to acquire the initial weights, preset sparsity, preset precision, calibration data, and standard data of the trained large-scale pre-trained model; A module splitting module, configured to split the pruned large-scale pre-trained model into multiple cascaded modules according to the cascade structure of the backbone, and split each module into multiple basic units, where each module is independently and serially compressed in the subsequent process; A compression module, configured to compress each of the basic units in the first module in the backbone of the large-scale pre-trained model to obtain a compressed single module; A weight update module, configured to update the weights in the second module in the backbone of the large-scale pre-trained model; An iteration module, configured to perform the same operations as the compression module in the second module in the backbone, and perform the same operations as the weight update module in the third module in the backbone until the last module in the backbone is executed, obtaining a compressed model.

[0024] In one embodiment, the data acquisition module is specifically configured to: Obtain a data set according to different application scenarios of the large-scale pre-trained model, obtain the original full-precision weights of the relevant model as the initial weights before compression, select the preset sparsity and preset precision, and rearrange the sparsity; Randomly sample the first preset number of samples in the dataset as calibration data, and the calibration data is updated along with the forward propagation; Before preprocessing the calibration data and performing forward inference to the backbone structure in the large-scale pre-trained model, form a calibration data composed entirely of intermediate activations; Duplicate a copy of the calibration dataset as the standard data, and the standard data is used for weight error compensation.

[0025] In one embodiment, the module splitting module is specifically used for: Divide the backbone structure in the large-scale pre-trained model into multiple modules with similar or identical structures but different weight values in sequence according to its tandem structure; Split a single module into multiple compression units, and the compression units include linear layers, convolutional layers and their variants.

[0026] In one embodiment, the compression module operates on the t-th module, starting from t = 0, and further includes: Perform forward inference on the standard data to the t-th module; Within the t-th module, respectively construct compression constraint conditions for each of the compression units according to the corresponding weights and calibration data, wherein the constraint conditions are used to solve the importance scores of each weight in each of the compression units; According to the importance scores, sort the weights in each compression unit from large to small, and select the least important weights to form a pruning 0-1 mask, where 0 corresponds to the weights not to be pruned, and 1 corresponds to the weights to be pruned; According to the compression constraint conditions and the pruning 0-1 mask, solve the weight change amount of each compression unit after weight compression relative to the weight before compression; Add the weight change amount to the weight matrix of each compression unit to obtain the compressed t-th module; Perform forward inference on the calibration dataset to the compressed t-th module.

[0027] In one embodiment, the weight update module is specifically used for: Perform forward inference on the standard data to the (t + 1)-th module; Use the calibration data as the training data, the standard data as the label, and the mean square error as the loss function, and train the (t + 1)-th module until convergence to obtain the (t + 1)-th module with updated weights; the expression of the mean square error is: ; Wherein, is the calibration data of the i-th layer, is the standard data; After forward-inferring the calibration dataset to the (t + 1)-th module.

[0028] See Figure 1 , Figure 1 which is a schematic flowchart of a method for quickly compressing a large-scale pre-trained model provided by an embodiment of the present invention, and includes the following steps: Step S1: Obtain the weights of the trained large-scale pre-trained model, a preset sparsity s, a preset accuracy, a small amount of calibration data, and standard data. Among them, for the preset sparsity and the preset accuracy, it is specified that the sparsity of all layers is s.

[0029] Specifically, step S1 further includes: Step S11: Obtain the FP16 weight information of the pre-trained model and the dataset in the corresponding scenario, select the compression degree (sparsity and preset accuracy) according to the deployment environment, and then re-arrange the sparsity according to the preset sparsity. The re-arrangement is divided into two parts: inter-layer sparsity re-arrangement and intra-layer sparsity re-arrangement.

[0030] Among them, the inter-layer sparsity re-arrangement is as shown in (a) in Figure 3 , which is to intercept a part of the preset sparsity of the last module and evenly distribute it to all other modules, and the distribution follows the following formula: ; where P is the number of weights to be pruned, is the module index, is the j-th module, is the number of all weights of the j-th module, s is the preset sparsity, and α is a preset hyperparameter. By adjusting α, the proportion of the pruning quantity of the previous module divided by the last module can be adjusted, and α is adjusted manually according to the effect of the pruned model.

[0031] Among them, the intra-layer sparsity re-arrangement is as shown in (b) in Figure 3 , which is to intercept a part of the preset sparsity of the QKV matrix in a Transformer and distribute it to the last fully connected layer, and the distribution follows the following formula: ; where represents any one of the three matrices of the j-th module, and β is a hyperparameter. By adjusting β, the degree to which a part of the preset sparsity of the QKV matrix in the Transformer is intercepted and distributed to the last fully connected layer can be adjusted.

[0032] Step S12: Randomly sample 128 samples from the dataset corresponding to the deployment scenario as the calibration dataset; Step S13: Input the calibration data into the model inference respectively before the model backbone, and update the standard data with this result , as shown in Figure 3 (a) therein; Step S14: Make an exact copy of the calibration data set and name it the standard data , as shown in Figure 3 (a) therein; Step S2: Split the pruned large-scale pre-trained model into a series of cascaded modules according to the backbone's cascading structure, and split a single module into multiple basic units. Among them, each module is independently and serially compressed in the subsequent process; Specifically, Step S2 further includes: Step S21: The large-scale pre-trained model usually contains a backbone with a cascading structure, which is composed of multiple modules with very similar structures. Divide the backbone into a series of modules with similar or identical structures but different weight values in order according to its cascading structure, as modules 1 to n, as shown in Figure 2 and Figure 3 (a) therein; Step S22: Split a single module into one-by-one compression units, and one matrix is one compression unit. Figure 3 Taking the classic Transformer as an example in (b) therein, it is divided into , , , , , six basic compression units; Step S3: Compress each of the basic units respectively within the first module in the backbone of the large-scale pre-trained model to obtain a compressed single module; Specifically, taking module 1 in Figure 3 (a) as an example, Step S3 further includes: Step S31: Forward-inference the standard data through the position before module 1 to after module 1 and before module 2, and update the standard data to ; Step S32: Within module 1, construct compression constraint conditions for each of the compression units respectively according to the weights and the calibration data. Among them, this constraint condition is used to solve the importance score of each weight in each of the compression units. Here, the SparseGPT method described in the background technology is used to compress each of the compression units. By constructing the constraint condition, the constrained optimization problem is: ; Among them, L represents the defined loss function; W represents the weight matrix; E represents a selection matrix, each column of which contains only one 1 and the rest are 0; represents the weight matrix before pruning, represents the change in weight matrix, Represents a quantization function.

[0033] in is defined by the following function, ; Among them, h is the scale factor and z is the displacement factor. Its purpose is to limit the dynamic range of x to In the example, N is the number of quantization bits. If a 4-bit quantization method is used, N is 4. Is the rounding function. The function is defined by the following formula: ; Its role is to ensure that the dynamic range of x is strictly limited to Can be processed by direct estimator and The non-differentiable property of .

[0034] Here, you can choose to apply only pruning, only quantization, or both pruning and quantization to the model according to the actual application of model compression. In this case, the form of the constraint does not change much.

[0035] Usually, the quantization function is difficult to differentiate, so we use step-by-step optimization to combine the first two tasks to find an analytical solution: ; in, is the kth weight of the row, For Perform the inverse of the sparse matrix of the second-order terms after Taylor expansion, represents the element in the kth row and kth column of the matrix, represents the k-th row of the matrix.

[0036] In this way, the error caused by pruning each weight can be directly calculated, so we can directly use Do pruning based on the importance score of pruning, and then use it after completion Function quantization operation.

[0037] Step S33: sort the weights in each compression unit from large to small according to the importance score, and select the least important weight to form a pruning 0-1 mask, where 0 corresponds to the weight that is not pruned and 1 corresponds to the weight that is to be pruned; Step S34: Solve the weight change amount of each compression unit after weight compression relative to before compression according to the compression constraint condition and pruning 0-1 mask, that is, the result obtained in Step S32 ; Step S35: Add the weight change amount to the weight matrix of each compression unit to obtain the compressed Module 1; Step S36: After forward-inferring the calibration data to the compressed Module 1, obtain .

[0038] Step S4: Update the weights within the second module in the large-scale pre-trained model backbone; Specifically, the present invention uses the adjustment of the latter module to compensate for the error generated by the compression of the previous module. Therefore, taking Module 2 as an example, Step S4 further includes: Step S41: After forward-inferring the standard data to Module 2, obtain the standard data after passing through Module 2 ; Step S42: Use the calibration data as the training data, the standard data as the label, and the mean squared error as the loss function, and train Module 2 until convergence to obtain Module 2 with updated weights, where the formula of the loss function is: ; where represents the output result obtained after the calibration data is calculated through Module 2. This loss function measures the deviation of this output result from the standard output result, that is, . To save video memory overhead, the size of the mini-batch in training is usually 1, but an appropriate size can be selected according to the special hardware conditions in the scenario.

[0039] Step S5: Compress and update each module in a sliding window-style iterative process. Repeat the operation of Step S3 on Module 2 and repeat the operation of Step 4 on Module 3, then repeat the operation of Step S3 on Module 3 and repeat the operation of Step S4 on Module 4, and so on until all modules are completely compressed. As shown in Figure 3 (a).

[0040] In steps S31, S36, S41 and S43, forward inference involving all standard data and calibration data will result in a large amount of video memory occupation if a large number of intermediate results are saved in the video memory. Therefore, the present invention uses block intermediate storage, that is, all intermediate results are stored in the system memory. When using data for forward inference, the data is read into the video memory one by one. After the previous data is processed by the GPU, the next data will be read into the video memory. After a single data is processed by the GPU, it is returned to its original position in the system memory. This reduces the video memory overhead of the entire forward inference. Therefore, the three operations of reading the data to be processed into the GPU, the GPU processing the current data, and sending the data processed by the GPU back to the CPU can be performed at the same time. However, specific practices can be selected according to the scenario.

[0041] It should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions of each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A large-scale pre-trained model fast compression method, characterized in that: The following steps are involved: S1. Obtain the initial weights, preset sparsity, preset accuracy, calibration data and standard data of the large-scale pre-trained model after training; S2, splitting the pruned large-scale pre-trained model into multiple serially connected modules according to the cascade structure of the skeleton, and splitting each module into multiple basic units, wherein each module is independently and serially compressed in a subsequent process; S3. In the first module in the skeleton of the large-scale pre-trained model, compress each of the basic units respectively to obtain a compressed single module; S4, update the weights in the second module in the large-scale pre-trained model skeleton; S5. Perform the same operation as step S3 in the second module in the skeleton, and perform the same operation as step S4 in the third module in the skeleton, until step S3 is completed in the last module in the skeleton to obtain a compressed model.

2. The large-scale pre-training model fast compression method according to claim 1 is characterized in that: The step S1 specifically includes: S11. Acquire a data set according to different application scenarios of a large-scale pre-trained model, obtain the original full-precision weight of the relevant model as the initial weight before compression, select a preset sparsity and a preset precision, and rearrange the sparsity; S12, randomly sampling a first preset number of samples in the data set as calibration data, wherein the calibration data is updated along with forward propagation; S13, preprocessing the calibration data and forward reasoning to the skeleton structure in the large-scale pre-training model to form calibration data consisting entirely of intermediate activations; S14. Copy a calibration data set as standard data, where the standard data is used for weight error compensation.

3. The large-scale pre-training model fast compression method according to claim 1 is characterized in that: The step S2 specifically includes: S21, dividing the skeleton structure in the large-scale pre-trained model into a plurality of modules with similar or identical structures but different weight values ​​in sequence according to its serial structure; S22. Split a single module into multiple compression units, where the compression units include linear layers, convolutional layers, and their variants.

4. The large-scale pre-training model fast compression method according to claim 3 is characterized in that: The step S3 operates on the t-th module, starting from t=0, and further includes: S31, forward reasoning the standard data to the tth module; S32, in the tth module, constructing compression constraint conditions for each compression unit according to the corresponding weight and calibration data, wherein the constraint conditions are used to solve the importance score of each weight in each compression unit; S33, sorting the weights in each compression unit from large to small according to the importance score, selecting the least important weight to form a pruning 0-1 mask, where 0 corresponds to a weight that is not pruned and 1 corresponds to a weight that is to be pruned; S34, according to the compression constraint and the pruning 0-1 mask, solving the weight change amount after the weight compression relative to the weight before compression in each compression unit; S35, adding the weight change amount to the weight matrix of each compression unit to obtain the compressed t-th module; S36, forward reasoning the calibration data set to the compressed t-th module.

5. The large-scale pre-training model fast compression method according to claim 4 is characterized in that: The step S4 specifically includes: Step S41, forward reasoning the standard data to the t+1th module; Step S42, using the calibration data as training data, the standard data as labels, and the mean square error as the loss function, training the t+1th module until convergence, and obtaining the t+1th module after weight update; Step S43, forward inference of the calibration data set to the t+1th module.

6. A large-scale pre-trained model fast compression system, characterized in that: include: A data acquisition module is used to obtain the initial weights, preset sparsity, preset accuracy, calibration data and standard data of the large-scale pre-trained model after training; A module splitting module is used to split the pruned large-scale pre-trained model into multiple serially connected modules according to the cascade structure of the skeleton, and split each module into multiple basic units, wherein each module is compressed independently and serially in a subsequent process; A compression module, used to compress each of the basic units in the first module in the skeleton of the large-scale pre-trained model to obtain a compressed single module; The weight update module is used to update the weights in the second module in the large-scale pre-trained model skeleton; The iteration module is used to perform the same operation of the compression module in the second module in the skeleton, and to perform the same operation of the weight update module in the third module in the skeleton, until the last module in the skeleton is executed, thereby obtaining a compressed model.

7. The large-scale pre-trained model fast compression system according to claim 6, characterized in that: The data acquisition module is specifically used for: Obtain data sets according to different application scenarios of large-scale pre-trained models, obtain the original full-precision weights of relevant models as initial weights before compression, select preset sparsity and preset precision, and rearrange the sparsity; Randomly sampling a first preset number of samples in the data set as calibration data, wherein the calibration data is updated with forward propagation; Preprocessing the calibration data and forward reasoning it to a skeleton structure in a large-scale pre-trained model to form calibration data consisting entirely of intermediate activations; A copy of the calibration data set is used as standard data, which is used for weight error compensation.

8. The large-scale pre-trained model fast compression system according to claim 6, characterized in that: The module splitting module is specifically used for: The skeleton structure in the large-scale pre-trained model is divided into multiple modules with similar or identical structures but different weight values ​​in sequence according to its serial structure; A single module is split into multiple compression units, which include linear layers, convolutional layers and their variants.

9. The large-scale pre-trained model fast compression system according to claim 8, characterized in that: The compression module operates on the t-th module, starting from t=0, and further includes: Forward inference of standard data to the tth module; In the tth module, respectively construct compression constraints for each of the compression units according to the corresponding weights and calibration data, wherein the constraints are used to solve the importance score of each weight in each of the compression units; According to the importance score, the weights in each compression unit are sorted from large to small, and the least important weight is selected to form a pruning 0-1 mask, where 0 corresponds to the weight that is not pruned and 1 corresponds to the weight that is to be pruned; According to the compression constraint and the pruning 0-1 mask, solving the weight change amount after the weight compression relative to the weight before compression in each compression unit; Adding the weight change to the weight matrix of each compression unit to obtain the compressed t-th module; The calibration dataset is forward inferred to the compressed t-th module.

10. The large-scale pre-trained model fast compression system according to claim 9, characterized in that: The weight updating module is specifically used for: Forward inference of standard data to the t+1th module; Use the calibration data as training data, the standard data as labels, and the mean square error as the loss function to train the t+1th module until convergence, and obtain the t+1th module after weight update; Forward inference of the calibration dataset to after the t+1th module.