A model compression and inference acceleration method based on lightweight multi-export network
By adopting a lightweight multi-export network model in power grid applications, combined with the joint optimization method of static compression and dynamic early exit, the problem of large computing overhead in deep learning models is solved, and efficient model compression and inference acceleration is achieved, which is suitable for environments with poor hardware computing capabilities.
Patent Information
- Application Number
- CN202211194881.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-09-26
AI Technical Summary
Deep learning models have high computing overhead in power grid applications and slow deployment and inference speed. Especially in environments with poor hardware computing capabilities, it is difficult to meet the needs of fast inference and low resource consumption.
The model compression and inference acceleration method based on lightweight multi-exit network are adopted, and the joint optimization method of static model compression and dynamic early exit are combined to optimize time and space efficiency, reduce model volume, and dynamically adjust the calculation amount according to sample difficulty.
While retaining model performance, it greatly reduces storage and computing overhead, improves the difficulty of model deployment, and achieves dynamic acceleration on samples of different complexities, and improves inference speed.
Smart Images

Figure CN115600675B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a model compression and inference acceleration method based on a lightweight multi-export network, and belongs to the technical field of artificial intelligence. Background Art
[0002] With the development of deep learning technology, deep neural network models are becoming deeper and deeper, and the model complexity is increasing. Correspondingly, the computational overhead of deep learning models is also increasing. This brings certain challenges to the deployment and application of deep learning models. Especially in the field of intelligent processing of numerous data applications in power grids, on the one hand, power grid data from terminals to control terminals contains a large amount of text, numerical data and other data, requiring the deployed model to have reliable natural language and structured language processing capabilities; on the other hand, there is a large difference in hardware computing power from user terminals to central control terminals, requiring the deployed model to have faster inference speed and lower resource consumption.
[0003] Research on model compression and inference acceleration has been started early in the fields of computer vision such as image recognition and target detection, mainly for convolutional neural networks (CNN). Since 2019, with the rise of Transformer multi-layer transformers and large-scale pre-trained language models, some scholars have gradually conducted research on the compression and acceleration of pre-trained language models. For text data in power grid application data, good model compression and inference acceleration technology is needed.
[0004] Currently commonly used model compression and inference acceleration methods include pruning, matrix decomposition, quantization, knowledge distillation, and dynamic early exit.
[0005] The pruning compression method aims to remove some parameters or structures with lower importance in the original model. First, the importance of model parameters or structures is evaluated by weight values, Taylor expansion and other methods, and then the parameters or structures are sorted from high to low in importance, and the parameters or structures with higher importance are retained to achieve model compression. For pre-trained language models with multi-layer transformer (Transformer) structures such as Bidirectional Encoder Representations from Transformers (BERT), there are existing technologies for pruning the feed-forward network (FFN) neurons in the Transformer; there are also technical solutions that believe that the multi-head self-attention structure in the Transformer is redundant and achieve compression by reducing the number of self-attention heads. Both of the above studies are compression of model width. Fan, Grave, and Joulin proposed the LayerDrop pruning method to reduce the number of Transformer layers, that is, compress the model depth. The above pruning methods are all coarse-grained structured pruning. Fine-grained unstructured pruning can theoretically reduce the complexity and computational complexity of the model, but it is difficult to directly reduce the storage space and improve the computational efficiency of the pruned sparse matrix on general-purpose devices. In contrast, structured pruning is more feasible and can directly reduce the time and space overhead of actual calculations without the use of special computing equipment.
[0006] Like various neural networks, the main module of a pre-trained language model is usually a transformation matrix composed of weights. The dimensions of this type of matrix, that is, the input and output dimensions of the transformation, are often large. Therefore, the original parameter matrix can be approximately decomposed into several matrices with smaller dimensions by low-rank approximation methods such as singular value decomposition, thereby reducing the number of matrix parameters while being sufficient to restore accuracy. Based on this idea, the prior art performs singular value decomposition on each parameter matrix of the pre-trained BERT. When pre-training the ALBERT model, the word embedding matrix is decomposed into two small matrices that are multiplied, reducing the number of word embedding matrix parameters to 17%, and its output dimension does not need to be related to the hidden state dimension. In addition, the ALBERT model uses a parameter sharing method to make the parameters of each layer of the transformation matrix the same and only stored once, greatly reducing the number of model parameters. However, compared with BERT, the amount of computation and time required for ALBERT to achieve similar accuracy has increased rather than decreased.
[0007] Knowledge distillation aims to transfer the knowledge and behavior of a large "teacher network" model to a small "student network" model through training. There is an existing technology that uses the cosine loss of the hidden state to distill the BERT model with half the number of layers in the pre-training stage. There is also a technical solution that proposes the mean square error loss represented by the student and teacher models [CLS], so that the model learns the most important classification feature representation of BERT. Using data augmentation, the hidden state output of the embedding layer of BERT, the Transformer attention matrix, the hidden state output of each layer, and the model probability distribution output are distilled in the two stages of pre-training and fine-tuning. The trained 4-layer TinyBERT model is less than BERT. BASE The model achieves a 9.4-fold acceleration ratio with 1 / 7 of the model parameters, and the average accuracy is retained at 96.8%.
[0008] The core of dynamic early exit is to select the amount of calculation according to the difficulty of the sample. Simple samples only need to be calculated by a small number of model modules, while complex samples require more calculation. This idea first came from the technology in the field of vision. It introduces a network branch exit after a specific convolution layer of the ordinary convolutional neural network model, and makes confidence judgments on each branch. If the shallow layer can extract sample features and make high-confidence predictions, there is no need to go through the remaining modules. Some studies have applied this idea to pre-trained language models such as BERT, treating each layer of the Transformer in the model as a module and performing layer-level dynamic early exit. It should be noted that the dynamic early exit method can only accelerate inference, cannot reduce the model size, and will introduce additional branch exit parameters. Therefore, it makes sense to combine static model compression with dynamic early exit. However, it is difficult to find an optimal setting for the static compression technology currently used, which is both effective for simple input samples and accurate for complex samples. At the same time, early exit cannot reduce the redundancy of the model width and is powerless to reduce the actual size of the model. In addition, interpretability studies have shown that the attention and semantic features of each layer in BERT are different. Therefore, deriving a multi-outlet network model from a pre-trained single-outlet network model such as BERT will result in inconsistency in the training objectives. That is, each layer of BERT needs to simultaneously make output predictions and provide services to deeper layers, and there is a conflict between the two. According to experimental observations, the uncompressed BERT is less affected by this inconsistency, while the compressed small-capacity model has difficulty balancing the conflict between shallow and deep predictions. Implanting the outlet after compressing the model will lead to severe performance degradation, hindering the complementarity of the two optimizations. Summary of the invention
[0009] The purpose of the present invention is to provide a model compression and inference acceleration method based on a lightweight multi-exit network, combining the two optimization methods of static model compression and dynamic early exit, and using a width-compressed lightweight multi-exit network model to optimize time and space efficiency. At the same time, joint optimization is performed to reduce the performance degradation of the compressed lightweight multi-exit network model caused by inconsistencies in each layer.
[0010] The purpose of the present invention is achieved through the following technical solution: a model compression and inference acceleration method based on a lightweight multi-export network, comprising the following steps:
[0011] Step 1: Use the existing fine-tuning method to train the transformer-based pre-trained language model on the user-given dataset to obtain the teacher model, and use the teacher model to initialize the student model, which is a lightweight multi-export network model;
[0012] Step 2: Set the intermediate dimension of the word embedding matrix according to the number of teacher model parameters P′ and the expected number of parameters P of the lightweight multi-export network model Number of self-attention heads Feedforward network intermediate dimensions in is the floor operator, the intermediate dimension of the word embedding matrix of the teacher model, the number of self-attention heads A, and the intermediate dimension of the feedforward network are E′, A′, and F′ respectively;
[0013] Step 3: Use a joint optimization method that combines static compression and dynamic acceleration to train the target lightweight multi-export network model;
[0014] Step 4: Set or change the confidence threshold of the lightweight multi-exit network model according to actual needs before inference to achieve variable degree of acceleration.
[0015] The purpose of the present invention can also be further achieved by the following technical measures:
[0016] Furthermore, in step 2, the settings of the intermediate dimension E of the word embedding matrix, the number of self-attention heads A, and the intermediate dimension F of the feedforward network must satisfy the following constraints:
[0017] Constraint 1: The intermediate dimension E of the word embedding matrix must be a positive integer and smaller than the dimension E′ of the word embedding matrix of the fine-tuned pre-trained language model, that is, 0<E<E′, E∈N * , where N * is a set of positive integers;
[0018] Constraint 2: The number of self-attention heads A must be a positive integer and less than the number of self-attention heads A′ of the fine-tuned pre-trained language model, that is, 0<A<A′, A∈N * , where N * is a set of positive integers;
[0019] Constraint 3: The intermediate dimension F of the feedforward network must be a positive integer and less than the number of self-attention heads F′ of the fine-tuned pre-trained language model, that is, 0<F<F′, F∈N * , where N * is a set of positive integers; and the setting of the intermediate dimension F of the feedforward network satisfies
[0020] F=A×hs×θ
[0021] Among them, hs is the dimension of each self-attention head, which is set to 64; θ is the scaling factor, which is set to 2.
[0022] Further, in step 3, the joint optimization method of integrated static compression and dynamic acceleration includes word embedding matrix approximation based on singular value decomposition, transformation matrix compression based on iterative pruning, lightweight multi-exit network recovery training based on knowledge distillation, and inference acceleration based on dynamic early exit; the steps of the joint optimization of integrated static compression and dynamic acceleration are as follows:
[0023] Step 3.1: Use the truncated singular value decomposition method to approximate the word embedding matrix of the model based on singular value decomposition. First, embed the word embedding matrix W t The matrix is split into three matrices U, the product of ∑ and V of dimension E; then, the truncated singular value decomposition truncates the three matrices to dimension E-ΔE, where And multiply U and ∑ matrix as the first word embedding matrix W t1 , use the V matrix as the second word embedding matrix W t2 ;
[0024] Step 3.2: Perform transformation matrix compression based on iterative pruning on the student model (i.e., lightweight multi-export network model). The transformation matrix compression is reflected in the transformation matrix, that is, the least important rows or columns of the matrix are discarded to reduce the number of parameters and floating-point operations. For the multi-head self-attention structure in the student model, the pruning granularity is the attention head, that is, each pruning will reduce the number of self-attention heads from the current value by ΔA, where ΔA=1. For the feedforward network structure in the model, the intermediate dimension of the feedforward network is reduced from the current value by ΔF.
[0025] Step 3.3: Perform knowledge distillation-based lightweight multi-exit network recovery training on the student model: After the forward propagation of the teacher model and the student model is completed, obtain the predictions of each layer of the teacher model and hidden state is the predicted probability distribution of the output of the teacher model from layer 1 to layer L, is the hidden state output by the word embedding layer of the teacher model, is the hidden state of the output of layers 1 to L of the teacher model; obtains the predictions of each layer of the student model and hidden state is the predicted probability distribution of the output of the student model from layer 1 to layer L, is the hidden state output by the word embedding layer of the student model, Calculate the distillation loss function for the hidden states of the output of layers 1 to L of the student model
[0026]
[0027] Where L is the number of model layers, CELoss(·) is the cross entropy loss function, MSELoss(·) is the mean square error loss function, and ∑ is the summation operation;
[0028] Step 3.4: Use the gradient balancing method to scale the gradient of each layer of the network structure at different ratios to avoid the problem of excessive gradient caused by superimposing the loss function in formula (1); specifically, the gradient of the loss function for the k-th layer of the network structure is expressed as:
[0029]
[0030] where w k represents the weight of the k-th layer network structure, is the gradient of the i-th layer loss to the k-th layer weight before gradient equalization, is the gradient of the model loss with respect to the weight of the kth layer after gradient equalization;
[0031] Step 3.5: If the compression of the student model has been completed, that is, the intermediate dimension of the word embedding matrix of the student model is reduced from E′ to E, the number of self-attention heads is reduced from A′ to A, and the intermediate dimension of the feedforward network is reduced from F′ to F, then continue to run step 3.3 until the training data iteration is completed; otherwise, continue to run steps 3.1 to 3.5.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] The prior art is generally a shallow and wide compression model structure that does not maintain the overall structure of the pre-trained model. The present invention designs a deep and narrow width compressed multi-export model to optimize time and space efficiency, retaining the ability of the pre-trained model to extract high-level semantics, and significantly reducing storage and computing overhead while retaining most of the model's performance. The prior art is unable to dynamically adjust the amount of calculation for samples of different difficulty levels, there is redundancy for simple samples, and insufficient calculation for complex samples. The present invention overcomes this problem through dynamic inference. Experiments have shown that the technical solution of the present invention still has optimal performance under highly compressed and accelerated conditions.
[0034] The present invention analyzes the redundant locations and importance of each structure in the pre-trained language model, and compresses the model using a compression method based on pruning and matrix decomposition. Under the premise of small loss of precision, the model size is statically reduced, the time and space overhead of storage and calculation is reduced, and the difficulty of model deployment is reduced. Based on the dynamic early exit technology, an intermediate outlet is added to the original single-outlet network model, and the multi-outlet network model is fine-tuned, so that the model can dynamically judge the inference calculation amount for input samples of different complexities without destroying the overall structure, thereby further improving the model inference speed. The proposed joint optimization method adds outlet calibration to the width compression of the model trunk, removes structures that contribute less to each outlet and redundancy in width, reduces the inconsistency of the multi-outlet network model, and makes up for the problem that the combination of static compression and dynamic acceleration leads to a significant reduction in model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 Schematic diagram of the model compression and inference acceleration method based on lightweight multi-export network; W t is the word embedding matrix of the original model, W t1 and W t2 To decompose the two smaller word embedding matrices after approximation, W Q , W K , W V , W O is the transformation matrix in multi-head self-attention, W FI , W FO is the transformation matrix in the feedforward network;
[0036] Figure 2 A schematic diagram of the truncated singular value decomposition acting on the BERT word embedding matrix;
[0037] Figure 3 Schematic diagram of the effect of structured pruning on the neural network structure and transformation matrix; N i O represents the i-th output neuron, represents the jth input neuron, W ij represents the weight of the connection between the i-th output neuron and the j-th input neuron; the dotted line part is the structure removed after pruning, and the solid line part is the structure retained after pruning;
[0038] Figure 4 Schematic diagram of early exit of the 12-layer BERT model under three different confidence thresholds. DETAILED DESCRIPTION
[0039] The present invention uses a model compression and inference acceleration method based on a lightweight multi-exit network to reduce the size of the pre-trained language model with little loss of accuracy, reduce the time and space overhead of storage and calculation, reduce the difficulty of model deployment, and enable the model to further improve the model inference speed for input samples of different complexities without destroying the overall structure. In particular, the joint optimization method of the present invention removes structures that contribute less to each outlet of the lightweight multi-exit network, reducing the inconsistency of each layer in the lightweight multi-exit network model. The joint optimization process can be broken down into four main parts: word embedding matrix approximation based on singular value decomposition, transformation matrix compression based on iterative pruning, lightweight multi-exit network recovery training based on knowledge distillation, and inference acceleration based on dynamic early exit.
[0040] Figure 1 The model schematic diagram of the model compression and inference acceleration method based on the lightweight multi-exit network of the present invention comprises the following steps:
[0041] Step 1: Use the existing fine-tuning method to fine-tune the transformer-based pre-trained language model on the user-given data set to obtain an uncompressed single-outlet teacher model, and use the teacher model to initialize the student model, which is a lightweight multi-outlet network model);
[0042] Step 2: Set the intermediate dimension of the word embedding matrix according to the number of parameters P′ of the teacher model and the expected number of parameters P of the lightweight multi-export network model Number of self-attention heads Feedforward network intermediate dimensions in is the floor rounding operator. In this process, 1) the intermediate dimension E of the word embedding matrix must be a positive integer and less than the dimension E′ of the fine-tuned pre-trained language model word embedding matrix, that is, 0<E<E′, E∈N * , where N * is a set of positive integers; 2) The number of self-attention heads A must be a positive integer and less than the number of self-attention heads A′ of the fine-tuned pre-trained language model, that is, 0<A<A′, A∈N * , where N * is a set of positive integers; 3) The intermediate dimension F of the feedforward network must be a positive integer and less than the number of self-attention heads F′ of the fine-tuned pre-trained language model, that is, 0<F<F′, F∈N * , where N * is a set of positive integers; the setting of the intermediate dimension F of the feedforward network satisfies F = A × hs × θ, where hs is the dimension of each self-attention head, generally set to 64; θ is the scaling factor, generally set to 2.
[0043] Step 3: Use a joint optimization method of comprehensive static compression and dynamic acceleration to train the target lightweight multi-export network model, as follows;
[0044] Step 3.1: Word Embedding Matrix Approximation Based on Singular Value Decomposition
[0045] The present invention uses the truncated singular value decomposition method to approximate the first module of the pre-trained language model BERT, the word embedding matrix. The truncated singular value decomposition firstly transforms the word embedding matrix W t The matrix is split into three smaller matrices: the product of U, ∑, and V. After that, the truncated singular value decomposition truncates the three matrices to the desired dimension E (which needs to be smaller than the hidden state dimension of the model to be compressed), and multiplies U and ∑ matrices as the first word embedding matrix W. t1 , use the V matrix as the second word embedding matrix W t2 ,like Figure 2 shown.
[0046] Therefore, using truncated singular value decomposition to approximate the word embedding matrix can reduce the number of parameters from Compress to ( in is the vocabulary size, and H is the hidden state size. When E is small, the word embedding matrix can be greatly compressed. The model to be compressed selected by the present invention is BERT, and the vocabulary size The hidden state size is H=768. If the intermediate transformation dimension is set to E=128, the word embedding matrix parameters can be compressed from 23.4 million to 4.0 million.
[0047] Step 3.2: Transformation matrix compression based on iterative pruning:
[0048] Pruning can be divided into coarse-grained structured pruning and fine-grained unstructured pruning. Structured pruning removes all network structures such as neurons and all their connected weights, while unstructured pruning generally only removes specific weights. Considering that it is difficult for general equipment to store and calculate the sparse matrix obtained by unstructured pruning, the present invention chooses structured pruning. Figure 3 As shown, the structured pruning used in the present invention prunes unimportant neurons in the fully connected network, which is reflected in the transformation matrix by discarding certain rows or columns, thereby reducing the number of parameters and floating-point operations. Specifically, the present invention uses the following model structure importance calculation method to assist pruning.
[0049] The basis for pruning is the importance of the structure. This application hopes to prune the unimportant parts of the model, that is, the pruned model is closer to the loss of the model before pruning. If the importance of the structure is formally defined, the pruning goal is
[0050]
[0051] in is the loss function, X is the input vector, W is the weight matrix to be pruned, W′ is the weight matrix after pruning, |·| is the absolute value operation, It means to find the minimum value of the expression f when W′ is a variable. In other words, pruning can be regarded as a specific weight h i Set to 0 and minimize the difference between the loss of the pruned model and the loss of the pre-pruned model
[0052]
[0053] To minimize the difference in loss The optimization equation can be approximated using the Taylor expansion method. The loss function In h i = 0 can be expressed as
[0054]
[0055] Its leaves is the loss function For weight h i The derivative of 1 is the first-order remainder of the Taylor expansion. Ignore R 1 , and substituting formula (5) into formula (4) we can get
[0056]
[0057] Therefore, the importance of the structure can be expressed as the product of the weight and the gradient, i.e. After that, in each layer of Transformer, the importance of the weights connected to the structure The sum is calculated and pruned in ascending order to remove unimportant structures. In order to avoid irreparable damage to the model caused by a single pruning, the present invention uses iterative pruning. Iterative pruning is performed while training, and only a fixed proportion of structures set manually are pruned each time until the pruned model reaches the target size. Generally, the proportion of each pruning can be set to 10%. This method can enable the model to gradually restore the structure damaged by pruning during training and adjust the importance of the structure in a timely manner.
[0058] Step 3.3: Resume training of lightweight multi-export network model based on knowledge distillation:
[0059] The present invention uses two knowledge distillation methods: prediction distillation and feature distillation.
[0060] Predictive distillation is similar in concept to label smoothing, which uses a trained teacher model to predict the input samples and outputs its probability distribution as the “soft label” learned by the student model, aiming to solve the problem of imperfect training data annotation and improve the generalization of the student model. At the same time, predictive distillation can be combined with data enhancement. Data enhancement uses methods such as synonym replacement, random deletion, and mask generation to generate more data for limited data, enrich the distribution of training data, and further improve the generalization ability of the model. Usually, the enhanced data does not have a real label. If the label of the original data is directly used as the corresponding enhanced data label, problems such as semantic mismatch may occur. In order to avoid this problem, the present invention uses the probability distribution output of the teacher model As the label of the augmented data, and make the probability distribution output of the student model As close as possible to the probability distribution output of the teacher model. Therefore, the loss function of predictive distillation can be expressed as
[0061]
[0062] Where L is the number of model layers, CELoss(·) is the cross entropy loss function, and ∑ is the summation operation.
[0063] Since the lightweight multi-export network of the present invention is relatively small, using only prediction distillation may not achieve the expected effect, because even if the prediction of the teacher model is known, the lightweight multi-export network cannot truly learn the language representation. To this end, the present invention uses feature distillation to transfer the intermediate representation of the teacher model to the student model. The present invention selects the embedding layer and the Transformer layer hidden state output of the teacher model As features, and make the embedding layer and Transformer layer hidden state output of the student model As close as possible to the hidden state output of the teacher model. Therefore, the loss function of feature distillation can be expressed as
[0064]
[0065] Where L is the number of model layers, MSELoss(·) is the mean square error loss function, and ∑ is the summation operation.
[0066] In summary, in the lightweight multi-export network recovery training based on knowledge distillation of the present invention, the loss function is expressed as
[0067]
[0068] Step 3.4: Inference acceleration based on dynamic early exit:
[0069] Different from static model compression, early exit aims to dynamically determine the amount of computation by the model itself during inference based on the complexity of the input samples. Specifically, the present invention uses layer-level early exit, such as Figure 4 As shown in the figure, a classifier is added after the Transformer layer, which is used as an early exit. During inference, after passing through each layer of Transformer, the input sample temporarily flows to the exit classifier to determine whether its output reaches the user-preset confidence threshold. If the confidence of this layer meets the exit condition, it is considered that the model can make a reliable judgment on the input sample at this time, and there is no need to continue to enter the subsequent Transformer layer, and the inference process can be exited early.
[0070] The present invention uses the entropy of the probability distribution of the classifier output as the confidence judgment criterion, which is defined as
[0071]
[0072] Where P(i) is the probability of sample classification i, C is the number of categories, and ln is the natural logarithm operation. If the entropy value S is greater than a given threshold S T , indicating that the model classification probabilities are similar, and it is difficult to make credible predictions in this state; on the contrary, it means that the model tends to make a certain prediction with a larger probability value, and the probability values of other predictions are small, so a high-confidence prediction can be made.
[0073] When training a lightweight multi-export network model, it is necessary to optimize the loss function of each layer of the export. Specifically, the loss of the i-th layer export can be written as
[0074]
[0075] As shown in formula (1), the present invention superimposes the loss function of each layer's exit. In order to avoid the problem of excessive gradient caused by the superposition of loss functions, the present invention uses a gradient balancing method, that is, scaling the gradient of each layer of the network structure at different proportions so that the gradient of the subsequent back propagation of each layer is averaged. Specifically, the gradient of the loss function for the kth layer of the network structure can be expressed as
[0076]
[0077] where w k represents the weight of the k-th layer network structure, is the gradient of the i-th layer loss to the k-th layer weight before gradient equalization, is the gradient of the model loss with respect to the weight of the kth layer after gradient equalization.
[0078] Step 3.5: If the model compression is completed, that is, the intermediate dimension of the word embedding matrix of the student model is reduced from E′ to E, the number of self-attention heads is reduced from A′ to A, and the intermediate dimension of the feedforward network is reduced from F′ to F, then continue to run step 3.3 until the training data iteration is completed. Otherwise, continue to run steps 3.1 to 3.5.
[0079] Step 4: Set or change the confidence threshold of the lightweight multi-exit network model according to actual needs before inference to achieve variable degree of acceleration. Figure 4 In the figure, the dark black path represents the inference process of the conventional BERT model, which can also be regarded as setting the confidence threshold to the maximum value (in the binary classification scenario, the maximum value is 0.7). The gray and light gray paths are the cases of medium (in the binary classification scenario, about 0.3) and low (in the binary classification scenario, about 0.1) confidence thresholds, respectively.
[0080] Specific implementation cases:
[0081] Taking the classification of power grid announcements as an example, we first constructed a data set of power grid announcements and divided the announcements into four categories: "grid construction", "grid procurement", "technical support" and "grid service". The number of announcements in the above four categories is shown in Table 1.
[0082] Table 1 Details of the power grid announcement dataset
[0083]
[0084] First, use this dataset to train an uncompressed multi-exit BERT teacher model, and use the teacher model to initialize the student model. At this point, the student model occupies approximately 420 megabytes of storage space and the inference time is approximately 50 milliseconds. According to the actual computing requirements, the expected model requires approximately 50-70 megabytes of storage space and 10-30 milliseconds of inference time. Based on this, the expected word embedding matrix intermediate dimension E = 128, the number of self-attention heads A = 2, and the feedforward network intermediate dimension F = 256 of the student model are set. Thereafter, according to step 3, the student model is trained using a joint optimization method of comprehensive static compression and dynamic acceleration, resulting in a student model of approximately 63 megabytes, and a complete inference time of approximately 22 milliseconds, which meets the expected computing requirements. Afterwards, it is observed that the model classification accuracy is high, allowing the model to exit earlier, so the classification confidence threshold is set to a medium-low level of 0.3. Under this setting, the average inference time of the model is reduced to about 16 milliseconds. The test input announcement title "2021 Power Grid Safety Automatic Device and System Protection Control Strategy Adaptability Assessment Tender Announcement of the State Grid Central China Electric Power Control Center" and output its correct classification as "Technical Support", which took about 18 milliseconds; the input announcement title "2021 Fourth Service Tender Announcement of the Advanced Training Center of State Grid Corporation of China" and output its correct classification as "Power Grid Service", which took about 11 milliseconds.
[0085] The present invention evaluates the model of comprehensive static compression and dynamic acceleration on the General Language Understanding Evaluation (GLUE) dataset. One task is selected from each of the three types of tasks in the GLUE dataset, namely Stanford Sentiment Tree Classification (SST-2), Quora Question Pair Equivalence Classification (QQP), Question Natural Language Inference (QNLI), and another ternary classification task Multi-Genre Natural Language Inference (MNLI).
[0086] The evaluation indicators of this technology are accuracy and F1 value, where TP is the number of correct positive predictions, FP is the number of wrong negative predictions, TN is the number of correct negative predictions, and FN is the number of wrong positive predictions. The model method in the present invention is compared with the mainstream compression and acceleration methods, and the specific results are shown in Table 2. The experimental results show that the optimization method proposed in the present invention exceeds the performance of many mainstream models, proving the effectiveness of the method proposed in the present invention:
[0087] Table 2 Experimental results on the GLUE dataset. The bold content in the table is the best experimental result.
[0088]
[0089] In addition to the above embodiments, the present invention may also have other implementation modes. Any technical solutions formed by equivalent replacement or equivalent transformation shall fall within the protection scope required by the present invention.
Claims
1. A model compression and inference acceleration method based on a lightweight multi-export network, comprising the following steps: Step 1: Use the existing fine-tuning method to train the transformer-based pre-trained language model on the text dataset in the power grid application data given by the user to obtain the teacher model, and use the teacher model to initialize the student model, which is the lightweight multi-export network model; Step 2: Set the intermediate dimension of the word embedding matrix according to the number of teacher model parameters P′ and the expected number of parameters P of the lightweight multi-export network model Number of self-attention heads Feedforward network intermediate dimensions in is the floor operator, the intermediate dimension of the word embedding matrix of the teacher model, the number of self-attention heads, and the intermediate dimension of the feedforward network are E′, A′, and F′ respectively; Step 3: Use a joint optimization method that combines static compression and dynamic acceleration to train the target lightweight multi-export network model; Step 4: Set or change the confidence threshold of the lightweight multi-exit network model according to actual needs before inference to achieve variable degree of acceleration; It is characterized in that in step 2, the settings of the intermediate dimension E of the word embedding matrix, the number of self-attention heads A, and the intermediate dimension F of the feedforward network must satisfy the following constraints: Constraint 1: The intermediate dimension E of the word embedding matrix must be a positive integer and smaller than the dimension E′ of the word embedding matrix of the fine-tuned pre-trained language model, that is, 0 <E<E′,E∈N * , where N * is a set of positive integers; Constraint 2: The number of self-attention heads A must be a positive integer and less than the number of self-attention heads A′ of the fine-tuned pre-trained language model, i.e. 0 <A<A′,A∈N * , where N * is a set of positive integers; Constraint 3: The intermediate dimension F of the feedforward network must be a positive integer and smaller than the intermediate dimension F′ of the feedforward network of the fine-tuned pre-trained language model, that is, 0 <F<F′,F∈N * , where N * is a set of positive integers; and the setting of the intermediate dimension F of the feedforward network satisfies F = A × hs × θ Among them, hs is the dimension of each self-attention head, which is set to 64; θ is the scaling factor, which is set to 2.
2. A model compression and inference acceleration method based on a lightweight multi-exit network as claimed in claim 1, characterized in that: In step 3, the joint optimization method of comprehensive static compression and dynamic acceleration includes word embedding matrix approximation based on singular value decomposition, transformation matrix compression based on iterative pruning, lightweight multi-exit network recovery training based on knowledge distillation, and inference acceleration based on dynamic early exit; the steps of the joint optimization of comprehensive static compression and dynamic acceleration are as follows: Step 3.1: Use the truncated singular value decomposition method to approximate the word embedding matrix of the model based on singular value decomposition. First, embed the word embedding matrix W t The matrix is split into three matrices U, the product of ∑ and V of dimension E; then, the truncated singular value decomposition truncates the three matrices to dimension E-ΔE, where And multiply U and ∑ matrix as the first word embedding matrix W t1 , use the V matrix as the second word embedding matrix W t2 ; Step 3.2: Perform transformation matrix compression based on iterative pruning on the student model. The transformation matrix compression is reflected in the transformation matrix, that is, the least important rows or columns of the matrix are discarded to reduce the number of parameters and floating-point operations. For the multi-head self-attention structure in the model, the granularity of pruning is the attention head, that is, each pruning will reduce the number of self-attention heads from the current value by ΔA, where ΔA=1; for the feedforward network structure in the model, the intermediate dimension of the feedforward network is reduced from the current value by ΔF, Step 3.3: Perform knowledge distillation-based lightweight multi-exit network recovery training on the student model: After the forward propagation of the teacher model and the student model is completed, obtain the predictions of each layer of the teacher model and hidden state is the predicted probability distribution of the output of the teacher model from layer 1 to layer L, is the hidden state output by the word embedding layer of the teacher model, is the hidden state of the output of layers 1 to L of the teacher model; obtains the predictions of each layer of the student model and hidden state is the predicted probability distribution of the output of the student model from layer 1 to layer L, is the hidden state output by the word embedding layer of the student model, Calculate the distillation loss function for the hidden states of the output of layers 1 to L of the student model Where L is the number of model layers, CELoss(·) is the cross entropy loss function, MSELoss(·) is the mean square error loss function, and ∑ is the summation operation; Step 3.4: Use the gradient balancing method to scale the gradient of each layer of the network structure at different ratios to avoid the problem of excessive gradient caused by superimposing the loss function in formula (1); specifically, the gradient of the loss function for the k-th layer of the network structure is expressed as: where w k represents the weight of the k-th layer network structure, is the gradient of the i-th layer loss to the k-th layer weight before gradient equalization, is the gradient of the model loss with respect to the weight of the kth layer after gradient equalization; Step 3.5: If the compression of the student model has been completed, that is, the intermediate dimension of the word embedding matrix of the student model is reduced from e′ to e, the number of self-attention heads is reduced from A′ to A, and the intermediate dimension of the feedforward network is reduced from F′ to F, then continue to run step 3.3 until the training data iteration is completed; otherwise, continue to run steps 3.1 to 3.5.
Citation Information
Patent Citations
Automatic compression method and platform based on multi-level knowledge distillation pre-training language model
CN112241455A