Neural network sparse method, system, device and medium
By injecting calibration data into the large language model and generating a global cropping template, the weight matrix of the Transformer layer is trimmed and reconstructed, and the problem of difficult to balance resource requirements and performance of large language models in the prior art is solved, and efficient model compression and excellent performance are achieved.
Patent Information
- Application Number
- CN202510133297.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-06
AI Technical Summary
The existing large language models are difficult to find a balance between resource requirements and performance maintenance after pruning, resulting in shortcomings in practical applications and promotion.
By injecting calibration data into the neural network, the input and output similarity of attention modules and multi-layer perception modules in the Transformer layer is obtained, and a global cropping template is generated based on the model target sparseness and the module preset quota ratio, and the weight matrix is cropped and reconstructed.
The large language model is compressed to a high sparse state without retraining, while maintaining excellent performance and performing better benchmarks at high sparseness.
Smart Images

Figure CN120068943A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a neural network sparsification method, system, device, and medium. Background Art
[0002] With the rapid development of deep learning technology, large language models (LLMs) have demonstrated remarkable achievements in fields such as natural language processing (NLP). Large language models such as LLaMA, Palm, GLM, BLOOM, and GPT have achieved excellent performance in many natural language processing tasks such as text generation, machine translation, question answering systems, text summarization, and sentiment analysis.
[0003] The weight matrix dimensions of currently popular large language models (hereinafter referred to as LLMs) are relatively large. Taking LLaMA 3.18b as an example, the dimension of the V weight matrix of an attention module is 4k x 4k. To reduce storage and computational costs for easy deployment, the large language model can be made lightweight by pruning (or sparsifying) it.
[0004] Existing pruning techniques for large language models still have limitations in the application of large language models. Either they cannot achieve a sufficient compression ratio, or it is difficult to guarantee the performance of the model after pruning. Especially in the case of complex architectures and massive parameters of large language models, it is difficult to find a satisfactory balance between resource requirements and performance maintenance.
[0005] Therefore, a new pruning method applicable to large language models is an urgent problem to be solved in the practical application and popularization of current large language models. Summary of the Invention
[0006] In view of the above-mentioned disadvantages of the prior art, the purpose of this application is to provide a neural network sparsification method, system, device, and medium, which is used to solve the technical problem of balancing resource requirements and performance maintenance in the pruning of large language models in the prior art.
[0007] To achieve the above purpose and other related purposes, this application provides a neural network sparsification method. The neural network includes several stacked Transformer layers, and each Transformer layer includes an attention module and a multi-layer perceptron module;
[0008] The method includes:
[0009] Inject calibration data into the neural network to obtain the input feature data and output feature data of the attention module and the multi-layer perceptron module in each Transformer layer;
[0010] Obtain the input-output similarity of the attention module and the multi-layer perception module in each Transformer layer according to the input feature data and output feature data of the attention module and the multi-layer perception module in each Transformer layer;
[0011] Obtain the module target sparsity of the attention module and the multi-layer perception module in each Transformer layer according to the model target sparsity, the module preset quota ratio, and the input-output similarity of the attention module and the multi-layer perception module in each Transformer layer;
[0012] Generate a global pruning template for each weight matrix of the attention module and the multi-layer perception module in each Transformer layer according to the module target sparsity of the attention module and the multi-layer perception module in each Transformer layer;
[0013] Perform weight pruning and reconstruction on the weight matrices of the attention module and the multi-layer perception module in each Transformer layer according to the global pruning template of each weight matrix of the attention module and the multi-layer perception module in each Transformer layer, so as to update each weight matrix in the attention module and the multi-layer perception module in each Transformer layer.
[0014] In an optional embodiment of the present application, obtaining the input-output similarity of the attention module and the multi-layer perception module in each Transformer layer according to the input feature data and output feature data of the attention module and the multi-layer perception module in each Transformer layer includes:
[0015] Obtain the final output feature data of the attention module and the multi-layer perception module in each Transformer layer according to the input feature data and output feature data of the attention module and the multi-layer perception module in each Transformer layer;
[0016] Obtain the input-output similarity of the attention module and the multi-layer perception module in each Transformer layer according to the input feature data and the final output feature data of the attention module and the multi-layer perception module in each Transformer layer.
[0017] In an alternative embodiment of the present application, obtaining the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer according to the input feature data and the final output feature data of the attention module and the multi-layer perceptron module in each Transformer layer includes:
[0018] According to the input feature data and the final output feature data of the attention module and the multi-layer perceptron module in each Transformer layer, use a preset similarity algorithm to obtain the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer;
[0019] Wherein, the preset similarity algorithm includes cosine similarity algorithm, Euclidean similarity algorithm, Manhattan similarity algorithm, Pearson correlation coefficient algorithm.
[0020] In an alternative embodiment of the present application, obtaining the module target sparsity of the attention module and the multi-layer perceptron module in each Transformer layer according to the model target sparsity, the module preset quota ratio, and the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer includes:
[0021] Calculate and obtain the sparsity quota of the attention module and the multi-layer perceptron module according to the model target sparsity and the module preset quota ratio;
[0022] According to the sparsity quota of the attention module and the multi-layer perceptron module, and the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer, obtain the module target sparsity of the attention module and the multi-layer perceptron module in each Transformer layer.
[0023] In an alternative embodiment of the present application, obtaining the module target sparsity of the attention module and the multi-layer perceptron module in each Transformer layer according to the sparsity quota of the attention module and the multi-layer perceptron module, and the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer includes:
[0024] According to a preset amplification parameter, the sparsity quota of the attention module and the multi-layer perceptron module, and the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer, obtain the target module sparsity of the attention module and the multi-layer perceptron module in each Transformer layer.
[0025] In an alternative embodiment of the present application, the target module sparsity T of the attention module and the multi-layer perception module in each of the Transformer layers is calculated using the following formula:
[0026]
[0027] In the formula, S represents the input-output similarity of the current attention module or multi-layer perception module; turbo represents a preset amplification parameter; Mean(S) is the average input-output similarity of the attention module or multi-layer perception module in all Transformer layers; and base is the sparsity quota of the attention module or multi-layer perception module.
[0028] In an alternative embodiment of the present application, turbo is between 1.0 and 1.5.
[0029] In an alternative embodiment of the present application, according to the module target sparsity of the attention module and the multi-layer perception module in each of the Transformer layers, a global pruning template for each weight matrix of the attention module and the multi-layer perception module in each of the Transformer layers is generated, including:
[0030] Obtain the input-related Hessian matrix of the target weight matrix, where the target weight matrix is any weight matrix of the attention module or the multi-layer perception module;
[0031] Generate a Hessian inverse matrix according to the input-related Hessian matrix of the target weight matrix, where the lower half of the Hessian inverse matrix is a zero matrix;
[0032] Calculate the error introduced by setting each weight in the target weight matrix to zero according to the Hessian inverse matrix;
[0033] Generate a global pruning template for the target weight matrix according to the error introduced by setting each weight in the target weight matrix to zero and the module target sparsity corresponding to the target weight matrix.
[0034] In an alternative embodiment of the present application, generating a global pruning template for the target weight matrix according to the error introduced by setting each weight in the target weight matrix to zero and the module target sparsity corresponding to the target weight matrix includes:
[0035] Sort the weights of the target weight matrix according to the error introduced by setting each weight in the target weight matrix to zero;
[0036] Determine the weight pruning quantity according to the module target sparsity corresponding to the target weight matrix and the total number of weights of the target weight matrix;
[0037] Determine the weights to be pruned in the target weight matrix according to the sorting of the weights in the target weight matrix and the number of weight pruning;
[0038] Generate a global pruning template for the target weight matrix according to the weights to be pruned in the target weight matrix.
[0039] In an alternative embodiment of the present application, according to the global pruning templates of each weight matrix of the attention module and the multi-layer perceptron module in each Transformer layer, perform weight pruning and reconstruction on the weight matrices of the attention module and the multi-layer perceptron module in each Transformer layer, so as to update each weight matrix in the attention module and the multi-layer perceptron module in each Transformer layer, including:
[0040] According to the global pruning template and the Hessian inverse matrix of each weight matrix of the attention module and the multi-layer perceptron module in each Transformer layer, perform weight pruning on the weight matrices of the attention module and the multi-layer perceptron module in each Transformer layer column by column, and use the OBS method to reconstruct the weights of the remaining unpruned columns, so as to update each weight matrix in the attention module and the multi-layer perceptron module in each Transformer layer.
[0041] To achieve the above object and other related objects, the present application provides a neural network sparsification system, the neural network includes a plurality of stacked Transformer layers, and each Transformer layer includes an attention module and a multi-layer perceptron module;
[0042] The system includes:
[0043] A data acquisition module, configured to inject calibration data into the neural network to obtain input feature data and output feature data of the attention module and the multi-layer perceptron module in each Transformer layer;
[0044] A similarity calculation module, configured to obtain the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer according to the input feature data and output feature data of the attention module and the multi-layer perceptron module in each Transformer layer;
[0045] A sparsity calculation module, configured to obtain the module target sparsity of the attention module and the multi-layer perception module in each Transformer layer according to the model target sparsity, the module preset quota ratio, and the input-output similarity of the attention module and the multi-layer perception module in each Transformer layer;
[0046] A template acquisition module, configured to generate a global pruning template for each weight matrix of the attention module and the multi-layer perception module in each Transformer layer according to the module target sparsity of the attention module and the multi-layer perception module in each Transformer layer;
[0047] A pruning and reconstruction module, configured to perform weight pruning and reconstruction on the weight matrices of the attention module and the multi-layer perception module in each Transformer layer according to the global pruning template of each weight matrix of the attention module and the multi-layer perception module in each Transformer layer, so as to update each weight matrix in the attention module and the multi-layer perception module in each Transformer layer.
[0048] To achieve the above object and other related objects, the present application provides an electronic device, and the electronic device includes:
[0049] One or more processors;
[0050] A memory, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device is enabled to implement the above neural network sparsity method.
[0051] To achieve the above object and other related objects, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor of a computer, the computer is enabled to execute the above neural network sparsity method.
[0052] As described above, the neural network sparsity method, system, device, and medium of the present application have the following beneficial effects:
[0053] By injecting calibration data into the neural network to obtain the input feature data and output feature data of the attention module and the multi-layer perceptron module in each Transformer layer; according to the input feature data and output feature data of the attention module and the multi-layer perceptron module in each Transformer layer, obtaining the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer; according to the model target sparsity, the module preset quota ratio, and the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer, obtaining the module target sparsity of the attention module and the multi-layer perceptron module in each Transformer layer; according to the module target sparsity of the attention module and the multi-layer perceptron module in each Transformer layer, generating a global pruning template for each weight matrix of the attention module and the multi-layer perceptron module in each Transformer layer; according to the global pruning template of each weight matrix of the attention module and the multi-layer perceptron module in each Transformer layer, performing weight pruning and reconstruction on the weight matrices of the attention module and the multi-layer perceptron module in each Transformer layer, so as to realize compressing the large language model to a high sparsity state through one-time weight pruning, without retraining, and still maintaining excellent performance; compared with the sparseGPT method, it has better accuracy performance in the benchmark test of high sparsity. Description of Drawings
[0054] Figure 1 It shows a schematic flowchart of the neural network sparsification method provided by an embodiment of the present application;
[0055] Figure 2 It shows a schematic diagram of the position for calculating the similarity signal of a Transformer layer, where the attention module (points A and B), and the multi-layer perceptron module (points B and C);
[0056] Figure 3 It shows the similarity graph of the attention module and the multi-layer perceptron module of each Transformer layer of LLaMA3.1 8B;
[0057] Figure 4 It shows a structural block diagram of the neural network sparsification system provided by an embodiment of the present application;
[0058] Figure 5 It shows a schematic structural diagram of an electronic device according to an embodiment of the present application. Detailed Description of the Invention
[0059] The following uses specific specific examples to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0060] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner. Therefore, only the components related to the present application are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0061] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.
[0062] The present application provides a neural network sparsification method called shrinkGPT, which is a post-training pruning method specifically tailored for large language models with the Transformer architecture. The shrinkGPT method of the present application can be compressed to a highly sparse state through one-time weight pruning without retraining and can also maintain excellent performance; compared with the sparseGPT method, it shows better accuracy in benchmark test performance at high sparsity levels.
[0063] Figure 1 The flowchart of the neural network sparsification method in an exemplary embodiment of the present application is shown, including steps S10 - step S50. The following will be combined with Figure 1 to elaborate on the technical solution of the present application in detail.
[0064] First, step S10 is executed to inject calibration data into the neural network to obtain the input feature data and output feature data of the attention module and the multi-layer perceptron module in each Transformer layer.
[0065] In this application, the neural network is a large language model (LLM) based on the Transformer architecture, including a number of stacked Transformer layers. Each Transformer layer includes an attention module and a multi-layer perceptron module using residual connections. The large language model can include large models such as LLaMA, Palm, GLM, BLOOM, GPT, etc.
[0066] In this application, the calibration data is batch calibration data, including multiple calibration data. When the calibration data is input into the neural network model, the model will first tokenize each calibration data into a series of continuous tokens through a tokenizer, and then perform word embedding operations on each token to map each token into a high-dimensional word vector in a high-dimensional vector space before inputting it into the model for processing. During the processing, the input feature data and output feature data corresponding to each token of each attention module and each multi-layer perceptron module can be obtained.
[0067] Next, step S20 is executed. According to the input feature data and output feature data of the attention module and the multi-layer perceptron module in each Transformer layer, the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer is obtained.
[0068] In a specific embodiment, when obtaining the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer according to the input feature data and output feature data of the attention module and the multi-layer perceptron module in each Transformer layer, first, according to the input feature data and output feature data of the attention module and the multi-layer perceptron module in each Transformer layer, the final output feature data of the attention module and the multi-layer perceptron module in each Transformer layer can be obtained; then, according to the input feature data and the final output feature data of the attention module and the multi-layer perceptron module in each Transformer layer, the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer is obtained.
[0069] The following takes the calculation process of the input-output similarity of the attention module as an example for illustration:
[0070] When calculating the input-output similarity of the attention module, the input feature vector data and output feature vector data corresponding to each Token of the attention module can be obtained first; the input feature vector data of each Token is superimposed on the output feature vector data as the final output feature vector data, and then the similarity between the input feature vector data and the final output feature vector data is calculated using a preset similarity algorithm as the similarity of the attention module for each Token; the similarity of the attention module for each Token of each calibration data is averaged to obtain the similarity of the attention module for each calibration data; finally, the similarity of the attention module for each calibration data is averaged to obtain the input-output similarity of the attention module, where the preset similarity algorithm includes cosine similarity algorithm, Euclidean similarity algorithm, Manhattan similarity algorithm, Pearson correlation coefficient algorithm. It should be noted that since the calculation of the input-output similarity of the attention module and the multi-layer attention module is similar, it will not be elaborated here.
[0071] Taking the cosine similarity algorithm as an example, the similarity calculation formula of the attention module or the multi-layer perception module for each Token is:
[0072]
[0073] In the formula, S represents the cosine similarity, and P1 and P2 are the input feature vector data and the final output feature vector data corresponding to each Token respectively. As Figure 2 shown, when calculating the similarity of the attention module for each Token, P1 and P2 are taken from the signals at points A and B; when calculating the similarity of the multi-layer perception module for each Token, P1 and P2 are taken from the signals at points B and C. |P1| and |P2| represent the Euclidean norms (i.e., the lengths of the vectors) of vectors P1 and P2 respectively.
[0074] Figure 3 shows the similarity graphs of the attention module (the attn curve in the figure) and the multi-layer perception module (the mlp curve in the figure) of 32 Transformer layers of the LLaMA 3.18B model. From Figure 3 it can be seen that in the LLaMA 3.18B model, the input-output similarities of the attention module and the multi-layer perception module of the Transformer layers at both ends are relatively low, while the input-output similarities of the attention module and the multi-layer perception module of the Transformer layers in the middle are relatively high, and the closer to the end, the higher the similarity of the attention module and the multi-layer perception module of the middle Transformer layers.
[0075] Next, step S30 is executed. According to the model target sparsity, the module preset quota ratio, and the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer, the module target sparsity of the attention module and the multi-layer perceptron module in each Transformer layer is obtained.
[0076] Specifically, first, according to the model target sparsity and the module preset quota ratio, the sparsity quotas of the attention module and the multi-layer perceptron module are calculated; then, according to the sparsity quotas of the attention module and the multi-layer perceptron module, and the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer, the module target sparsity of the attention module and the multi-layer perceptron module in each Transformer layer is obtained.
[0077] In this application, the module preset quota ratio is the quota ratio alpha of the attention module and the quota ratio 1-alpha of the multi-layer perceptron module. By introducing alpha, the sparsity quotas of the attention module and the multi-layer perceptron module can be adjusted, thus providing a flexible resource allocation strategy for the sparsification of the model. It can perform pruning operations targeted at the redundancy of the module, which helps to achieve more efficient model compression and optimization while maintaining the model performance. It should be noted that since the redundancy of the attention module is higher than that of the multi-layer perceptron module, the value of alpha usually ranges from 0.6 to 0.7, such as 0.6, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.7 or other values.
[0078] Among them, the sparsity quotas of the attention module and the multi-layer perceptron module are calculated by the following formula:
[0079] base attn = t·alpha
[0080] base mlp = t·(1-alpha)
[0081] In the formula, base attn represents the sparsity quota of the attention module, base mlp represents the sparsity quota of the multi-layer perceptron module, and t represents the model target sparsity.
[0082] To further make the pruning more targeted and help achieve more effective model compression while ensuring model performance, a preset amplification parameter can be introduced. According to the preset amplification parameter, the sparsity quota of the attention module and the multi-layer perceptron module, and the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer, the target module sparsity of the attention module and the multi-layer perceptron module in each Transformer layer is obtained.
[0083] Specifically, the following formula can be used to calculate the target module sparsity T of the attention module and the multi-layer perceptron module in each Transformer layer:
[0084]
[0085] In the formula, S represents the input-output similarity of the current attention module or multi-layer perceptron module; turbo represents the preset amplification parameter, and turbo ranges from 1.0 to 1.5, such as 1.0, 1.1, 1.2, 1.3, 1.4, 1.5; Mean(S) is the average input-output similarity of the attention module or multi-layer perceptron module in all Transformer layers in the model; base is the sparsity quota of the attention module or multi-layer perceptron module.
[0086] By introducing the preset amplification parameter, more pruning can be performed more specifically on the weight matrices of the attention module and the multi-layer perceptron module with high input-output similarity because they may contain more redundant information; while less pruning is performed on the weight matrices of the attention module and the multi-layer perceptron module with low input-output similarity because they may contain more unique or important information, and excessive pruning may have a greater impact on model performance. The present application can amplify the difference between the input-output similarity and the mean similarity of the attention module and the multi-layer perceptron module by introducing the preset amplification parameter to make this difference more obvious, so as to more clearly distinguish which module weight matrices should be pruned more or less, making the pruning more targeted and helping to achieve more effective model compression while ensuring model performance.
[0087] Next, step S40 is executed to generate a global pruning template for each weight matrix of the attention module and the multi-layer perceptron module in each Transformer layer according to the module target sparsity of the attention module and the multi-layer perceptron module in each Transformer layer. It should be noted that the global pruning template remains unchanged during the subsequent weight pruning process for each weight matrix.
[0088] In a specific embodiment of the present application, according to the module target sparsity of the attention module and the multi-layer perceptron module in each Transformer layer, generating a global pruning template for each weight matrix of the attention module and the multi-layer perceptron module in each Transformer layer may further include steps S41 - S44.
[0089] In step S41, obtain the input-related Hessian matrix H (also known as the Hessian matrix, Hesse matrix) of the target weight matrix W, where the target weight matrix W is any weight matrix of the attention module or the multi-layer perceptron module.
[0090] It should be noted that each attention module includes four weight matrices, namely the Q weight matrix of the Q processing unit, the K weight matrix of the K processing unit, the V weight matrix of the V processing unit, and the O weight matrix of the O processing unit, which correspond respectively. Among them, the Q weight matrix is also called the query weight matrix, the K weight matrix is also called the key weight matrix, the V weight matrix is also called the value weight matrix, and the O weight matrix is also called the output weight matrix; the multi-layer perceptron module includes three weight matrices, namely the up weight matrix of the up processing unit, the gate weight matrix of the gate processing unit, and the down weight matrix of the down processing unit. Among them, the up weight matrix is also called the upsampling weight matrix or upscaling weight matrix, the gate weight matrix is also called the gating weight matrix or control weight matrix, and the down weight matrix is the downsampling weight matrix or downscaling weight matrix.
[0091] When obtaining the input-related Hessian matrix H of the target weight matrix W, in step S10, after injecting calibration data into the neural network, obtain the input feature data X of the processing unit corresponding to the target weight matrix W, and construct the input-related Hessian matrix H of the target weight matrix W according to the input feature data by the following formula:
[0092] H = (XX T + λI)
[0093] In the formula, X represents the input feature data, X T represents the transpose of the input feature data, I represents the identity matrix, λ is a scalar, λI is a diagonal matrix, whose diagonal elements are λ and the rest are 0. Adding λI is a means of regularization, which can prevent the matrix from being singular or close to singular in optimization.
[0094] In step S42, generate the inverse Hessian matrix H -1 according to the input-related Hessian matrix H of the target weight matrix W, where the inverse Hessian matrix H -1A matrix with the lower half being zero. Specifically, the inverse Hessian matrix H of the Hessian matrix H can be generated through Cholesky decomposition -1 , and ensure that the lower half of the inverse Hessian matrix H -1 is 0, so that during the subsequent reconstruction process after weight pruning, the weights contained in the pruned columns are frozen, ensuring that the structure of the weight matrix is not damaged.
[0095] In step S43, according to the inverse Hessian matrix H -1 calculate the error introduced by setting each weight in the target weight matrix W to zero, where the calculation formula is as follows:
[0096]
[0097] In the formula, L q represents the error introduced by setting the q-th weight of the weight matrix to zero; q represents the weight index, and w q represents the q-th weight of the weight matrix; H -1 represents the inverse Hessian matrix of the input-related Hessian matrix of the weight matrix.
[0098] In step S44, according to the error introduced by setting each weight in the target weight matrix W to zero, and the module target sparsity corresponding to the target weight matrix W, generate the global pruning template M of the target weight matrix W. It should be noted that the module target sparsity corresponding to the target weight matrix W refers to the module target sparsity of the attention module or multi-layer perceptron module to which the target weight matrix W belongs. In other words, the weight matrices of each attention module or multi-layer perceptron module share the module target sparsity.
[0099] In an optional embodiment of the present application, when generating the global pruning template of the target weight matrix according to the error introduced by setting each weight in the target weight matrix to zero and the module target sparsity corresponding to the target weight matrix, the weights of the target weight matrix can be sorted first according to the error introduced by setting each weight in the target weight matrix to zero; then, determine the weight pruning quantity according to the module target sparsity corresponding to the target weight matrix and the total number of weights of the target weight matrix; then, determine the weights to be pruned in the target weight matrix according to the sorting of the weights in the target weight matrix and the weight pruning quantity; finally, generate the global pruning template of the target weight matrix according to the weights to be pruned in the target weight matrix. The generated global pruning template can be used to guide the pruning of the target weight matrix. After setting a large number of unimportant weights to zero, only the non-zero weights and their position information need to be saved when storing the model, greatly reducing the storage space of the model.
[0100] As an example, when sorting the weights of the target weight matrix from smallest to largest in terms of error, weights that are ranked higher and equal in number to the weights to be pruned can be selected as the weights to be pruned in the target weight matrix to generate a matrix as the global pruning template. This matrix has the same size as the target weight matrix and can be formed by setting the weights to be pruned in the target weight matrix to 0 and the corresponding weights not to be pruned to 1.
[0101] It should be noted that by considering the error introduced by setting each weight to zero, it is possible to accurately determine which weights have less impact on the performance of the model. Setting the corresponding positions of these weights to zero in the global pruning template can remove redundant information in the model without significantly reducing the model accuracy, making the model more lightweight, thereby improving the running speed and efficiency of the model.
[0102] It should be noted that since the module target sparsity reflects the importance of the module in the entire model and its influence on the output result, generating a global pruning template in combination with the module target sparsity can ensure that modules important for the model performance will not be over-pruned during the pruning process, and can maintain the accuracy and generalization ability of the model to the greatest extent while optimizing the model structure.
[0103] Finally, perform step S50 to perform weight pruning and reconstruction on the weight matrices of the attention module and the multi-layer perceptron module in each Transformer layer according to the global pruning template of each weight matrix of the attention module and the multi-layer perceptron module in each Transformer layer, so as to update each weight matrix in the attention module and the multi-layer perceptron module in each Transformer layer.
[0104] In a specific embodiment of the present application, according to the global pruning template and the Hessian inverse matrix of each weight matrix of the attention module and the multi-layer perceptron module in each Transformer layer, weight pruning can be performed column by column on the weight matrices of the attention module and the multi-layer perceptron module in each Transformer layer, and the weights of the remaining unpruned columns can be reconstructed using the classical OBS method to update each weight matrix in the attention module and the multi-layer perceptron module in each Transformer layer.
[0105] Use the classical OBS (Optimal Brain Surgeon) method to reconstruct and update the weights of the remaining unpruned columns. The preset formula is
[0106]
[0107] Where, δw represents the compensation value of the weight; q represents the weight index, and w q represents the q-th weight of the weight matrix; H -1 represents the Hessian inverse matrix of the input-related Hessian matrix of the weight matrix; e q represents w q corresponding unit vector.
[0108] In a specific embodiment of the present application, the pseudo-code of the shrinkGPT algorithm is as follows:
[0109]
[0110]
[0111] The pruning performance tests of the shrinkGPT method of the present application and the existing sparseGPT method are shown in Table 1 and Table 2. Among them, Table 1 and Table 2 are the benchmark test results of shrinkGPT and sparseGPT on LLaMA3.1 8B at 40% and 50% model target sparsity, respectively, using the GSM8K and TruthfulQA benchmarks.
[0112] Table 1 llama3.1 8B 40% Pruning Benchmark Test
[0113]
[0114] Table 2 llama3.1 8B 50% Pruning Benchmark Test
[0115]
[0116] Note: 1. sparseGPT uses the C4 dataset, 64 samples, block size 1024;
[0117] 2. shrinkGPT uses alpha = 0.65, C4 dataset, turbo = 1.0;
[0118] 3. GSM8K is tested with a length of 60 questions, enabling CoT, n_shots = 3;
[0119] 4. TruthfulQA health has 55 single-choice questions, and TruthfulQA history has 24 single-choice questions.
[0120] 5. The values in the Avg column are weighted averages.
[0121] In the benchmark test, compared with the sparseGPT method, the shrinkGPT method of the present application shows more excellent performance at high sparsity. Moreover, as the sparsity continues to increase, the benchmark test performance of the shrinkGPT method becomes increasingly outstanding.
[0122] Based on the same concept, as Figure 4 shown, the present application also provides a neural network sparsity system 11, and the neural network sparsity system 11 includes a data acquisition module 111, a similarity calculation module 112, a sparsity calculation module 113, a template acquisition module 114, and a pruning and reconstruction module 115.
[0123] Among them, the data acquisition module 111 is used to inject calibration data into the neural network to obtain the input feature data and output feature data of the attention module and the multi-layer perceptron module in each Transformer layer;
[0124] The similarity calculation module 112 is used to obtain the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer according to the input feature data and output feature data of the attention module and the multi-layer perceptron module in each Transformer layer;
[0125] The sparsity calculation module 113 is used to obtain the module target sparsity of the attention module and the multi-layer perceptron module in each Transformer layer according to the model target sparsity, the module preset quota ratio, and the input-output similarity of the attention module and the multi-layer perceptron module in each Transformer layer;
[0126] The template acquisition module 114 is used to generate a global pruning template for each weight matrix of the attention module and the multi-layer perceptron module in each Transformer layer according to the module target sparsity of the attention module and the multi-layer perceptron module in each Transformer layer;
[0127] The pruning and reconstruction module 115 is used to perform weight pruning and reconstruction on the weight matrices of the attention module and the multi-layer perceptron module in each Transformer layer according to the global pruning template of each weight matrix of the attention module and the multi-layer perceptron module in each Transformer layer, so as to update each weight matrix in the attention module and the multi-layer perceptron module in each Transformer layer.
[0128] It should be noted that the neural network sparsity system 11 provided in the above embodiments and the neural network sparsity method provided in the above embodiments belong to the same concept. The specific ways in which each module and unit perform operations have been described in detail in the method embodiments, and will not be repeated here. In practical applications, the neural network sparsity system 11 provided in the above embodiments can allocate the above functions to different functional modules as needed, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. No limitation is imposed herein.
[0129] As Figure 5 shown, it is a schematic structural diagram of an electronic device for implementing the neural network sparsity method of the present application.
[0130] The electronic device 1 may include a memory 12, a processor 13, and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a neural network sparsity program.
[0131] Among them, the memory 12 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 12 may be an internal storage unit of the electronic device 1 in some embodiments, such as the mobile hard disk of the electronic device 1. The memory 12 may also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 12 may include both the internal storage unit and the external storage device of the electronic device 1. The memory 12 can not only be used to store application software installed in the electronic device 1 and various types of data, such as the code for neural network sparsity, but also be used to temporarily store data that has been output or will be output.
[0132] In some embodiments, the processor 13 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control core of the electronic device 1. It uses various interfaces and circuits to connect all components of the entire electronic device 1. By running or executing programs or modules (such as neural network sparsity programs, etc.) stored in the memory 12, and by calling the data stored in the memory 12, it executes various functions of the electronic device 1 and processes data.
[0133] The processor 13 executes the operating system of the electronic device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above neural network sparsity method.
[0134] Exemplarily, the computer program may be divided into one or more modules. The one or more modules are stored in the memory 12 and executed by the processor 13 to complete this application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into a data acquisition module 111, a similarity calculation module 112, a sparsity calculation module 113, a template acquisition module 114, and a clipping and reconstruction module 115.
[0135] The above integrated units implemented in the form of software function modules can be stored in a computer-readable storage medium. The computer-readable storage medium may be non-volatile or volatile. The above software function modules are stored in a storage medium and include several instructions to enable a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute some functions of the neural network sparsity method described in various embodiments of this application.
[0136] In summary, a neural network sparsification method, system, device, and medium disclosed in this application inject calibration data into the neural network to obtain the input feature data and output feature data of the attention module and the multi-layer perception module in each Transformer layer; according to the input feature data and output feature data of the attention module and the multi-layer perception module in each Transformer layer, obtain the input-output similarity of the attention module and the multi-layer perception module in each Transformer layer; according to the model target sparsity, the module preset quota ratio, and the input-output similarity of the attention module and the multi-layer perception module in each Transformer layer, obtain the module target sparsity of the attention module and the multi-layer perception module in each Transformer layer; according to the module target sparsity of the attention module and the multi-layer perception module in each Transformer layer, generate a global pruning template for each weight matrix of the attention module and the multi-layer perception module in each Transformer layer; according to the global pruning template of each weight matrix of the attention module and the multi-layer perception module in each Transformer layer, perform weight pruning and reconstruction on the weight matrices of the attention module and the multi-layer perception module in each Transformer layer, so as to achieve compressing the large language model to a highly sparse state through a single weight pruning, without retraining and still maintaining excellent performance; compared with the sparseGPT method, it shows better accuracy in the benchmark test performance at high sparsity. Therefore, this application effectively overcomes various drawbacks in the prior art and has high industrial utilization value.
[0137] The above embodiments are only illustrative of the principles and effects of this application and are not intended to limit this application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed in this application should still be covered by the claims of this application.
Claims
1. A neural network sparse method, characterized in that: The neural network includes a plurality of stacked Transformer layers, each Transformer layer includes an attention module and a multi-layer perception module; The method comprises: Injecting calibration data into the neural network to obtain input feature data and output feature data of the attention module and the multi-layer perception module in each of the Transformer layers; According to the input feature data and output feature data of the attention module and the multi-layer perception module in each of the Transformer layers, the input and output similarities of the attention module and the multi-layer perception module in each of the Transformer layers are obtained; According to the model target sparsity, the module preset quota ratio, and the input-output similarity of the attention module and the multi-layer perception module in each of the Transformer layers, the module target sparsity of the attention module and the multi-layer perception module in each of the Transformer layers is obtained; Generate a global clipping template for each weight matrix of the attention module and the multi-layer perception module in each Transformer layer according to the module target sparsity of the attention module and the multi-layer perception module in each Transformer layer; According to the global clipping template of each weight matrix of the attention module and the multi-layer perception module in each Transformer layer, the weight matrices of the attention module and the multi-layer perception module in each Transformer layer are weight clipped and reconstructed.
2. The neural network sparse method according to claim 1, characterized in that: According to the input feature data and the output feature data of the attention module and the multi-layer perception module in each of the Transformer layers, the input and output similarity of the attention module and the multi-layer perception module in each of the Transformer layers is obtained, including: According to the input feature data and output feature data of the attention module and the multi-layer perception module in each of the Transformer layers, obtaining the final output feature data of the attention module and the multi-layer perception module in each of the Transformer layers; According to the input feature data and final output feature data of the attention module and the multi-layer perception module in each Transformer layer, the input and output similarity of the attention module and the multi-layer perception module in each Transformer layer is obtained.
3. The neural network sparse method according to claim 2, characterized in that: According to the input feature data and the final output feature data of the attention module and the multi-layer perception module in each of the Transformer layers, the input and output similarity of the attention module and the multi-layer perception module in each of the Transformer layers is obtained, including: According to the input feature data and final output feature data of the attention module and the multi-layer perception module in each of the Transformer layers, a preset similarity algorithm is used to obtain the input and output similarity of the attention module and the multi-layer perception module in each of the Transformer layers; The preset similarity algorithms include cosine similarity algorithm, Euclidean similarity algorithm, Manhattan similarity algorithm, and Pearson correlation coefficient algorithm.
4. The neural network sparse method according to claim 1, characterized in that: According to the model target sparsity, the module preset quota ratio, and the input-output similarity of the attention module and the multi-layer perception module in each of the Transformer layers, the module target sparsity of the attention module and the multi-layer perception module in each of the Transformer layers is obtained, including: According to the target sparsity of the model and the preset quota ratio of the module, the sparse quota of the attention module and the multi-layer perception module are calculated and obtained; According to the sparse quotas of the attention module and the multi-layer perception module and the input-output similarities of the attention module and the multi-layer perception module in each Transformer layer, the module target sparsity of the attention module and the multi-layer perception module in each Transformer layer is obtained.
5. The neural network sparse method according to claim 4, characterized in that: According to the sparse quotas of the attention module and the multi-layer perception module, and the input-output similarities of the attention module and the multi-layer perception module in each of the Transformer layers, the module target sparsity of the attention module and the multi-layer perception module in each of the Transformer layers is obtained, including: According to the preset amplification parameters, the sparse quotas of the attention module and the multi-layer perception module, and the input-output similarity of the attention module and the multi-layer perception module in each Transformer layer, the target module sparsity of the attention module and the multi-layer perception module in each Transformer layer is obtained.
6. The neural network sparse method according to claim 5, characterized in that: The target module sparsity T of the attention module and the multi-layer perception module in each Transformer layer is calculated using the following formula: Where S represents the input-output similarity of the current attention module or multi-layer perception module; turbo represents the preset amplification parameter; Mean(S) is the average input-output similarity of the attention module or multi-layer perception module in all Transformer layers; base is the sparse quota of the attention module or multi-layer perception module.
7. The neural network sparse method according to claim 1, characterized in that: According to the module target sparsity of the attention module and the multi-layer perception module in each of the Transformer layers, a global clipping template of each weight matrix of the attention module and the multi-layer perception module in each of the Transformer layers is generated, including: Obtaining an input-related Hessian matrix of a target weight matrix, wherein the target weight matrix is any weight matrix of the attention module or the multi-layer perception module; Generate a Hessian inverse matrix according to the input-related Hessian matrix of the target weight matrix, wherein the lower half of the Hessian inverse matrix is a matrix of zeros; Calculate the error introduced by resetting each weight in the target weight matrix to zero according to the Hessian inverse matrix; A global clipping template of the target weight matrix is generated according to the error introduced by resetting each weight in the target weight matrix and the module target sparsity corresponding to the target weight matrix.
8. The neural network sparse method according to claim 7, characterized in that: According to the error introduced by resetting each weight in the target weight matrix and the module target sparsity corresponding to the target weight matrix, a global clipping template of the target weight matrix is generated, including: Sorting the weights of the target weight matrix according to the error introduced by resetting each weight in the target weight matrix; Determine the weight clipping amount according to the module target sparsity corresponding to the target weight matrix and the total number of weights of the target weight matrix; Determine the weights to be pruned in the target weight matrix according to the order of the weights in the target weight matrix and the amount of weight pruned; A global clipping template of the target weight matrix is generated according to the weights to be clipped in the target weight matrix.
9. The neural network sparse method according to claim 7, characterized in that: According to the global clipping template of each weight matrix of the attention module and the multi-layer perception module in each of the Transformer layers, weight clipping and reconstruction are performed on the weight matrix of the attention module and the multi-layer perception module in each of the Transformer layers, including: According to the global pruning template and the Hessian inverse matrix of each weight matrix of the attention module and the multi-layer perception module in each Transformer layer, the weight matrices of the attention module and the multi-layer perception module in each Transformer layer are weight pruned column by column, and the weights of the remaining unpruned columns are reconstructed using the OBS method.
10. A neural network sparse system, characterized in that: The neural network includes a plurality of stacked Transformer layers, each Transformer layer includes an attention module and a multi-layer perception module; The system comprises: A data acquisition module, used to inject calibration data into the neural network to obtain input feature data and output feature data of the attention module and the multi-layer perception module in each of the Transformer layers; A similarity calculation module, used to obtain the input and output similarity of the attention module and the multi-layer perception module in each of the Transformer layers according to the input feature data and output feature data of the attention module and the multi-layer perception module in each of the Transformer layers; A sparsity calculation module, used to obtain the module target sparsity of the attention module and the multi-layer perception module in each of the Transformer layers according to the model target sparsity, the module preset quota ratio, and the input-output similarity of the attention module and the multi-layer perception module in each of the Transformer layers; A template acquisition module, used to generate a global clipping template of each weight matrix of the attention module and the multi-layer perception module in each of the Transformer layers according to the module target sparsity of the attention module and the multi-layer perception module in each of the Transformer layers; A cropping and reconstruction module is used to perform weight cropping and reconstruction on the weight matrices of the attention module and the multi-layer perception module in each Transformer layer according to a global cropping template of each weight matrix of the attention module and the multi-layer perception module in each Transformer layer.
11. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the neural network sparse method as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the neural network sparse method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Non-intrusive load decomposition method based on mathematical morphology and improved Transform
CN115905857A
Sparse attention computation model and method, electronic device, and storage medium
WO2023221940A1