Matrix decomposition and reconstruction method and system for large model reasoning scene and application
Through layer-by-layer computing link sensitivity analysis and progressive low-rank settings, the matrix decomposition and reconstruction method for large-model inference scenarios solves the problem of insufficient resources when deploying large-models to end-side devices, and achieves efficient model deployment and application performance improvement.
Patent Information
- Application Number
- CN202410770430.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-14
- Publication Date
- 2025-05-09
AI Technical Summary
When deploying large models to end-side devices, the prior art is limited by insufficient resources, resulting in reduced model performance and increased application costs. The existing error analysis methods cannot effectively evaluate the loss impact of the overall network.
A matrix decomposition and reconstruction method for large model inference scenarios is proposed. The decomposed weight nodes are filtered through the sensitivity analysis strategy of the layer-by-layer calculation link, and the appropriate low-rank value is selected using the progressive low-rank setting mode, the model parameters are reduced, and the network reconstruction is carried out based on the decomposed matrix.
This method does not require the introduction of a training process, and directly performs low-rank decomposition of the pre-training parameters of the large model, reduces memory usage and calculation time, improves the deployment and application performance of the large model, and is suitable for end-side devices with resource limitations.
Smart Images

Figure BDA0004893918280000021 
Figure BDA0004893918280000041 
Figure BDA0004893918280000101
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large model matrix decomposition, and relates to a matrix decomposition and reconstruction method, system and application for large model reasoning scenarios. Background Art
[0002] Matrix decomposition technology is currently widely used in model training scenarios. During training, matrices with large dimensions are reduced in dimension through low-rank decomposition technology to reduce the training parameters of the model and speed up the training. However, for large models, due to their huge parameters, generally in the hundreds of millions, they are currently mainly used in inference scenarios for deployment on end-side devices; and end-side devices generally have problems such as limited resources. Insufficient memory and computing power will significantly affect model performance, thereby increasing application costs. Therefore, how to efficiently deploy large models to end-side devices is crucial, and matrix decomposition technology is a feasible method because it can significantly reduce model parameters, thereby reducing the device's demand for memory and computing power. Most of the current methods based on large models combined with matrix decomposition require the introduction of retraining steps, and the training process requires a lot of time and resources, including data collection and organization, training parameter adjustment, and subsequent processing of model parameters. It involves too much manpower and material costs and does not utilize the deployment and application of large models in actual scenarios.
[0003] The number of network parameters of large AI models is generally over 100 million, and these parameters are redundant. That is, the model itself does not need to be represented by too many parameters and can be well applied to various downstream tasks. As a result, various compression technologies have emerged, such as quantization, pruning, distillation and other methods, which reduce network parameters and save storage and computing space.
[0004] In addition, the existing common error analysis methods generally refer to calculating the difference between the modified value and the original value for a certain parameter value, and then calculating the error percentage. However, the above error analysis method can only evaluate the error rate of the current single node, and cannot effectively evaluate the impact on the loss of the entire network. The current mainstream AI large models, such as Llama-7B, are based on the Decoder structure in the Transformer model. Most of the structure is matrix multiplication and addition operations. In addition, it also contains a large number of normalization operations and nonlinear transformation operations. Normalization and nonlinear transformation operations adjust the data distribution of the input value, thereby affecting the change in value and the output result. In other words, even if the error value of a node itself is large, it may not affect the subsequent value of the network after normalization and activation operations, so it can be decomposed. Summary of the invention
[0005] In order to solve the deficiencies in the prior art, the purpose of the present invention is to provide a matrix decomposition and reconstruction method, system and application for large model reasoning scenarios, which can be used in scenarios such as artificial intelligence, natural language processing, large model applications, model reasoning acceleration, etc. The present invention proposes a matrix decomposition and reconstruction method for large model reasoning scenarios. This method does not need to introduce a training process, directly performs low-rank decomposition on the existing large model pre-training parameters, and performs subsequent reasoning processes based on the decomposed model parameters and model structure. The overall process does not require too much manual participation and can be efficiently applied to various AI large model reasoning scenarios, while reducing memory usage and computing time, and improving large model deployment and application performance.
[0006] The present invention proposes a matrix decomposition method for accelerating large model reasoning. First, the weight nodes that can participate in the decomposition are screened out by calculating the link strategy layer by layer, and the nodes are reasonably selected through error analysis; then, the most suitable low-rank value is searched for the matrix to be decomposed by using a progressively reduced compression ratio, which can reduce the model parameters as much as possible while ensuring the accuracy of the model, and improve the reasoning performance; finally, based on the acquired matrix and the corresponding low-rank value, the network of the original large model is reconstructed, and the model reasoning is performed based on the reconstructed network, which effectively improves the deployment efficiency and reduces the consumption of memory and computing resources for large models. Compared with the current mainstream retraining method combined with matrix decomposition, the method proposed in the present invention is convenient and efficient, more friendly to resource-constrained scenarios, equivalent to the process of post-processing the model structure and weight parameters, without too much manual intervention, and simple to deploy.
[0007] Specifically,
[0008] The present invention proposes a matrix decomposition and reconstruction method for large model reasoning scenarios, the method comprising the following steps:
[0009] Step 1: Perform error statistics on conventional nodes in the network layer of the large model, and select weight nodes that can perform low-rank decomposition through sensitivity analysis strategy;
[0010] Step 2: progressively perform low-rank setting, decompose the weight matrices in the weight nodes that can be decomposed by low-rank decomposition obtained by screening in step 1, and obtain and record the optimal low-rank values of the decomposable weight matrices;
[0011] Step 3: Reconstruct the network of the original model based on the decomposable weight matrix and the corresponding low-rank values.
[0012] In step 1, when performing error statistics, each network layer in the large model is used as the basic unit and statistics are performed separately; the sensitivity analysis strategy refers to comparing the error of the conventional node with the preset sensitivity threshold S. If the error of the conventional node is less than the sensitivity threshold, the weight matrix in the weight node closest to the conventional node and using the approximate matrix to participate in the calculation can be low-rank decomposed; conversely, if the error of the conventional node is greater than the sensitivity threshold, the weight matrix cannot be low-rank decomposed; only the weight node has a decomposition error problem; the error of the conventional node comes from the numerical offset caused by the decomposition of its preceding weight node;
[0013] The error of the conventional node is calculated by the following formula:
[0014]
[0015] In a specific implementation, in step one, the preset sensitivity threshold is 3-5%; since different layers have similar data processing operations, processing in layers can improve processing efficiency and ensure data integrity.
[0016] The step 1 sets the screening compression ratio first, and the value of the screening compression ratio is small, for example, 2 times compression. The corresponding low-rank value can be obtained according to the compression ratio, and the specific values of the two decomposed matrices A and B can be calculated by a common low-rank decomposition algorithm; and can further include the following steps:
[0017] Step 1.1, generating simulation data according to the distribution characteristics of the data received by the model, inputting the simulation data into the original model, and obtaining and recording the input and output of all conventional nodes;
[0018] Step 1.2: Take each network layer in the model as a unit to perform weight matrix screening: traverse along the calculation link of the current network layer, and when encountering the first weight matrix, compress it based on the preset screening compression ratio, and replace the original weight matrix W value with the value of the sub-weight matrix A*B;
[0019] Step 1.3, continue traversing using the value of the sub-weight matrix A*B obtained in step 1.2, perform linear calculations on this value and the input value of the current weight node to generate output data, which is used as the input value of the regular node. For each regular node that receives this value, calculate the error percentage between its output value and the recorded original output value; if the error percentage is within the sensitivity threshold S, record the name and value of the previously decomposed weight matrix and continue traversing; if the error percentage is greater than the sensitivity threshold, continue traversing along the calculation link until a regular node or the next weight node that meets the requirements is encountered; if the error percentages of all regular nodes in the current network layer do not meet the requirements, the previously decomposed weight matrix does not meet the requirements and does not need to be recorded;
[0020] Step 1.4: Get the names and values of the decomposable weight matrices in all network layers.
[0021] In step 2, a range of compression multiples of the weight matrix is set, and the range of the compression multiples is divided into equal intervals to obtain one or more compression multiples; the obtained compression multiples are calculated in descending order to obtain low-rank values, the decomposable weight matrix is decomposed, and the weight matrix error percentage is calculated;
[0022] The compression multiple is greater than 1 and less than or equal to 4; the number of compression multiples after equal interval division can be set arbitrarily; in a specific implementation scheme, the compression multiple range can be equally divided into 3-5;
[0023] In a specific implementation, the maximum compression factor is set to 4 and the minimum compression factor is set to 1.5;
[0024] The calculation method of the low rank value corresponding to each compression factor is as follows:
[0025] Assume compression factor = P, then P = (m*n) / (m*k+k*n), m and n are the dimensions of the known weight matrix W to be compressed, from which the size of the low rank value (k value) can be calculated, k = (m*n) / [P*(m+n)];
[0026] When the error percentage of the weight matrix is less than the preset error percentage threshold J, the weight matrix supports the current compression multiple, and the weight matrix name and the low-rank value corresponding to the current compression multiple are recorded, and the above operation is repeated until the optimal low-rank value of all weight matrices is obtained;
[0027] The preset error percentage threshold can generally be set to 3%-5%, which is used to evaluate the numerical error percentage before and after the weight matrix decomposition;
[0028] Since the compression ratio list is arranged in descending order, when traversing in order, if the current larger compression ratio is used and the error percentage of the weight matrix meets the requirement, only the low-rank value corresponding to the compression ratio at this time is recorded, and there is no need to traverse the remaining compression ratios again. The purpose of this operation is to compress the weight matrix as much as possible while meeting a certain error condition, so only the low-rank value corresponding to the largest compression ratio is retained.
[0029] The calculation method of the weight matrix error percentage is shown in the following formula:
[0030]
[0031] Among them, A and B are sub-matrices obtained by decomposing the weight matrix W, and m and n are the dimensions of the weight matrix W.
[0032] After obtaining the optimal low-rank value of all weight matrices, the final parameter compression rate of the large model can be estimated. The total number of parameters in the original model is divided by the total number of parameters in the compressed and reconstructed model to obtain the final model compression rate.
[0033] In step three, the decomposed matrix corresponding to the optimal low-rank value replaces the original weight matrix to reconstruct the original model.
[0034] The present invention also provides a system for implementing the above-mentioned matrix decomposition and reconstruction method, the system comprising: an error statistics module, a low-rank decomposition module, a model reconstruction module, and a parameter compression rate estimation module;
[0035] The error statistics module is used to perform error statistics on conventional nodes of the network layer in the large model, and to screen weight nodes capable of low-rank decomposition through a sensitivity analysis strategy;
[0036] The low-rank decomposition module is used to perform low-rank decomposition on the weight matrix in the screened weight node, and obtain and record the optimal low-rank value of each decomposable weight matrix;
[0037] The model reconstruction module is used to reconstruct the network of the original model based on the decomposable matrix and the corresponding low-rank value;
[0038] The parameter compression rate estimation module is used to compare the total parameter amount in the original model with the total parameter amount in the compressed and reconstructed model to obtain the final model compression rate.
[0039] The present invention also provides the above-mentioned matrix decomposition and reconstruction method, or the application of the above-mentioned system in large-model reasoning acceleration, large-model end-side device deployment, large-model application, etc.
[0040] The present invention also provides a hardware system for implementing the above-mentioned matrix decomposition and reconstruction method, and the hardware system includes: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the above-mentioned matrix decomposition and reconstruction method is implemented.
[0041] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned matrix decomposition and reconstruction method is implemented.
[0042] The beneficial effects of the present invention include: the matrix decomposition and reconstruction method of the present invention for large model reasoning scenarios, the sensitivity analysis strategy based on the layer-by-layer calculation link can effectively screen out weight matrices with low decomposition sensitivity, the progressive low rank setting mode based on the decomposition error can effectively select appropriate low rank values, ensure model accuracy while reducing model parameters as much as possible, and the matrix decomposition method for large model reasoning acceleration scenarios can effectively reduce the memory usage of large models and improve deployment efficiency.
[0043] The method of the present invention does not require manual participation in the screening of decomposable weight matrices and the selection of optimal low-rank values, and the overall process is faster and more efficient. Although the method of the present invention modifies the original network structure, it reduces some redundant parameters and does not affect the representation ability of the model. Therefore, it can be combined with other large model acceleration methods such as quantization and distillation to improve the running speed of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0045] Figure 1 It is a schematic diagram of the sensitivity analysis process based on layer-by-layer link calculation of the present invention.
[0046] Figure 2 It is a schematic diagram of the progressive low-rank setting mode based on decomposition error of the present invention.
[0047] Figure 3 It is a schematic diagram of a network reconstruction method based on linear layer replacement according to the present invention. DETAILED DESCRIPTION
[0048] The present invention is further described in detail with reference to the following specific examples and drawings. The process, conditions, experimental methods, etc. for implementing the present invention, except for the contents specifically mentioned below, are all common knowledge and common common sense in the art and are not particularly limited by the present invention.
[0049] The present invention proposes a matrix decomposition and reconstruction method for large model reasoning scenarios, and the matrix decomposition and reconstruction method includes the following steps: step 1, error statistics of conventional nodes in the network layer in the large model, and screening weight nodes that can be low-rank decomposition through sensitivity analysis strategy; step 2, progressive low-rank setting, decomposing the weight matrix in the weight node that can be low-rank decomposition obtained by screening in step 1, and obtaining and recording the optimal low-rank value of each decomposable weight matrix; step 3, based on the decomposable weight matrix and the corresponding low-rank value, the original model is reconstructed. The present invention also provides a system for implementing the above method, and corresponding applications, which have a wide range of application scenarios.
[0050] The present invention is oriented to the application scenario of large model reasoning acceleration, and is used to reduce the reasoning time of large models on devices. First, an error analysis module is added behind the conventional nodes of each network layer of the large model, and the module is used to calculate the error percentage of the current conventional node, so as to evaluate the sensitivity of the current conventional node; then, according to the set sensitivity threshold, it is determined whether the weight matrix in the weight node closest to the current conventional node participates in the low-rank decomposition operation, and the sensitivity threshold is used to measure whether the error percentage meets the requirements, that is, whether the error percentage is within the set sensitivity threshold range. In a specific implementation scheme, the sensitivity threshold can be set to 3%; when the error percentage is less than 3%, it can be considered that the operation of low-rank pre-decomposition of the weight matrix will not significantly affect the output value of the subsequent node and the final accuracy of the model; then, for the decomposable weight matrix obtained in the above steps, an automatic search strategy is adopted to obtain the optimal low-rank value of each weight matrix, and the automatic search strategy is: by calculating the decomposition error value corresponding to the different compression rates of the weight matrix, the weight matrix parameters are compressed as much as possible under the condition of ensuring that the error is controllable, so as to obtain the optimal low-rank value; finally, matrix decomposition is performed based on the obtained low-rank dimension, and the large model network structure is reconstructed. The above invention technology can effectively reduce the network parameters of large models, thereby reducing the dependence on memory space and computing resources during model reasoning, thereby speeding up model reasoning, improving the efficiency of large model deployment, and reducing the application cost of large models.
[0051] The large model is formed by copying and stacking multiple identical network layers, and the network layer contains multiple nodes. The nodes can be divided into regular nodes and weight nodes according to their functions. Regular nodes: refers to general numerical calculations, such as tensor multiplication and addition, tensor splicing, numerical normalization, nonlinear mapping of activation functions, etc. The node receives multiple inputs and performs multiple operations between each input; Weight nodes: refers to nodes containing weight matrices, which are generally initialized using nn.Linear() in large models. The node receives an input, which performs linear operations with the weight matrix, and the output result is used for subsequent calculations. In the present invention, the error of the output value of the regular node is first calculated, and then it is determined whether to decompose the weight matrix in the corresponding weight node based on the error result.
[0052] The present invention first mainly decomposes the weight matrix of the weight nodes in each network layer of the large model, and then uses the decomposed low-dimensional matrix to participate in the multiplication and addition operation. Therefore, the more compressed matrices, the greater the benefit of the method. Therefore, a sensitivity analysis strategy based on layer-by-layer calculation of links is proposed, that is: taking each network layer in the large model as the basic unit, the error statistics of the conventional nodes therein are performed. If the error percentage of the conventional node is less than the sensitivity threshold set in advance, it can be considered that the weight matrix in the weight node closest to the conventional node and using the approximate matrix to participate in the calculation can be subjected to low-rank decomposition operation, because the approximate operation of the front calculation path does not significantly affect the subsequent output value of the conventional node, and most of the errors before the node position have been eliminated. At this time, the name and value of the weight matrix are recorded. This method focuses on the overall network and is more convenient and efficient.
[0053] Specifically,
[0054] For the step of screening the weight matrix, first set a smaller screening compression ratio, such as 2 times compression. According to the compression ratio, the corresponding low-rank value can be obtained, and the specific values of the two sub-weight matrices A and B can be calculated through the common low-rank decomposition algorithm. In the process of screening the weight matrix, refer to the following process:
[0055] a. First, according to the distribution characteristics of the data received by the model, a batch of simulation data is generated, and the simulation data is input into the original model to obtain and record the input and output values of all regular nodes;
[0056] b. Then, the weight matrix in the layer is screened based on the network layer. It should be noted that the input data of each layer is the original data that has been recorded before, rather than the erroneous value output by the previous layer due to weight decomposition, so as to ensure the stability of the network;
[0057] Specifically, a sequential traversal is performed along the computing link of the current network layer until the first weight matrix is encountered; the matrix is compressed based on a preset screening compression rate, such as a 2-fold compression rate, and the value of the sub-weight matrix A*B is used to replace the original W value. At this time, the approximate value is erroneous; the weight matrix in the weight node is pre-decomposed, and the original weight value is replaced by the value of the decomposed matrix product A*B. For the weight node, it also needs to receive an input value, which is linearly calculated with the value of A*B, and then used as the output value of the weight node and input into the next node; the traversal continues using the output value of the weight node as the input value of the regular node; for each regular node that receives the value, the output value of the regular node and the pre-recorded value of the regular node are calculated. The error percentage of the original output value is calculated. If the error percentage is within the set sensitivity threshold, it can be considered that the weight decomposition of the previous order of the regular node does not significantly affect its output. At this time, the name and value of the previously decomposed weight matrix are recorded, and the output value of the regular node is replaced with the previously recorded original output value, and the traversal is continued according to the above steps; if the error percentage is greater than the set sensitivity threshold, the calculation link is continued to traverse until a regular node or the next weight node that meets the sensitivity threshold range is encountered; if the error percentage of the regular nodes in between does not meet the requirements, it means that the previously decomposed weight matrix does not meet the requirements, and there is no need to record it. The output value of the weight node is replaced with the recorded original output value, and the traversal is continued according to the above steps until all weight matrices in the layer have been evaluated;
[0058] d.Finally, following the above steps, the names and values of the decomposable weight matrices in all network layers can be obtained.
[0059] In the present invention, layer by layer means that each network layer in the large model is used as the basic unit, error statistics are performed on the conventional nodes in the layer, and the weight matrix in the weight node in the layer that meets the requirements is decomposed. The consideration of taking the layer as the basic unit is that the large model is generally copied and stacked by multiple identical network layers, so each layer generally has similar operations for processing data. Taking the layer as the unit can not only ensure the integrity of the data to a certain extent, but also achieve direct code reuse, thereby improving the deployment efficiency of the present invention.
[0060] The sensitivity threshold is set according to the percentage of error. Generally speaking, the smaller the error, the smaller the impact on the model accuracy. In a specific embodiment of the present invention, the sensitivity threshold can be set to 3%. Since the number of parameters of large models is huge and there is redundancy, it is believed that errors within the threshold range will not significantly affect the accuracy of the model.
[0061] The approximate operation refers to: performing low-rank decomposition on the weight matrix W in the weight node, decomposing it into two sub-weight matrices A and B, and using the product of the two sub-weight matrices A and B to replace the original weight matrix W, thereby realizing numerical approximate operation and reducing the number of parameters.
[0062] If necessary, the original weight matrix can also be decomposed into sub-weight matrices with a number greater than two, ensuring that the matrix dimension after the final multiplication of the sub-weight matrices is consistent with the original dimension; however, too many decomposed matrices will increase the complexity of the model structure, and it is necessary to screen out good low-rank values for each matrix, which reduces efficiency;
[0063] The present invention mainly decomposes the weight matrix in the model, replaces the original large matrix with two small matrices with fewer parameters to participate in multiplication and addition operations, and improves the model deployment performance. The key to the above problem is to find the optimal low-rank value of the matrix to be decomposed. The optimal low-rank value means: if it is lower than this value, the matrix is small, but the decomposition error is large; if it is higher than this value, the matrix is large. Although the decomposition error is small, the matrix can still be further compressed to save space. A common method is to customize the low-rank value and calculate the specific value according to the manually set compression rate. However, in the process of compression rate calculation, the unknown parameter redundancy will cause the low-rank value to be set unreasonably. Therefore, the above method cannot reasonably set the low-rank value, which is easy to cause the problem of excessive decomposition error or too low compression rate, and lacks flexibility. The present invention proposes a progressive low-rank setting mode based on decomposition error. Referring to the 8-bit quantization algorithm, it is defaulted that when the network parameters are compressed 4 times (32bit / 8bit=4), the model accuracy is lossless. Considering that the low-rank decomposition has a greater impact on the parameters than quantization, a 4-fold compression rate is used as the maximum compression amount of the matrix. First, set the current maximum compression amount to X, and divide X into an arithmetic progression containing N elements at equal intervals, for example: X h =4,X l=2, N=5, then the sequence is [4.0, 3.5, 3.0, 2.5, 2.0], where each element represents the compression multiple; then the corresponding low-rank value is calculated in order of compression ratio from large to small, and the obtained decomposable matrix is decomposed. If the error percentage of the weight matrix is less than the preset error percentage threshold J, it is considered that the weight matrix supports the current compression ratio, and the matrix name and low-rank value are recorded. Repeat the above steps until all matrices obtained in step 1 are recorded. In a specific embodiment, the preset error percentage threshold is set to 3%; since the optimal low-rank value is selected, the transition is from a large ratio to a small ratio. If a certain ratio meets the error percentage, the low-rank value calculated based on the current ratio is the optimal one. Finally, the optimal low-rank value corresponding to each decomposable matrix is obtained, which can be used to estimate the final parameter compression ratio of the large model. The above method is a progressive search mode, which makes it easier to find the optimal low-rank value of the matrix and guarantees the accuracy and performance of the model to a certain extent.
[0064] The 8-bit quantization algorithm is the current mainstream large model compression algorithm. Most models can convert parameters from fp32 format to int8 format, which reduces the parameter storage space while ensuring that the model accuracy is almost lossless. The present invention refers to this algorithm, and the model accuracy is lossless under the default 4-fold parameter compression condition; the method in the present invention reduces the model parameters, which is conducive to deployment to resource-constrained end-side devices, such as mobile phones, computers, etc., without other complicated steps and processes; while the existing 8-bit quantization algorithm only reduces the model size, and in actual operations, there are still problems such as the need to convert low-bit data into high-precision data for calculation, and it is necessary to introduce more numerical type conversion operations.
[0065] Calculating the compression ratio list requires determining three values: maximum compression ratio, minimum compression ratio, and ratio interval. Compression ratio calculation formula: Assume that the m*n dimensional matrix W needs to be low-rank decomposed into matrices A and B with dimensions of m*k and k*n, respectively. Use the result of A*B to approximate the original matrix W. At this time, the parameter changes from the previous m*n to the current (m*k+k*n), then the compression ratio = (m*n) / (m*k+k*n). When the compression ratio = 1, it means that the number of parameters has not decreased, and the risk of matrix decomposition affecting model accuracy is increased. Therefore, the compression ratio needs to be greater than 1. The present invention sets the maximum ratio to 4 and the minimum ratio to 1.5. N can be set arbitrarily, generally 3-5 is appropriate.
[0066] The calculation method of the low rank value corresponding to each compression factor is as follows: Under the condition of known compression ratio, calculate the decomposition error of the matrix, for example, compression ratio = P, then P = (m*n) / (m*k+k*n), m and n are the dimensions of the known matrix W, from which the size of the k value can be calculated, k = (m*n) / [P*(m+n)]. At this time, the low rank value k is known, and the mainstream low rank decomposition algorithm, such as SVD, can be used to calculate the values of matrices A and B. Then, the difference between W and A*B is used to calculate the decomposition error of the weight matrix. The decomposition error needs to be evaluated according to a new threshold, that is, the preset error percentage threshold mentioned above, which is similar to the sensitivity threshold set in step 1. The range of the preset error percentage threshold is generally set to 3%-5%.
[0067] The parameter compression rate estimation is to first calculate the size of all weight parameters in the original model. For example, if the dimension of W is m*n, then its parameter amount is m*n. The parameter values of all weight matrices are calculated and superimposed in the above manner, and the parameters that are not involved in the decomposition in the original model, such as the parameters in the Embedding layer, are added to get the original parameter amount. Then calculate the size of the weight parameters in the reconstructed model. For example, W is decomposed into matrices A and B with dimensions of m*k and k*n, then the parameter amount at this time is (m*k+k*n), and so on. Similarly, the parameters that are not involved in the decomposition in the original model, such as the parameters in the Embedding layer, are added to get the compressed model parameters. Finally, the original parameter amount is divided by the compressed parameter amount to get the final model compression rate.
[0068] Taking the LLaMa-7b model as an example, the model parameters were compressed by 1.21 times and the model accuracy loss was 5% when the method of the present invention was used on the MMLU data set. In actual application deployment, this loss is acceptable compared to the decrease in the number of parameters.
[0069] The above steps have obtained the matrices involved in the decomposition and the corresponding low-rank values. In the large model, the weight matrix is represented in the form of a linear layer, for example:
[0070] W = torch.nn.Linear(in_features, out_features, bias = False) / / Weight matrix W. In_features is the number of input features (P), out_features is the number of output features (Q), and bias = False means that this layer does not use a bias term. The dimension of the weight matrix W of this linear layer is P × Q.
[0071] If the matrix involved in the decomposition is W, the dimension is P*Q, and the low rank value is K, then W needs to be decomposed into two small matrices A and B, with dimensions of P*K and K*Q respectively. The two matrices participate in the matrix product operation to approximate the original matrix W and reduce the memory space occupied. Therefore, W is decomposed in the following way:
[0072] W_l = torch.nn.Linear(P, K, bias = False) / / Creates a new linear layer W_l, representing the decomposed matrix A. The number of input features of this layer is P, the number of output features is K, and no bias term is used. The dimension of the weight matrix of this layer (i.e., matrix A) is P×K.
[0073] W_r = torch.nn.Linear(K, Q, bias = False) / / Another new linear layer W_r is created, representing the decomposed matrix B. The number of input features of this layer is K, the number of output features is Q, and no bias term is used. The dimension of the weight matrix of this layer (i.e., matrix B) is K×Q.
[0074] W_l.weight=A
[0075] W_r.weight = B / / Set matrices A and B to the weights of the two decomposed linear layers W_l and W_r. A and B have been obtained through some matrix decomposition algorithm.
[0076] The decomposed network layers store their respective weight matrices A and B, and use the result of A*B to approximate the original W value. The above method reconstructs the large model into a low-rank structure, which can effectively reduce the storage space of the large model and improve the performance of the model during inference.
[0077] Specifically,
[0078] 1. The sensitivity analysis strategy based on calculating the links layer by layer can effectively screen out the weight matrix with low decomposition sensitivity:
[0079] The existing technology generally decomposes and approximates a single weight matrix to calculate the decomposition error. However, considering that there are many normalization and nonlinear operations in large models, the errors caused by decomposition may be eliminated. Therefore, even if the decomposition error of a node itself is large, it will not affect the subsequent output results. Taking the weight matrix W as an example, its dimension is P*Q. The low rank value is set to K. Then, the matrix decomposition method is used to decompose W into two matrices A and B with dimensions of P*K and K*Q respectively, and the product of matrix A and matrix B is used to approximate W. At this time, only P*K+K*Q=K*(P+Q) parameters are stored in the model. Compared with the previous P*Q parameters, the compression rate of the model = (P*Q) / (K*(P+Q)). If the K value is much smaller than P and Q, the compression ratio will take a larger value, but it will reduce the representation ability of the matrix to a certain extent. The error percentage calculation method for conventional nodes and weight matrices is as follows:
[0080]
[0081] m*n is the number of parameters of the weight matrix W;
[0082] Among them, in step one, the weight matrix is screened by calculating the error percentage of the output value of the conventional node; in step two, the decomposition error percentage of the weight matrix is calculated to obtain the optimal low-rank value of the matrix.
[0083] The elements in the above formula are calculated by position. For the Llama-7B model, it contains 32 self-attention layers, each layer contains 4 weight matrices, namely q_proj, k_proj, v_proj, o_proj, the dimensions of the above matrices are 4096*4096, if the compression ratio is set to 2, the low rank value is 1024. Taking q_proj as an example, if the q_proj matrix in the 32 layers is decomposed, the corresponding matrix decomposition error value in each layer is:
[0084]
[0085] The matrix decomposition error value is calculated as:
[0086]
[0087] This method can clearly show the error value before and after matrix decomposition:
[0088] Based on the 2-fold compression rate, the decomposition error values of most q matrices are within the range of 0.4-0.5, while the ideal decomposition error should be below 0.1. Further analysis found that even if the decomposition error value does not meet the requirements, considering that there are many normalization and nonlinear operations in the network layer, such as the softmax activation function, the numerical error caused by the decomposition of the pre-order weight matrix may be alleviated to a certain extent. Therefore, the present invention does not only consider the decomposition error of a single weight matrix, but is based on the output result of the conventional node calculation participated in by the approximate matrix of the matrix. In each layer in the network, the output value of the conventional node is recorded and analyzed, and the value is the output result after the approximate value of the weight matrix participates in the calculation. The error of the output result is analyzed to determine the weight that can finally participate in the decomposition. The analysis of each layer is independent, that is: the input value of each layer is the original data recorded, and these data come from the reasoning results of the original model, so as to prevent the error accumulation problem and ensure the stability and accuracy of each layer calculation. If the error percentage of the output value of the conventional node is within the set sensitivity threshold, then based on the current conventional node, the weight matrix of the nearest pre-decomposition and approximate replacement is traced back to the nearest pre-decomposition and approximate replacement, and the name of the matrix and the original value are added to the decomposable matrix list for subsequent work.
[0089] In the specific implementation process, the method of the present invention pays more attention to the large model of the Decoder-Only architecture, which is composed of multiple identical network layers stacked together. Each network layer includes a self_attn layer and an mlp layer. These two layers contain multiple weight matrices. The present invention focuses on analyzing the decomposability of these matrices.
[0090] In actual use, since there is often a lot of redundancy in the parameters of large models, the method of the present invention can obtain more significant benefits. It can reduce the model parameters as much as possible while ensuring the accuracy of the model, and can also be used for other large model products.
[0091] 2. The progressive low-rank setting mode based on decomposition error can effectively select the appropriate low-rank value, ensuring the accuracy of the model while reducing the model parameters as much as possible:
[0092] In the prior art, the low-rank value K of the weight matrix is generally set in advance according to the compression rate. For example, if the compression rate is set to R, the following formula needs to be satisfied:
[0093]
[0094] There is a problem with the above method. For any model, it is impossible to predict the redundancy of its parameters, so it is impossible to set a reasonable value in advance to ensure the final accuracy of the model and compress the parameters of the large model as much as possible. In response to the above problem, the present invention proposes a progressive low-rank setting mode based on decomposition error, which further searches for more suitable low-rank values among the selected weight matrices with low decomposition sensitivity. First, starting from a larger compression ratio, gradually reduce the compression ratio of the weight matrix. Based on a certain compression ratio, when the decomposition error percentage is less than the set threshold, record the name of the matrix and the corresponding low-rank value for subsequent model reconstruction. For example, the weight matrix G meets the conditions of the previous step and can participate in matrix decomposition; in this step, first set the error rate threshold to J, and the compression ratio list to [4.0, 3.5, 3.0, 2.5, 2.0]. Then, starting from a compression ratio of 4.0, the decomposition error percentage of the G matrix is calculated in sequence. If it is less than the threshold J, the current compression ratio is recorded, and the corresponding low-rank value is calculated and stored. If it is greater than the threshold J, the compression ratio list is continued to be traversed, and the above steps are repeated until the compression ratio of this stage is 2.0. At this time, it is equivalent to that there is no room for further compression of the G matrix, but the conditions of the previous step are met (the compression ratio of 2.0 is used in the previous step for preliminary matrix screening), indicating that it is the optimal compression ratio and no further compression is required.
[0095] 3. The matrix decomposition method for large model reasoning acceleration scenarios can effectively reduce the memory usage of large models and improve deployment efficiency:
[0096] Based on the above steps, the weight names that can participate in the decomposition and the corresponding optimal low-rank values have been obtained. After that, the current mainstream low-rank decomposition technology, such as singular value decomposition SVD, random singular value decomposition RSVD and other methods can be used to process the matrix, and the original large matrix is replaced with two small matrices for storage. At the same time, the structure of the large model is modified, and the original linear layer is replaced with two stacked linear layers to participate in matrix operations, so as to achieve the requirements of lightweight parameters, reduce the memory requirements of the large model, and speed up the efficiency of model deployment reasoning. The method proposed in the present invention does not require a training step, and performs parameter search and correction by simulating the input data of the real scene. It has a higher matching degree for tasks that cannot obtain original data and computing resources, and is more widely applicable.
[0097] Figure 1It represents the sensitivity analysis process of the present invention based on layer-by-layer link calculation. Among them, the circle represents the linear layer, and each linear layer contains a weight matrix (that is, the circle can also represent the weight node); the square represents the conventional node, which will perform linear or nonlinear operations on the input data and then output the result. First, the input data of the randomly generated simulated real scene is used to reason based on the original model to obtain the standard input and output values of each conventional node (using the original undecomposed weight matrix); then the approximate weight matrix (the product value of the decomposed submatrix) needs to be used to replace it when the model is traversed again. At the output position of each conventional node, it is necessary to compare the difference between the current result and the standard result and calculate the error rate. If the preset sensitivity threshold S is not met, the backward traversal is continued based on the current node; if the preset sensitivity threshold S is met, the weight matrix replaced by the approximate matrix is traced back to the nearest weight matrix, and the name and value of the matrix are recorded. In addition, the output result of the node needs to be replaced with the recorded original result to prevent error accumulation. This operation also needs to be set at the beginning of each network layer. N in the above figure indicates that the large model is stacked by similar structures. After the above steps, a list of decomposable weight matrices can be obtained. In addition, the present invention only decomposes the weights in the self_attn layer and the mlp layer, and does not consider the Embedding parameters. Taking the Llama-7B model as an example, the parameters in the self_attn layer and the mlp layer account for a large proportion, about 95% or more, so there is a greater compression benefit. The following list shows the proportion of parameters in each layer of the model:
[0098] Network Layer Weight parameter Weight ratio Embedding 32000*4096 0.019 Linear(attention) (4096*4096*4)*32 0.32 Linear(FFN) (4096*11008*3)*32 0.64 RMSNORM (attention norm) 4096 0 RMSNORM(FFNnorm) 4096 0 RMSNORM(outputnorm) 4096 0 linear 32000*4096 0.019
[0099] Figure 2 Indicates the progressive low-rank setting mode based on decomposition error of the present invention. For the weights in the decomposable matrix list obtained in the above steps, the compression rates in the compression rate list are used for compression in descending order, while paying attention to the decomposition error percentage of the current weight matrix. If the preset error percentage threshold J is not met, continue to traverse the compression ratio list until the traversal is completed; if the error value meets the preset error percentage threshold J, record the name of the matrix and the corresponding low-rank value for use in subsequent steps.
[0100] Figure 3 This represents the network reconstruction method based on linear layer replacement of the present invention. The original linear layer is decomposed into two small linear layers for stacking, and the model is reconstructed according to the optimal low-rank value obtained in the above steps; in addition, the original weight matrix is stored in a low-rank manner and can be directly loaded during model inference.
[0101] Example
[0102] 1. Get a list of decomposable matrices. The specific process is as follows: Figure 1As shown. First, set the sensitivity threshold S in advance to control the final accuracy of the model within a certain range. For the Llama-7B large model, set the sensitivity threshold to 3%-5%; since the large model is stacked by the same network layers, it is equivalent to copying the same network layer into multiple copies, so only a single network layer can be inserted into the error statistics module, and then the error collection of the entire network can be completed by copying. Then, error statistics are performed. First, a set of random tensors simulating real data is generated, the number is n, [X1, X2, X3, ... Xn], and they are batch-input into the model to complete the reasoning, and the input and output values of each regular node are recorded and saved as standard data; then the low-rank value of the weight matrix is calculated with a compression ratio of 2.0, so as to perform low-rank decomposition on the weights in the network, and the product result of the decomposed sub-matrix is used to replace the original weight matrix. The above process is based on layers. In order to prevent error accumulation, the input value of each layer is guaranteed to be standard data. When the error rate of the output result of a regular node is less than the sensitivity threshold S, the weight matrix of the nearest pre-decomposition and approximate replacement is traced back based on the node, and the name and value of the matrix are recorded. Then continue to traverse backwards and replace the output data of the regular node with the original standard data until the traversal of this layer is completed. Repeat the above steps until the self_attn layer and mlp layer of the large model are traversed.
[0103] 2. Obtain the optimal low-rank value of the decomposable matrix. The specific process is as follows Figure 2 As shown, the present invention sets the compression ratio interval and traverses the compression ratio table from high to low. If the current ratio meets the preset decomposition error percentage threshold J, the low rank value corresponding to the ratio is used as the optimal low rank value of the current weight matrix. The optimal low rank search is performed according to the information list obtained in the above steps. The information list is stored in the form of a dictionary, for example: info_dict = {"layers.0.self_attn.q_proj.weight": attn_weights}, where "layers.0.self_attn.q_proj.weight" is the weight name, and attn_weights is the weight node name. After that, the weight matrix in the list is compressed in descending order according to the compression ratio list, while paying attention to the error percentage of the weight decomposition. If the error rate is less than the preset error percentage threshold J, the current compression ratio is recorded, and the corresponding low rank value is calculated and added to the weight-low rank value list; if the error rate is greater than or equal to the preset error percentage threshold J, the compression ratio list continues to be traversed until the error percentage meets the requirements. The preset error percentage threshold J in this step and the sensitivity threshold S in the above step can be set to a uniform value or different. The weight-low rank value list is stored in the form of a dictionary, for example:
[0104] weight_rank_dict={"layers.0.self_attn.q_proj.weight":1024},
[0105] Where "layers.0.self_attn.q_proj.weight" is the weight name and 1024 is the low rank value.
[0106] 3. According to the weight-low rank value dictionary obtained in the above steps, the original large model structure is reconstructed. The specific process is as follows Figure 3 As shown. The original linear layer is replaced by two stacked linear layers, and the product of the weight values of the two layers is used to approximate the original weight matrix. At the same time, the low-rank values in the above list are used to reconstruct the network. Finally, subsequent reasoning and prediction are performed based on the reconstructed model structure and the weight values represented by the dimensionality reduction.
[0107] The present invention proposes a matrix decomposition method for large model reasoning acceleration scenarios, aiming to reduce model storage space, speed up model reasoning speed, and reduce model deployment costs. Based on the method described in the present invention, verification is performed on the Llama-7B model and WikiText dataset (the PPL value represents the model perplexity, the smaller the value, the better). The experimental results show that this method can effectively reduce model parameters without significantly affecting the model reasoning accuracy.
[0108] Model / Dataset Model parameters Compression Reasoning time Test accuracy PPL Llama-7B 7000559616 — 486.05s 5.37 Low rank decomposition model of the present invention 5781511883 1.21 411.37s 5.64
[0109] Comparative Example
[0110] Existing methods:
[0111] SVD is used to decompose the weight matrix in the large model, and the compression ratio of each weight matrix is uniformly set to 1.25. The results are verified based on the Llama-7B model and the WikiText dataset:
[0112] Model / Dataset Model parameters Compression Reasoning time Test accuracy PPL Llama-7B 7000559616 — 486.05s 5.37 Existing low-rank decomposition models 5475921920 1.27 398.57s 18.84 (complete loss of precision)
[0113] From the comparative examples, we can see that the existing low-rank decomposition model will pre-set the weight matrix to be decomposed and the corresponding compression ratio, and the effect is relatively poor. After decomposing the weight matrix in the Llama-7B model, the test accuracy will be significantly reduced, which significantly affects the model performance.
[0114] The method in the present invention is more flexible and efficient. By simulating a small part of the input data of the real scene, it is possible to obtain information such as whether the weight matrix can participate in the decomposition and the optimal low-rank value during the decomposition. It is more effective than manual random settings.
[0115] Various aspects of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products of the embodiments of the present invention. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0116] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0117] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0118] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to the embodiment of the present invention.In this regard, each square frame in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the module, program segment or instruction part contain one or more executable instructions for realizing the logical function of the specification.In some implementations as replacement, the function marked in the square frame can also occur in a sequence different from that marked in the accompanying drawings.For example, two continuous square frames can actually be executed substantially in parallel, and they can also be executed in the opposite order sometimes, depending on the function involved.It should also be noted that each square frame in the block diagram and / or flow chart, and the combination of the square frames in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0119] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the present invention, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the attached claims.
Claims
1. A matrix decomposition and reconstruction method for large model reasoning scenarios, characterized in that: The matrix decomposition and reconstruction method comprises the following steps: Step 1: Perform error statistics on conventional nodes in the network layer of the large model, and select weight nodes that can perform low-rank decomposition through sensitivity analysis strategy; Step 2: progressively perform low-rank setting, decompose the weight matrices in the weight nodes that can be decomposed by low-rank decomposition obtained by screening in step 1, and obtain and record the optimal low-rank values of the decomposable weight matrices; Step 3: Reconstruct the network of the original model based on the decomposable weight matrix and the corresponding low-rank values.
2. The matrix decomposition and reconstruction method according to claim 1, characterized in that: In step 1, when performing error statistics, each network layer in the large model is used as the basic unit and statistics are performed separately; The sensitivity analysis strategy refers to comparing the error of a regular node with a preset sensitivity threshold. If the error of the regular node is less than the sensitivity threshold, the weight matrix in the weight node that is closest to the regular node and uses an approximate matrix to participate in the calculation can be subjected to a low-rank decomposition operation; the error of the regular node comes from the numerical offset caused by the decomposition of its preceding weight node.
3. The matrix decomposition and reconstruction method according to claim 1, characterized in that: In step 2, a range of compression multiples of the weight matrix is set, and the range of compression multiples is divided into equal intervals to obtain one or more compression multiples; The obtained compression multiples are calculated in descending order to obtain low-rank values, the decomposable weight matrix is decomposed, and the weight matrix error percentage is calculated. When the error percentage of the weight matrix is less than the preset error percentage threshold, the weight matrix supports the current compression multiple, and the low-rank value corresponding to the weight matrix name and the current compression multiple is recorded. Repeat the above operation until the optimal low-rank value of all weight matrices is obtained.
4. The matrix decomposition and reconstruction method according to claim 3, characterized in that: In step 2, the calculation method of the error percentage of the weight matrix is as follows: Wherein, A and B are sub-weight matrices obtained by pre-decomposing the weight matrix W, and m and n are the dimensions of the weight matrix W; The calculation method of the low rank value corresponding to each compression factor is as follows: Assume compression factor = P, then P = (m*n) / (m*k+k*n), m and n are the dimensions of the known weight matrix W to be compressed, from which the size of the low rank value k can be calculated, k = (m*n) / [P*(m+n)].
5. The matrix decomposition and reconstruction method according to claim 1, characterized in that: After obtaining the optimal low-rank value of all weight matrices, the method also includes a step of estimating the final parameter compression rate of the large model, dividing the total number of parameters in the original model by the total number of parameters in the compressed and reconstructed model to obtain the final model compression rate.
6. The matrix decomposition and reconstruction method according to claim 1, characterized in that: In step three, the decomposed matrix corresponding to the optimal low-rank value replaces the original weight matrix to reconstruct the original model.
7. A system for implementing the matrix decomposition and reconstruction method according to any one of claims 1 to 6, characterized in that: The system comprises: an error statistics module, a low-rank decomposition module, a model reconstruction module, and a parameter compression rate estimation module; The error statistics module is used to perform error statistics on conventional nodes of the network layer in the large model, and to screen weight nodes capable of low-rank decomposition through a sensitivity analysis strategy; The low-rank decomposition module is used to perform low-rank decomposition on the weight matrix in the screened weight node, and obtain and record the optimal low-rank value of each decomposable weight matrix; The model reconstruction module is used to reconstruct the network of the original model based on the decomposable matrix and the corresponding low-rank value; The parameter compression rate estimation module is used to compare the total parameter amount in the original model with the total parameter amount in the compressed and reconstructed model to obtain the final model compression rate.
8. Application of the matrix decomposition and reconstruction method as described in any one of claims 1 to 6, or the system as described in claim 7 in large-model reasoning acceleration, large-model end-side device deployment, and large-model applications.
9. A hardware system for implementing the method according to any one of claims 1 to 6, characterized in that: The hardware system comprises: a memory and a processor; a computer program is stored in the memory, and when the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Response text generation method and device
CN121189294A
Answer text generation method and device
CN121189294B