Large language model compression method, device, equipment and storage medium

By clustering and residual decomposing the weights of large language models, combined with singular value decomposition and fine-tuning technology, the problems of low model compression efficiency and performance loss in existing technologies are solved, and efficient model compression and performance preservation are achieved.

CN120123803BActive Publication Date: 2025-09-23BEIJING WUWEN CORE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510612690.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-09-23
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

Existing large language model compression methods mainly focus on weight compression of a single linear layer, which is difficult to meet the high requirements of model storage and computing efficiency in practical applications, and there is a performance loss problem.

Method used

By clustering the weights of each linear layer of the large language model, calculating the residuals and decomposing them, compressing them using cluster centers and decomposition residuals, and combining singular value decomposition and fine-tuning techniques to optimize the model structure.

Benefits of technology

Significantly improve model compression efficiency, maintain model performance, reduce storage and computing costs, improve the computing and storage efficiency of computing devices, and enhance user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123803B_ABST
    Figure CN120123803B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of artificial intelligence technology, and in particular to a large language model compression method, apparatus, device and storage medium. The method comprises: clustering the weights of each linear layer of a large language model to obtain a plurality of cluster centers; for each of the weights, calculating the residual between the weight and a target cluster center, and decomposing the residual to obtain a decomposed residual, wherein the target cluster center is the cluster center closest to the weight; and compressing the large language model based on the plurality of cluster centers and the decomposed residual of each weight. The embodiment of the present disclosure achieves efficient compression of the weights of a large language model by applying clustering and residual processing to the weights of each linear layer, while maintaining the performance of the model as much as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a large language model compression method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of artificial intelligence (AI), large language models (LLMs) have achieved remarkable success in natural language processing tasks and are widely used in fields such as machine translation, text generation, and sentiment analysis. However, these models typically have billions or even hundreds of billions of parameters, resulting in significant storage and computational challenges in practical applications. To address these challenges, researchers have proposed a variety of model compression techniques, such as quantization, pruning, low-rank decomposition, and knowledge distillation.

[0003] Current compression methods primarily focus on compressing the weights of a single linear layer. For example, quantization reduces storage requirements by converting floating-point parameters into a low-bit-width representation, but this approach can introduce quantization errors, impacting model performance. Pruning reduces model complexity by removing unimportant weights, but excessive pruning can lead to performance degradation. Furthermore, while methods such as low-rank decomposition and knowledge distillation improve compression efficiency to a certain extent, they also have limitations. A rational and effective method for compressing large language models has yet to be established in the relevant art. Summary of the Invention

[0004] In view of this, the present disclosure proposes a large language model compression method, apparatus, device and storage medium, aiming to significantly reduce the number of parameters of the large language model while maintaining the performance of the model as much as possible through clustering and residual decomposition technology, save the hardware and software resources required for calculating and storing the large language model, and improve the computing and storage efficiency of computing devices deployed with large language models.

[0005] According to one aspect of the present disclosure, a large language model compression method is provided, the method comprising:

[0006] Cluster the weights of each linear layer of the large language model to obtain multiple cluster centers;

[0007] For each of the weights, calculating the residual between the weight and the target cluster center, and decomposing the residual to obtain a decomposed residual, wherein the target cluster center is the cluster center closest to the weight;

[0008] The large language model is compressed according to the multiple cluster centers and the decomposition residual of each weight.

[0009] In one possible implementation, clustering the weights of each linear layer of the large language model to obtain multiple cluster centers includes:

[0010] Combining weights of the same size among the weights of each linear layer of the large language model into a weight group, wherein the weights are a two-dimensional matrix, and the size includes the length and width of the two-dimensional matrix;

[0011] For each of the weight groups, rearrange each weight therein into a one-dimensional weight vector, and concatenate the weight vectors into a two-dimensional concatenated weight matrix;

[0012] For each of the weight groups, a preset clustering algorithm is used to cluster the splicing weight matrix to obtain a preset number of cluster centers, where the preset number is smaller than the number of weights in the weight group.

[0013] In another possible implementation, for each weight, calculating the residual between the weight and the target cluster center, and decomposing the residual to obtain a decomposed residual includes:

[0014] For each of the weights, calculating the distance between the weight vector of the weight and each cluster center in the multiple cluster centers, and determining the cluster center with the closest distance as the target cluster center;

[0015] Calculating a difference vector between the weight vector of the weights and the target cluster center as the residual;

[0016] Rearranging the residuals into a residual matrix of the same size as the weights;

[0017] The residual matrix is ​​decomposed based on a preset matrix processing method to obtain the decomposed residual.

[0018] In another possible implementation, the preset matrix processing method includes singular value decomposition (SVD), and decomposing the residual matrix based on the preset matrix processing method to obtain the decomposition residual includes:

[0019] Performing singular value decomposition on the residual matrix to obtain a plurality of singular values ​​and a plurality of left singular vectors and a plurality of right singular vectors corresponding to the plurality of singular values;

[0020] Selecting, in descending order, a preset number of singular values ​​and a corresponding preset number of left singular vectors and a preset number of right singular vectors from the plurality of singular values, wherein the preset number is smaller than the length and width of the weight;

[0021] Based on the preset number of singular values ​​and the corresponding preset number of left singular vectors and the preset number of right singular vectors, two decomposition residual matrices are constructed as the decomposition residuals.

[0022] In another possible implementation, compressing the large language model according to the multiple cluster centers and the decomposition residual of each weight includes:

[0023] For each of the weight groups, the decomposition residual of each weight and the preset number of cluster centers are used as the compressed weight group, thereby completing the compression of the large language model.

[0024] In another possible implementation, after calculating, for each weight, a residual between the weight and the target cluster center and decomposing the residual to obtain a decomposed residual, the method further includes:

[0025] For each of the weights, an approximate weight is calculated based on its corresponding target cluster center and decomposition residual;

[0026] Replacing each of the weights in the large language model with a corresponding approximate weight, and fine-tuning the large language model to fine-tune the two decomposition residual matrices of each weight;

[0027] The two decomposition residual matrices after fine-tuning are used as the decomposition residuals of the weights.

[0028] In another possible implementation, the clustering algorithm includes a K-Means clustering algorithm.

[0029] According to another aspect of the present disclosure, a large language model compression device is provided, the device comprising:

[0030] The clustering module is used to cluster the weights of each linear layer in the large language model to obtain multiple cluster centers;

[0031] a residual processing module, configured to calculate, for each weight, a residual between the weight and a target cluster center, and decompose the residual to obtain a decomposed residual, wherein the target cluster center is the cluster center closest to the weight;

[0032] A compression module is used to compress the large language model according to the multiple cluster centers and the decomposition residual of each weight.

[0033] According to another aspect of the present disclosure, a computing device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above method.

[0034] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented.

[0035] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program implements the above method when executed by a processor.

[0036] The disclosed embodiment provides a large language model compression method, which clusters the weights of each linear layer of the large language model to obtain multiple cluster centers. For each weight, the residual between the weight and the target cluster center closest to the weight is calculated, and the residual is decomposed to obtain a decomposed residual. Finally, the large language model is compressed based on the multiple cluster centers and the decomposed residual of each weight. Thus, the large language model is compressed by clustering and residual processing, which fully utilizes the structural consistency of the large language model. Compared with the method of compressing only a single linear layer weight in the related art, the model compression efficiency can be greatly improved. In addition, during the model compression process, by retaining the residual information, the model performance loss caused by compression can be effectively reduced, ensuring that the compressed model can still maintain a high performance in actual applications. Using this method to compress a large language model can significantly reduce the number of parameters of the large language model, save the hardware and software resources required for calculating and storing the large language model, and improve the computing and storage efficiency of the computing device deployed with the large language model.

[0037] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.

[0039] Figure 1 A flowchart of a large language model compression method provided by an exemplary embodiment of the present disclosure is shown.

[0040] Figure 2 A flowchart of a clustering method provided by an exemplary embodiment of the present disclosure is shown.

[0041] Figure 3 A flowchart of a residual processing method provided by an exemplary embodiment of the present disclosure is shown.

[0042] Figure 4 A flowchart of a model fine-tuning method provided by an exemplary embodiment of the present disclosure is shown.

[0043] Figure 5 A schematic structural diagram of a large language model compression device provided by an exemplary embodiment of the present disclosure is shown.

[0044] Figure 6 It is a block diagram of a device according to an exemplary embodiment. DETAILED DESCRIPTION

[0045] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0046] As used herein, the terms "comprises," "comprising," "having," or variations thereof are open ended and include one or more stated features, integers, elements, steps, parts, or functions, but do not preclude the presence or addition of one or more other features, integers, elements, steps, parts, functions, or groups thereof.

[0047] When an element is referred to as being "connected," "coupled," "responsive" or variations thereof to another element, it can be directly connected, coupled or responsive to the other element or intervening elements may be present.

[0048] Although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another element / operation. Therefore, without departing from the teachings of the present invention, the first element / operation in some embodiments may be referred to as the second element / operation in other embodiments.

[0049] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0050] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0051] Current large language model compression methods focus solely on weight compression of a single linear layer. Consequently, the compression ratio is limited, making it difficult to meet the high demands for model storage and computational efficiency in practical applications. To address this issue, the present disclosure proposes a large language model compression method based on clustering and residual processing. This method leverages the structural consistency of the large language model, significantly improving model compression efficiency while ensuring model performance.

[0052] The large language model compression method proposed in the embodiment of the present disclosure can be widely used in scenarios where a large language model needs to be deployed, especially on computing devices with limited resources. The computing device can be a terminal or a server. The terminal includes a mobile terminal or a fixed terminal, such as a mobile phone, a tablet computer, a laptop computer, a desktop computer, etc. The server can be a single server, or a server cluster consisting of several servers, or a cloud computing service center. By adopting the large language model compression method provided in the embodiment of the present disclosure, the storage space and computing resource consumption of the model can be significantly reduced, so that the computing device can efficiently run the large language model, and the compressed large language model can achieve fast inference under limited hardware resources, providing a smoother user experience.

[0053] The large language model itself has the ability to understand and generate human language, and its application areas are extremely wide, including but not limited to text summarization, translation, sentiment analysis, and dialogue systems. For example, in the field of text summarization, the large language model is used to automatically extract or generate the core content of the text, helping users to quickly obtain the key points of information. In the field of translation, the large language model can perform high-quality text translation, and is used to convert text in one language into text in another language while maintaining semantic accuracy and fluency. In the field of sentiment analysis, the large language model is used to identify and analyze emotional tendencies in text, and is widely used in market analysis, etc. The large language model in the field of dialogue systems is used to construct dialogue texts that can communicate naturally with human users, and is used in customer service, entertainment and other occasions. It should be emphasized that the embodiments of the present disclosure do not impose any restrictions on the specific application scenarios of the large language model. The universality and flexibility of its compression method enable it to meet the diverse needs in different fields and scenarios.

[0054] The core concept of the method provided by the embodiments of the present disclosure is to leverage the structural consistency of each layer of a large language model. Large language models typically have a highly consistent multi-layer structure, with each layer containing similar modules (such as self-attention modules and feedforward networks). This structural consistency leads to similar distributions of weights across different layers. Clustering algorithms can be used to group similar weights together, thus achieving efficient compression. Furthermore, the weight distribution of large language models often exhibits redundancy. Due to the large scale of the model, many weights are functionally similar or exhibit certain regularities in their values. By applying a preset clustering algorithm to the multiple weights of all linear layers, the cluster centers of each weight cluster can be efficiently obtained. Based on this, the original weights are assigned to the target cluster center with the closest distance to them, and the residuals between the weights and the assigned target cluster center are accurately calculated. Subsequently, the residuals are decomposed using a preset residual processing technique, further reducing storage requirements while preserving key information. This residual processing method effectively balances compression rate and model accuracy. This method not only significantly reduces the model's storage space requirements but also optimizes computational efficiency, providing strong support for the efficient deployment and fast inference of large language models in resource-constrained environments.

[0055] The large language model compression method provided by the embodiments of the present disclosure is introduced below using several exemplary embodiments.

[0056] Please refer to Figure 1 , which shows a flowchart of a large language model compression method provided by an exemplary embodiment of the present disclosure. This embodiment uses the method used in a computing device as an example. The method includes the following steps.

[0057] Step 101: cluster the weights of each linear layer of the large language model to obtain multiple cluster centers.

[0058] In large language models, weights are parameters of the neural network that define the relationship between input features and output. The weights of a linear layer determine how the input data is transformed at that layer. In this application, unless otherwise specified, weights refer to the set of all parameters of each linear layer (usually in the form of a two-dimensional matrix), not a single parameter.

[0059] Clustering is an unsupervised learning method used to divide data points into different clusters so that the similarity between data points in the same cluster is high, while the similarity between data points in different clusters is low.

[0060] The cluster center is the representative point of each cluster, which can usually be the mean or other statistics of all data points in the cluster. It is used to describe the central position or characteristics of the cluster.

[0061] In some embodiments, the computing device extracts the weights of each linear layer in the large language model. A preset clustering algorithm may be used to cluster the weights of all linear layers, such as through the group clustering method described below. During the clustering process, the weights are assigned to different clusters so that the weights within the same cluster are as similar as possible. After clustering, each cluster has a cluster center, which is a representative point of the weights. The cluster center can typically be calculated using the mean, median, or other statistical method of all weights within the cluster.

[0062] In the embodiments of the present disclosure, the choice of clustering algorithm is highly flexible. The K-means clustering algorithm is an optional solution, but the present solution is not limited to this algorithm. Other types of clustering techniques, as long as they can meet the requirements, can be considered and applied in the embodiments of the present disclosure.

[0063] Step 102: For each weight, calculate the residual between the weight and the target cluster center, and decompose the residual to obtain a decomposed residual, wherein the target cluster center is the cluster center closest to the weight.

[0064] For each weight, the nearest cluster center is called the target cluster center. The target cluster center is used to approximate the weight. The residual is used to indicate the difference between the weight and the nearest target cluster center, and is used to indicate the information lost in the weight clustering process.

[0065] In some embodiments, for each weight, the computing device calculates the difference between the weight and the nearest target cluster center as a residual. The residual indicates the difference between the weight and the target cluster center, that is, the residual reflects the degree of deviation between the weight and the target cluster center. The residual is decomposed using a preset residual processing method to obtain a decomposed residual. The residual processing method may include a matrix processing method or other residual processing methods.

[0066] The choice of residual processing method in the embodiments of the present disclosure also offers a high degree of flexibility. Singular value decomposition is an optional matrix processing method, but the present solution is not limited to this method. Other residual processing methods, as long as they meet the requirements, may be considered and applied in the embodiments of the present disclosure.

[0067] Step 103: compress the large language model based on the multiple cluster centers and the decomposition residuals of each weight.

[0068] During model inference, for each original weight, its target cluster center plus the corresponding decomposition residual is used to approximate the weight. This can reduce storage space and computational complexity while preserving the performance of the original model as much as possible.

[0069] After processing the weights of each linear layer in the above manner, the large language model is compressed to obtain a compressed large language model. This compressed large language model can achieve fast inference while maintaining high performance under limited hardware resources.

[0070] In summary, the embodiments of the present disclosure provide a large language model compression method, which brings the following beneficial effects: 1. Significantly improve compression efficiency: Through clustering and residual processing, the structural consistency of the large language model is fully utilized. Compared with the method of only compressing the weight of a single linear layer in the related art, the model compression efficiency can be greatly improved. 2. Maintain model performance: During the compression process, by retaining the residual information, the model performance loss caused by compression can be effectively reduced, ensuring that the compressed model can still maintain a high performance in actual applications. 3. Reduce storage and computing costs: The storage requirements of the model are greatly reduced, so that the large language model can run on devices with limited resources, which improves the computing and storage efficiency of computing devices deployed with large language models, reduces inference costs, and enhances user privacy and security. 4. Improve the scalability of the model: The compressed model is easier to deploy in different devices and scenarios, which improves the scalability and application scope of the model.

[0071] The following further introduces each step in the above embodiment.

[0072] In some embodiments, as Figure 2 As shown, the above step 101 can be replaced by the following steps:

[0073] In step 201 , weights of the same size in the weights of each linear layer of the large language model are grouped into a weight group, where the weights are a two-dimensional matrix, and the size includes the length and width of the two-dimensional matrix.

[0074] The computing device obtains the weights of each linear layer in the large language model to be processed, and divides the multiple weights of all linear layers into one or more weight groups based on their size. Each weight group includes multiple weights, and all weights in each weight group have the same size. The weights are two-dimensional matrices. The same size means that the length and width of the two-dimensional matrix are exactly the same. The length of the two-dimensional matrix generally refers to the number of rows in the two-dimensional matrix, and the width of the two-dimensional matrix generally refers to the number of columns in the two-dimensional matrix.

[0075] The number of weights in each weight group may be the same, or at least one weight group may be different from the other groups, which is not limited in the present embodiment.

[0076] The weight of each linear layer is a two-dimensional matrix, that is, the weight of each linear layer is also called the two-dimensional weight matrix of each linear layer. Usually, each linear layer corresponds to a two-dimensional weight matrix.

[0077] Schematically, in a large language model, the weights of a linear layer are represented as a two-dimensional matrix , where M is the length of the two-dimensional matrix and N is the width of the two-dimensional matrix. The weights of each linear layer in the large language model are divided into multiple weight groups according to their size, so that the size of the weights in each weight group remains consistent.

[0078] Step 202: for each weight group, rearrange each weight therein into a one-dimensional weight vector.

[0079] For each weight group, the computing device rearranges each weight therein into a one-dimensional weight vector.

[0080] The length of the rearranged one-dimensional weight vector is the product of the number of rows and columns of the original two-dimensional matrix. The rearrangement process can be performed as follows: each row of the original two-dimensional matrix is ​​sequentially taken out to obtain multiple row vectors equal to the number of rows. All row vectors are concatenated to form a one-dimensional weight vector. Alternatively, each column of the original two-dimensional matrix is ​​sequentially taken out to obtain multiple column vectors equal to the number of columns. All column vectors are concatenated to form a one-dimensional weight vector.

[0081] Schematically, taking one of the weight groups as an example, the weight group includes L weights in the form of a two-dimensional matrix of the same size , each weight Can be rearranged into a one-dimensional weight vector, denoted as a one-dimensional weight vector The length of the rearranged one-dimensional weight vector is MN, which is the product of the number of rows and columns of the original two-dimensional matrix. Therefore, an M×N two-dimensional weight matrix corresponds to a one-dimensional weight vector of length MN. For example, if there is a 3×4 two-dimensional weight matrix, the length of the one-dimensional weight vector after the two-dimensional weight matrix is ​​rearranged is 3×4=12.

[0082] Step 203: For each weight group, the rearranged weight vectors are concatenated into a two-dimensional concatenated weight matrix.

[0083] For each weight group, the one-dimensional weight vectors of each weight are re-arranged and concatenated to form a two-dimensional concatenated weight matrix. Schematically, each one-dimensional weight vector is concatenated as a row or column of the concatenated weight matrix (i.e., a new two-dimensional matrix).

[0084] Schematically, L one-dimensional weight vectors Perform splicing to obtain the splicing weight matrix .

[0085] Step 204: For each weight group, use a preset clustering algorithm to cluster the spliced weight matrix to obtain a preset number of cluster centers, where the preset number is less than the number of weights in the weight group.

[0086] For each weight group, cluster the spliced weights using a preset clustering algorithm. The weights in the large language model are compressed through the clustering algorithm. That is, the original weights in each weight group are flattened and spliced into a new two-dimensional matrix, namely the spliced weight matrix, and then the spliced weight matrix is clustered to obtain a preset number of cluster centers, where the preset number is less than the number of weights in the weight group, that is, the number of multiple cluster centers corresponding to each weight group is less than the number of weights in that weight group. These cluster centers are used to approximately represent the original weights, thereby achieving model compression.

[0087] Schematically, for the spliced weight matrix use a preset clustering algorithm (such as the K-means clustering algorithm) to cluster and obtain a set of cluster centers , where C < L, preferably, C << L. C is the number of multiple cluster centers corresponding to the weight group, and L is the number of weights in that weight group. This set of cluster centers can be understood as a set of C cluster centers (vectors) of length MN . For example, assume that a weight group contains 100 weights. Through clustering, 10 cluster centers can be obtained. These 10 cluster centers will represent these 100 weights, thereby achieving weight compression and simplification.

[0088] In some embodiments, as Figure 3 shown, the above step 102 can be replaced with the following steps:

[0089] Step 301: For each weight, determine the nearest cluster center as the target cluster center.

[0090] For each weight, after the computing device converts each weight into a one-dimensional weight vector, calculate the distance between each weight vector and each of the multiple cluster centers, and determine the nearest cluster center as the target cluster center of the weight vector. The cluster center can be the cluster center obtained by clustering the weight group where the weight vector is located. Optionally, the Euclidean distance formula can be used to calculate the distance between the weight vector and the cluster center. The embodiments of the present disclosure do not limit this.

[0091] Step 302: Calculate the residual between the weight and the target cluster center.

[0092] After converting each weight into a one-dimensional weight vector and determining the corresponding target cluster center, the computing device assigns each one-dimensional weight vector to the target cluster center. For each weight, the difference vector between each weight vector and the target cluster center is calculated as the residual, that is, the residual is a one-dimensional vector. Schematically, the one-dimensional weight vector is calculated by the following formula and the corresponding target cluster center The difference vector is used as the residual r:

[0093]

[0094] Step 303: rearrange the residuals into a residual matrix with the same size as the weights.

[0095] The computing device rearranges the residuals to obtain a residual matrix of the same size as the original two-dimensional weight matrix. Further, the multiple residuals are rearranged in rows or columns to form a two-dimensional residual matrix. For example, for the residual , the corresponding original two-dimensional weight matrix has a dimension of M×N, then the residual Arranged into an M×N two-dimensional residual matrix .

[0096] Step 304 : Decompose the residual matrix based on a preset matrix processing method to obtain a decomposed residual.

[0097] The computing device decomposes the residual matrix using a preset matrix processing method to obtain a decomposed residual representation, i.e., a decomposed residual. The matrix processing method includes a method for decomposing a matrix into a product of several low-rank matrices for dimensionality reduction, feature extraction, and data compression.

[0098] The preset matrix processing method may include singular value decomposition. Schematically, the residual matrix Take as an example, perform singular value decomposition on it, select the largest K singular values ​​and the corresponding K left singular vectors and K right singular vectors, and construct two decomposition residual matrices and As the decomposition residual, to approximate the residual matrix .

[0099] In some embodiments, decomposing the residual matrix based on a preset matrix processing method to obtain a decomposed residual may include the following steps:

[0100] 1. Perform singular value decomposition on the residual matrix to obtain multiple singular values ​​and multiple left singular vectors and multiple right singular vectors corresponding to the multiple singular values.

[0101] Schematically, the residual matrix Decomposed into the product of three matrices:

[0102]

[0103] Among them, U is the left singular vector matrix. U is an orthogonal matrix with size M×M, and its column vectors (with size M×1) are called left singular vectors, which are the left singular vectors of the residual matrix The number of all left singular vectors is equal to the number of rows M of the weights. Σ is the singular value diagonal matrix. Σ is a matrix with size M×N, and the elements on its diagonal are singular values. The singular values are non-negative real numbers, arranged in descending order, reflecting the "stretching" degree of the two-dimensional residual matrix in different directions. The number of singular values is equal to the rank of the residual matrix That is, the number of non-zero singular values. V is the right singular vector matrix, with size N×N, and its column vectors (with size N×1) are called right singular vectors, which are the right singular vectors of the residual matrix The number of all right singular vectors of the residual matrix is equal to the number of columns N of the weights.

[0104] 2. Select a preset number of singular values, the corresponding preset number of left singular vectors, and the corresponding preset number of right singular vectors from the multiple singular values in descending order. The preset number is less than the length and width of the weights.

[0105] Schematically, in descending order, select the largest K singular values from Σ, as well as the corresponding K left singular vectors (K column vectors of the left singular vector matrix) and K right singular vectors (K column vectors of the right singular vector matrix V). K is the preset number. Let be the diagonal matrix (with size K×K) of the retained K singular values, be the corresponding left singular vector matrix (with size M×K), be the corresponding right singular vector matrix (with size N×K). Among them, K < M and K < N. Preferably, K << M and K << N, that is, K is much smaller than the number of rows M of the residual matrix (which is also the number of rows M of the original two-dimensional weight matrix) and much smaller than the number of columns N of the residual matrix (which is also the number of columns N of the original two-dimensional weight matrix). Use the selected K singular values and the corresponding singular vectors to approximately represent the two-dimensional residual matrix r:

[0106]

[0107] 3. Based on the preset number of singular values, the corresponding preset number of left singular vectors, and the corresponding preset number of right singular vectors, construct two decomposed residual matrices as the decomposition residuals.​​​​​​​​​

[0108] Schematically, the Constructed as the decomposition residual matrix ,Will Constructed as At this time, the residual matrix It can be approximately expressed as:

[0109]

[0110] Among them, the decomposition residual matrix It is the product of the left singular vector matrix and the diagonal matrix of singular values, which is used to represent the features in the low-dimensional space, that is, ; Decomposition of the residual matrix is the transpose of the right singular vector matrix, which is used to remap the low-dimensional features back to the original space, i.e. .

[0111] Through the above steps, the computing device can efficiently decompose the residual matrix, thereby reducing storage and computational complexity while retaining key features of the original data.

[0112] In some embodiments, the above step 103 can be replaced by the following step: for each weight group, the decomposition residual of each weight and a preset number of cluster centers are used as the compressed weight group, thereby completing the compression of the large language model.

[0113] During the compression process of a large language model, for each weight group, the decomposition residuals of each weight in that weight group (such as the two decomposition residual matrices mentioned above) and a preset number of cluster centers (such as the C cluster centers mentioned above) are combined to form the compressed weight group, thus completing the compression of the large language model. This compression method can effectively reduce the model's storage requirements while preserving key model information, thereby achieving model lightweighting without significantly reducing model performance.

[0114] After the above steps, for the aforementioned L weights in the form of two-dimensional matrices of the same size , and cluster to get C cluster centers The weight group, its parameter quantity can be compressed from the original L×M×N to C×M×N+L×(K×M+K×N). Where L×M×N is L two-dimensional weight matrices The number of parameters, C×M×N is the C cluster centers The number of parameters, and L×(K×M+K×N) is the L decomposition residual matrix of the L two-dimensional weight matrix and L decomposition residual matrices The number of parameters. C and K are hyperparameters, and their selection is crucial for the compression effect and performance recovery of the model. C and K satisfy C < L, K < M, and K < N, and preferably, C ≪ L, K ≪ M, and K ≪ N.

[0115] As mentioned above, C is the number of cluster centers corresponding to the weight group. is the number of weights in the weight group. The reason for choosing C much smaller than L is that clustering can combine similar weights into a single cluster center, thereby reducing the number of model parameters. If C is close to or equal to L, the clustering effect is not significant and effective compression cannot be achieved.

[0116] K represents the preset number of singular values ​​selected in the matrix decomposition, while M and N represent the length (i.e., the number of rows in a two-dimensional matrix) and width (i.e., the number of columns in a two-dimensional matrix), respectively, of the weights. K is chosen to be significantly smaller than M and N because, by selecting the primary singular values, we can approximate the original weights while significantly reducing the number of parameters. If K is close to or equal to M or N, the compression effect is not significant, and effective parameter reduction cannot be achieved.

[0117] Therefore, choosing the hyperparameter settings of C≪L, K≪M and K≪N can ensure that the model still maintains key features after compression, while significantly reducing the number of parameters and improving the efficiency and scalability of the model.

[0118] In some embodiments, as Figure 4 As shown, after the above step 102, the method further includes the following steps:

[0119] Step 401 : For each weight, an approximate weight is calculated based on its corresponding target cluster center and decomposition residual.

[0120] Schematically, for each weight, first its target cluster center Rearrange into a two-dimensional matrix of M×N. Then, based on the target cluster center in the form of a two-dimensional matrix And the two decomposed residual matrices U and V are used to calculate the approximate weights by the following formula to approximate the original weights :

[0121]

[0122] In step 402 , each weight in the large language model is replaced with a corresponding approximate weight, and the large language model is fine-tuned to fine-tune the two decomposed residual matrices of each weight.

[0123] For example, when fine-tuning a large language model, a Low-Rank Adaptation (LoRA) technique can be used. The LoRA training mechanism involves fixing the target cluster center for each weight and fine-tuning the two decomposed residual matrices for each weight.

[0124] In model compression methods, model weights are clustered onto several "cluster centers." These cluster centers can be considered "representative points" for the weights, with each weight mapped to the nearest cluster center. Fixed target cluster centers mean that their values ​​remain unchanged during model fine-tuning. In other words, the cluster centers remain constant during training, serving as a fixed reference point for the model.

[0125] Furthermore, during model compression, after weights are mapped to cluster centers, residuals (i.e., the difference between the original weights and the target cluster centers) are generated. To preserve the important information in these residuals, they can be further compressed into low-rank matrices (such as the two decomposed residual matrices described above) or other forms. During model fine-tuning, gradient updates are performed only on these compressed residuals, rather than updating all weights.

[0126] The main purpose of this fine-tuning strategy is to maintain the model's lightweight by fixing the cluster centers and ensuring that the model's compression structure remains unchanged. By updating the residuals, important information lost in the original weights is retained, mitigating performance degradation caused by compression.

[0127] Schematically, for each weight, its corresponding target cluster center As W0, the two decomposition residual matrices U and V are B and A respectively, and the update formula can be expressed as:

[0128]

[0129] Where W0 is the pre-trained weight matrix, ΔW is the update amount, B and A are two matrices, and their product BA represents the weight update amount. During LoRA training, W0 is fixed, and only A and B are training parameters that are continuously optimized and adjusted.

[0130] By this method, it is possible to maintain W0 (target cluster center ) is stable, and by fine-tuning A and B (decomposing the residual matrices U and V), the performance degradation caused by approximate expression is alleviated, thereby achieving effective fine-tuning of large language models.

[0131] Step 403: Use the two fine-tuned decomposition residual matrices as weight decomposition residuals.

[0132] After fine-tuning, the two fine-tuned decomposition residual matrices are used as the weight decomposition residuals. This step lays the foundation for subsequent model compression. During the model compression process, these fine-tuned decomposition residuals are used to efficiently compress the large language model.

[0133] In summary, the method provided by the embodiment of the present disclosure, on the one hand, clusters weights of the same size, greatly reducing the number of parameters of the model. On the other hand, the residual between the original weight and the target cluster center is calculated, and then the residual is decomposed, which further reduces the number of parameters while retaining key feature information, thereby achieving a balance between model compression and network performance. The compressed model is more efficient in storage and computing, reduces memory usage and computing resource consumption, and is suitable for deployment on resource-constrained devices. On the other hand, by performing gradient updates on the approximate expression of the residual, i.e., the two decomposed residual matrices, during the fine-tuning stage, the performance loss caused by approximate compression can be effectively restored, ensuring that the model can still maintain high performance after compression. This solution achieves a good balance between model compression and network performance, achieving a significant reduction in the number of parameters and ensuring the performance of the model through fine-tuning optimization. It is suitable for the efficient deployment and optimization of large-scale deep learning models.

[0134] The following is an apparatus embodiment of the present disclosure. For parts not described in detail in the apparatus embodiment, reference may be made to the technical details disclosed in the above method embodiment.

[0135] Please refer to Figure 5 , which shows a schematic diagram of the structure of a large language model compression device provided by an exemplary embodiment of the present disclosure. The device can be implemented in whole or in part as a computing device using software, hardware, or a combination of both. The device includes a clustering module 51, a residual processing module 52, and a compression module 53.

[0136] A clustering module 51 is used to cluster the weights of each linear layer in the large language model to obtain multiple cluster centers;

[0137] The residual processing module 52 is used to calculate the residual between each weight and the target cluster center, and decompose the residual to obtain a decomposed residual, wherein the target cluster center is the cluster center closest to the weight;

[0138] The compression module 53 is used to compress the large language model according to the multiple cluster centers and the decomposition residual of each weight.

[0139] In a possible implementation, the clustering module 51 is further configured to:

[0140] The weights of each linear layer of the large language model with the same size are grouped into a weight group, where the weights are two-dimensional matrices, and the size includes the length and width of the two-dimensional matrix;

[0141] For each weight group, rearrange each weight into a one-dimensional weight vector, and concatenate the weight vectors into a two-dimensional concatenated weight matrix;

[0142] For each weight group, a preset clustering algorithm is used to cluster the spliced ​​weight matrix to obtain a preset number of cluster centers, which is smaller than the number of weights in the weight group.

[0143] In another possible implementation, the residual processing module 52 is further configured to:

[0144] For each weight, calculating the distance between the weight vector of the weight and each cluster center in the multiple cluster centers, and determining the cluster center with the closest distance as the target cluster center;

[0145] Calculate the difference vector between the weight vector of the weight and the target cluster center as the residual;

[0146] Rearrange the residuals into a residual matrix of the same size as the weights;

[0147] The residual matrix is ​​decomposed based on a preset matrix processing method to obtain the decomposed residual.

[0148] In another possible implementation, the preset matrix processing method includes singular value decomposition, and the residual processing module 52 is further used to:

[0149] Performing singular value decomposition on the residual matrix to obtain a plurality of singular values ​​and a plurality of left singular vectors and a plurality of right singular vectors corresponding to the plurality of singular values;

[0150] Selecting, in descending order, a preset number of singular values ​​and a corresponding preset number of left singular vectors and a preset number of right singular vectors from the plurality of singular values, where the preset number is less than the length and width of the weight;

[0151] Based on a preset number of singular values ​​and a corresponding preset number of left singular vectors and a preset number of right singular vectors, two decomposition residual matrices are constructed as decomposition residuals.

[0152] In another possible implementation, the compression module 53 is further configured to:

[0153] For each weight group, the decomposition residual of each weight and a preset number of cluster centers are used as the compressed weight group, thereby completing the compression of the large language model.

[0154] In another possible implementation, the apparatus further includes a fine-tuning module configured to:

[0155] For each weight, the approximate weight is calculated based on its corresponding target cluster center and decomposition residual;

[0156] Replace each weight in the large language model with the corresponding approximate weight and fine-tune the large language model to fine-tune the two factorized residual matrices for each weight;

[0157] The two decomposition residual matrices after fine-tuning are used as the decomposition residuals of the weights.

[0158] In another possible implementation, the clustering algorithm includes a K-means clustering algorithm.

[0159] It should be noted that, when the device provided in the above embodiment realizes its function, it only uses the division of the above-mentioned functional modules as an example. In actual application, the above-mentioned functions can be assigned to different functional modules according to actual needs, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0160] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0161] An embodiment of the present disclosure further provides a computing device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above method.

[0162] The embodiment of the present disclosure further provides a non-volatile computer-readable storage medium having a computer program stored thereon, and the computer program implements the above method when executed by a processor.

[0163] An embodiment of the present disclosure further provides a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program implements the above method when executed by a processor.

[0164] Figure 6 1 is a block diagram of an apparatus 1900 according to an exemplary embodiment. For example, the apparatus 1900 can be used to execute the above method and can be provided as a server or a terminal device. Figure 6The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.

[0165] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958 (I / O interface). The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2003. TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or similar.

[0166] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the apparatus 1900 to perform the above-described method.

[0167] A computer-readable storage medium can be a tangible device that can hold and store programs / instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or raised structure within a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted via wires.

[0168] The computer programs (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage medium in the respective computing / processing device for storage.

[0169] The computer program (or computer program instructions) used to perform the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, the state information of the computer-readable program instructions is used to personalize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), so that the electronic circuit can execute the computer-readable program instructions, thereby implementing various aspects of the present disclosure.

[0170] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0171] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0172] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0173] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0174] While various embodiments of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A large language model compression method, characterized in that: The method comprises: Clustering weights of linear layers of a large language model to obtain a plurality of cluster centers, wherein the large language model is used to understand and generate human language, and the weights are used to define a relationship between input features and output; For each of the weights, calculating the residual between the weight and the target cluster center, and decomposing the residual to obtain a decomposed residual, wherein the target cluster center is the cluster center closest to the weight; According to the plurality of cluster centers and the decomposition residuals of each weight, storing the decomposition residual of each weight and a preset number of cluster centers as a compressed weight group, thereby compressing the large language model; Deploying the compressed large language model on a computing device, so that the computing device runs the compressed large language model for reasoning, the computing device including a terminal or a server; The step of calculating the residual between each weight and the target cluster center and decomposing the residual to obtain the decomposed residual includes: For each of the weights, calculating the distance between the weight vector of the weight and each cluster center in the multiple cluster centers, and determining the cluster center with the closest distance as the target cluster center; Calculating a difference vector between the weight vector of the weights and the target cluster center as the residual; Rearranging the residuals into a residual matrix of the same size as the weights; The residual matrix is ​​decomposed based on a preset matrix processing method to obtain the decomposed residual.

2. The method according to claim 1, characterized in that The weights of each linear layer of the large language model are clustered to obtain multiple cluster centers, including: Combining weights of the same size among the weights of each linear layer of the large language model into a weight group, wherein the weights are a two-dimensional matrix, and the size includes the length and width of the two-dimensional matrix; For each of the weight groups, rearrange each weight therein into a one-dimensional weight vector, and concatenate the weight vectors into a two-dimensional concatenated weight matrix; For each of the weight groups, a preset clustering algorithm is used to cluster the splicing weight matrix to obtain a preset number of cluster centers, where the preset number is smaller than the number of weights in the weight group.

3. The method according to claim 2, characterized in that The preset matrix processing method includes singular value decomposition, and the decomposing the residual matrix based on the preset matrix processing method to obtain the decomposition residual includes: Performing singular value decomposition on the residual matrix to obtain a plurality of singular values ​​and a plurality of left singular vectors and a plurality of right singular vectors corresponding to the plurality of singular values; Selecting, in descending order, a preset number of singular values ​​and a corresponding preset number of left singular vectors and a preset number of right singular vectors from the plurality of singular values, wherein the preset number is smaller than the length and width of the weight; Based on the preset number of singular values ​​and the corresponding preset number of left singular vectors and the preset number of right singular vectors, two decomposition residual matrices are constructed as the decomposition residuals.

4. The method according to claim 3, characterized in that The compressing the large language model according to the multiple cluster centers and the decomposition residual of each weight includes: For each of the weight groups, the decomposition residual of each weight and the preset number of cluster centers are used as the compressed weight group, thereby completing the compression of the large language model.

5. The method according to claim 4, characterized in that After calculating the residual between each weight and the target cluster center and decomposing the residual to obtain the decomposed residual, the method further includes: For each of the weights, an approximate weight is calculated based on its corresponding target cluster center and decomposition residual; Replacing each of the weights in the large language model with a corresponding approximate weight, and fine-tuning the large language model to fine-tune the two decomposition residual matrices of each weight; The two decomposition residual matrices after fine-tuning are used as the decomposition residuals of the weights.

6. The method according to any one of claims 2 to 5, characterized in that The clustering algorithm includes a K-means clustering algorithm.

7. A large language model compression device, characterized in that: The device comprises: a clustering module for clustering the weights of each linear layer in a large language model used to understand and generate human language, and obtaining a plurality of cluster centers, wherein the large language model is used to understand and generate human language, and the weights are used to define the relationship between input features and output; a residual processing module, configured to calculate, for each weight, a residual between the weight and a target cluster center, and decompose the residual to obtain a decomposed residual, wherein the target cluster center is the cluster center closest to the weight; a compression module, configured to store the decomposition residual of each weight and a preset number of cluster centers as a compressed weight group based on the multiple cluster centers and the decomposition residual of each weight, thereby compressing the large language model; The apparatus deploys the compressed large prediction model on a computing device, so that the computing device runs the compressed large language model for reasoning, wherein the computing device includes a terminal or a server; Wherein, the residual processing module is further used for: For each of the weights, calculating the distance between the weight vector of the weight and each cluster center in the multiple cluster centers, and determining the cluster center with the closest distance as the target cluster center; Calculating a difference vector between the weight vector of the weights and the target cluster center as the residual; Rearranging the residuals into a residual matrix of the same size as the weights; The residual matrix is ​​decomposed based on a preset matrix processing method to obtain the decomposed residual.

8. A computing device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the method according to any one of claims 1 to 6.

9. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Media account recommendation method and device based on artificial intelligence and electronic equipment

    CN112861009A

  • Switch cabinet voiceprint fault detection method based on deep neural network

    CN117219124A