Model compression method and device, equipment, storage medium and product

By adjusting parameter distribution and orthogonal transformation of the attention layer weight matrix of the neural network, the problem of parameter redundancy and computing resource requirements in model compression is solved, and efficient model compression and performance improvement is achieved.

CN120278216APending Publication Date: 2025-07-08TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510344604.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

With the growth of artificial intelligence model parameters, the demand for computing resources and storage resources has increased significantly, and it is difficult for the existing technology to achieve high-quality model compression, resulting in poor hardware adaptability and complex pruning process.

Method used

By adjusting the parameter distribution and orthogonal transformation of the attention layer weight matrix of the neural network, dimensionality reduction processing is performed, a high-quality compression model is generated to maintain the output consistency and computational efficiency of the model.

Benefits of technology

It realizes high-quality compression of the model, reduces parameter redundancy, improves computing efficiency, reduces storage requirements, and simplifies the model retraining process, maintains the model's inference accuracy and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278216A_ABST
    Figure CN120278216A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model compression method and device, equipment, a storage medium and a product. The method comprises the steps that a weight matrix contained in a to-be-compressed model is acquired, the weight matrix is extracted from an attention layer in a neural network contained in the to-be-compressed model, and feature distribution adjustment is carried out on the weight matrix, the feature density of the adjusted matrix in the target matrix area is higher than the feature density of the weight matrix in the target matrix area, dimension reduction processing is carried out on the adjusted matrix to obtain a compressed matrix, and a compression model corresponding to the to-be-compressed model is generated based on the compressed matrix. Therefore, by performing feature distribution adjustment on the weight matrix, the distribution of the features in the target matrix region is more dense, and the distribution of the features in the non-target matrix region is more loose, so that the performance loss of the model in the dimension reduction processing process is reduced, and the high-quality compression of the model is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a model compression method, a model compression device, a computer device, a computer-readable storage medium, and a model compression product. Background Art

[0002] With the progress of scientific and technological research, artificial intelligence technology has developed rapidly, and more and more models have demonstrated excellent performance in natural language processing tasks. It has been found that as the performance increases, the parameters of the models also increase; for example, the parameters of large language models (LLMs) usually reach billions or more. The growth of model parameters poses extremely high requirements for both computing resources and storage resources. How to perform high-quality compression on model parameters has become a hot research issue currently. Summary of the Invention

[0003] Embodiments of this application provide a model compression method, device, equipment, computer-readable storage medium, and product, which can perform high-quality compression on a model.

[0004] On the one hand, embodiments of this application provide a model compression method, including:

[0005] Obtain the weight matrix of the model to be compressed, where the model to be compressed includes a neural network, and the weight matrix is extracted from the attention layer in the neural network;

[0006] Adjust the parameter distribution of the weight matrix to obtain an adjusted matrix, and the parameter density of the adjusted matrix in the target matrix region is higher than that of the weight matrix in the target matrix region;

[0007] Perform dimensionality reduction processing on the adjusted matrix to obtain a compressed matrix;

[0008] Generate a compressed model corresponding to the model to be compressed based on the compressed matrix.

[0009] On the one hand, embodiments of this application provide a model compression device, and the model compression device includes:

[0010] An obtaining unit, configured to obtain the weight matrix of the model to be compressed, where the model to be compressed includes a neural network, and the weight matrix is extracted from the attention layer in the neural network;

[0011] A processing unit, configured to adjust the feature distribution of the weight matrix to obtain an adjusted matrix, and the feature density of the adjusted matrix in the target matrix region is higher than that of the weight matrix in the target matrix region;

[0012] and for performing dimensionality reduction on the adjusted matrix to obtain a compressed matrix, where the dimension of the compressed matrix is smaller than the dimension of the weight matrix;

[0013] and for generating a compressed model corresponding to the model to be compressed based on the compressed matrix.

[0014] In one implementation, the processing unit is configured to adjust the parameter distribution of the weight matrix to obtain an adjusted matrix, specifically:

[0015] calculate the covariance of the weight matrix to obtain a first covariance matrix, where the first covariance matrix is used to indicate the correlation between any two parameters of the weight matrix;

[0016] determine an orthogonal transformation matrix corresponding to the weight matrix based on the correlation between the parameters in the weight matrix;

[0017] perform matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain the adjusted matrix.

[0018] In one implementation, the processing unit is configured to determine an orthogonal transformation matrix corresponding to the weight matrix based on the correlation between the parameters in the weight matrix, specifically:

[0019] perform eigenvalue decomposition on the first covariance matrix to obtain a first eigenvector matrix and a first eigenvalue matrix, where the first eigenvector matrix is used to indicate the change direction of each parameter in the weight matrix, and the first eigenvalue matrix is used to indicate the importance degree of each change direction;

[0020] determine the i-th column eigenvector in the first eigenvector matrix as the orthogonal transformation matrix corresponding to the weight matrix, where the eigenvalue corresponding to the i-th column eigenvector is greater than the eigenvalue threshold.

[0021] In one implementation, the processing unit is configured to adjust the parameter distribution of the weight matrix to obtain an adjusted matrix, specifically:

[0022] generate at least one candidate orthogonal transformation matrix based on the weight matrix;

[0023] determine a screening strategy for the candidate orthogonal transformation matrix according to the parameter distribution rule of the weight matrix;

[0024] screen at least one candidate orthogonal transformation matrix according to the screening strategy to obtain at least one orthogonal transformation matrix corresponding to the weight matrix;

[0025] perform matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain the adjusted matrix.

[0026] In one embodiment, the processing unit is configured to determine a screening strategy for a candidate orthogonal transformation matrix according to the parameter distribution law of the weight matrix, specifically:

[0027] If there is a preset distribution law whose similarity to the parameter distribution law of the weight matrix is greater than the similarity threshold, then determine the screening strategy associated with the preset distribution law as the screening strategy for the candidate orthogonal transformation matrix;

[0028] If there is no preset distribution law whose similarity to the parameter distribution law of the weight matrix is greater than the similarity threshold, then configure the screening strategy for the candidate orthogonal transformation matrix as a random selection strategy.

[0029] In one embodiment, the weight matrix includes an input parameter weight matrix and an output parameter weight matrix; the processing unit is configured to perform matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain an adjusted matrix, specifically:

[0030] Take the orthogonal transformation matrix as the left multiplication matrix and the input parameter weight matrix as the right multiplication matrix, and calculate to obtain the adjusted input parameter weight matrix;

[0031] Take the transposed matrix of the orthogonal transformation matrix as the right multiplication matrix and the output parameter weight matrix as the left weight matrix, and calculate to obtain the adjusted output parameter weight matrix.

[0032] In one embodiment, the processing unit is configured to perform dimensionality reduction processing on the adjusted matrix to obtain a compressed matrix, specifically:

[0033] Perform parameter contribution analysis on the adjusted matrix to obtain the contribution degree of each parameter in the adjusted matrix to the model to be compressed;

[0034] Based on the contribution degree of each parameter to the model to be compressed and the contribution degree threshold, perform parameter pruning on the adjusted matrix to obtain a compressed matrix; or, based on the contribution degree of each parameter to the model to be compressed and a preset ratio, perform parameter pruning on the adjusted matrix to obtain a compressed matrix.

[0035] In one embodiment, the processing unit is configured to perform parameter contribution analysis on the adjusted matrix to obtain the contribution degree of each parameter in the adjusted matrix to the model to be compressed, specifically:

[0036] Perform mean normalization processing on the adjusted matrix to obtain a matrix after mean normalization processing;

[0037] Calculate the second covariance matrix based on the matrix after mean normalization processing;

[0038] Perform eigenvalue decomposition on the second covariance matrix to obtain a second eigenvector matrix and a second eigenvalue matrix, where the second eigenvector matrix and the second eigenvalue matrix are used to indicate the contribution degrees of the respective parameters in the adjusted matrix in the model to be compressed.

[0039] In one implementation, the processing unit is configured to perform parameter pruning on the adjusted matrix based on the contribution degrees of the respective parameters in the model to be compressed and a contribution degree threshold to obtain a compressed matrix, specifically:

[0040] Remove the eigenvectors in the eigenvector matrix whose eigenvalues are less than the contribution degree threshold to obtain a compressed eigenvector matrix;

[0041] Perform matrix multiplication on the adjusted matrix and the compressed eigenvector matrix to obtain a compressed matrix.

[0042] In one implementation, the processing unit is further configured to:

[0043] Verify the output stability of the compressed model;

[0044] If the compressed model fails the output stability verification, then adjust the weight matrix using an updated parameter adjustment strategy, and generate a new compressed model based on the adjustment result; or, remove redundant parameters in the adjusted matrix using an updated parameter removal strategy, and generate a new compressed model based on the removal result; or, perform lightweight fine-tuning processing on the compressed model.

[0045] In one implementation, the processing unit is configured to verify the output stability of the compressed model, specifically:

[0046] Respectively call the model to be compressed and the compressed model to process M pieces of data to be processed, where M is a positive integer;

[0047] If the processing results of the jth piece of data to be processed are inconsistent, then verify the validity of the jth processing result of the compressed model, where j is a positive integer less than or equal to M;

[0048] Generate a stability verification result of the compressed model based on the number of identical processing results and the validity verification result of the jth processing result.

[0049] Correspondingly, the present application provides a computer device, which includes:

[0050] A memory, in which a computer program is stored;

[0051] A processor, configured to load the computer program to implement the above model compression method.

[0052] Accordingly, the present application provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor to perform the above model compression method.

[0053] Accordingly, the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above model compression method.

[0054] In an embodiment of the present application, a weight matrix included in a model to be compressed is obtained. The weight matrix is extracted from an attention layer in a neural network included in the model to be compressed. The parameter distribution of the weight matrix is adjusted so that the parameter density of the adjusted matrix in a target matrix region is higher than that of the weight matrix in the target matrix region. The adjusted matrix is dimension-reduced to obtain a compressed matrix, and a compressed model corresponding to the model to be compressed is generated based on the compressed matrix. It can be seen that by adjusting the parameter distribution of the weight matrix, the distribution of parameters can be made denser in the target matrix region and looser in the non-target matrix region, thereby reducing the performance loss during the dimension reduction process of the model and achieving high-quality compression of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1 FIG. is a scenario diagram of model compression provided by an embodiment of the present application;

[0057] Figure 2 FIG. is a flowchart of a model compression method provided by an embodiment of the present application;

[0058] Figure 3a FIG. is a schematic diagram of a Transformer network structure provided by an embodiment of the present application;

[0059] Figure 3b FIG. is a schematic diagram of a processing result provided by an embodiment of the present application;

[0060] Figure 3c FIG. is a schematic diagram of an orthogonal transformation provided by an embodiment of the present application;

[0061] Figure 3d Another schematic diagram of orthogonal transformation provided by an embodiment of the present application;

[0062] Figure 3e Yet another schematic diagram of orthogonal transformation provided by an embodiment of the present application;

[0063] Figure 4 A flowchart of another model compression method provided by an embodiment of the present application;

[0064] Figure 5a A schematic diagram of the principle of weight matrix compression provided by an embodiment of the present application;

[0065] Figure 5b A flowchart of matrix pruning provided by an embodiment of the present application;

[0066] Figure 5c An architecture diagram of the model compression process provided by an embodiment of the present application;

[0067] Figure 5d A schematic diagram of the compression effect provided by an embodiment of the present application;

[0068] Figure 6 A structural schematic diagram of a model compression device provided by an embodiment of the present application;

[0069] Figure 7 A structural schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0070] It should be noted in advance that, in order to enable those skilled in the art to better understand the technical solutions proposed in the embodiments of the present application, the embodiments of the present application will describe clearly and completely the implementation manners of the technical solutions proposed in the embodiments of the present application in combination with one or more attached drawings. And, each attached drawing shown in the embodiments of the present application is only for exemplary illustration. For example, the execution sequence of each step in the attached drawing can be adaptively adjusted according to the actual application scenario. In addition, in the embodiments of the present application, the block diagrams shown in each attached drawing are only functional entities, and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or, these functional entities can be implemented in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0071] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0072] It should be noted that: "a plurality of" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0073] The embodiments of the present application relate to technologies related to models (such as large language models). The following briefly introduces the relevant terms and concepts of models:

[0074] Transformer network: It is a deep learning model architecture and the basis of large language models. It is mainly used to process sequence data, especially in natural language processing tasks. The core innovation of the Transformer network is the self-attention mechanism, which allows the model to focus on information at different positions when processing the input sequence, thereby capturing long-range dependencies.

[0075] In the Transformer network, the structure of the weight matrix directly affects the inference effect and computational efficiency of the model. The model compression method provided by the present application mainly realizes model compression by adjusting the weight matrix (adjusting the parameter distribution and removing redundant parameters).

[0076] Large Language Model (LLM): It is a natural language processing model constructed by the Transformer architecture. This natural language processing model can understand and generate natural language, and perform various language tasks, such as text generation, translation, question answering, and dialogue, through training on a large amount of text data. The characteristic of the large language model is that it has a huge number of parameters, usually containing hundreds of millions to hundreds of billions of parameters, enabling it to have powerful language understanding and generation capabilities. In the embodiments of the present application, the large language model can be the model to be compressed.

[0077] Based on the above technologies related to models, the embodiments of the present application provide a model compression scheme, which can compress the model with high quality. Figure 1 A model compression scenario diagram provided for the embodiments of the present application, such asFigure 1 As shown in the figure, the model compression scenario provided by this application includes a terminal device 101 and a server 102. The model compression solution provided by this application can be executed by the terminal device 101 or the server 102. If the model to be compressed is a large language model or other model with a large number of parameters, the model compression solution provided by this application is usually executed by the server 102. It can be understood that when the model compression solution provided by this application is executed by the terminal device 101, the server 102 may not be included in the model compression scenario. Among them, the terminal device may include, but is not limited to: smart phones (such as Android phones, IOS phones, etc.), tablet computers, portable personal computers, mobile Internet devices (MID), intelligent voice interaction devices, intelligent household appliances, vehicle-mounted terminals, aircraft, wearable devices, etc. This application embodiment does not make any limitations in this regard; the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. This application embodiment does not make any limitations in this regard.

[0078] It should be noted that Figure 1 the numbers of the terminal device and the server in the figure are only for illustration and do not constitute an actual limitation of this application. The terminal device 101 and the server 102 can be connected by wired or wireless means, and this application does not make any restrictions in this regard.

[0079] The general process of the model compression solution provided by this application is as follows:

[0080] (1) The server 102 obtains the weight matrix of the model to be compressed; among them, the model to be compressed can be any artificial intelligence model that includes a neural network (such as a Transformer network), and the weight matrix is extracted from the attention layer in the neural network; for example, the model to be compressed can be a deep learning model or a large language model. The weight matrix can be separate (such as a separate input parameter weight matrix or output parameter weight matrix); it can also be correlated (such as an input parameter weight matrix and an output parameter weight matrix used in combination). The number of weight matrices can be one or more; for example, one weight matrix can be obtained by extracting the attention layer of a certain neural network in the model to be compressed, or multiple weight matrices can be obtained by extracting the attention layers of all neural networks in the model to be compressed. This application does not make any restrictions in this regard.

[0081] (2) The server 102 adjusts the parameter distribution of the weight matrix to obtain an adjusted matrix. The parameter density of the adjusted matrix in the target matrix region is higher than that of the weight matrix in the target matrix region. For example, assuming that both the weight matrix and the adjusted matrix are composed of a first matrix region and a second matrix region, after adjusting the weight matrix, the parameter density of the first matrix region decreases, and the parameter density of the second matrix region increases (i.e., the parameters in the matrix are aggregated towards the second matrix region). It should be noted that the parameters here specifically refer to valid parameters (such as non-zero parameters), or important parameters (such as parameters whose influence on the model to be compressed exceeds the influence threshold). The parameter distribution adjustment includes at least one of the following: adjusting the position of the parameters in the weight matrix, adjusting the values of the parameters in the weight matrix. The adjustment methods can include, but are not limited to: orthogonal transformation, permutation transformation. In one implementation, the server 102 performs an orthogonal transformation on the weight matrix to obtain an adjusted matrix.

[0082] (3) The server 102 performs dimensionality reduction on the adjusted matrix to obtain a compressed matrix. Dimensionality reduction means reducing the dimension of the matrix. For example, reducing a 5*5 matrix to a 3*3 matrix. The purpose of dimensionality reduction is to remove redundant parameters in the adjusted matrix. Redundant parameters can be understood as parameters in the adjusted matrix whose contribution degree (influence on the data processing process of the model to be compressed) is less than the contribution threshold. The contribution threshold can be a preset value or a dynamic value (such as determined based on a preset ratio and the contribution degree sorting result of all parameters in the matrix). The dimensionality reduction methods can include, but are not limited to: pruning, low-rank factorization. In one implementation, the server 102 performs a pruning process on the adjusted matrix to obtain a compressed matrix.

[0083] (4) The server 102 generates a compressed model corresponding to the model to be compressed based on the compressed matrix. In one implementation, the server 102 outputs model parameters based on the compressed matrix and replaces the parameters at the corresponding positions in the model to be compressed with the output model parameters to obtain a compressed model.

[0084] In an embodiment of the present application, a weight matrix included in a model to be compressed is obtained. The weight matrix is extracted from an attention layer in a neural network included in the model to be compressed. The parameter distribution of the weight matrix is adjusted so that the parameter density of the adjusted matrix in a target matrix region is higher than that of the weight matrix in the target matrix region. The adjusted matrix is dimensionally reduced to obtain a compressed matrix. Based on the compressed matrix, a compressed model corresponding to the model to be compressed is generated. It can be seen that by adjusting the parameter distribution of the weight matrix, the distribution of parameters can be made denser in the target matrix region and looser in the non-target matrix region, thereby reducing the performance loss during the dimensionality reduction process of the model and achieving high-quality compression of the model.

[0085] Based on the above model compression scheme, an embodiment of the present application proposes a more detailed model compression method. The model compression method proposed in the embodiment of the present application will be introduced in detail below with reference to the accompanying drawings.

[0086] Please refer to Figure 2 , Figure 2 which is a flowchart of a model compression method provided by an embodiment of the present application. The model compression method can be executed by a computer device; for example, it can be executed by the Figure 1 terminal device 101 or the server 102 shown in Figure 2 . As shown in

[0087] S201. Obtain a set of weight matrices of the model to be compressed.

[0088] The model to be compressed includes a plurality of Transformer networks. Each Transformer network includes an attention layer. The weight matrix is obtained by extracting the attention layer. The weight matrix may specifically include at least one of an input parameter weight matrix and an output parameter weight matrix. Figure 3a which is a schematic diagram of a Transformer network structure provided by an embodiment of the present application. As shown in Figure 3a , the Transformer network may include one or more attention layers. The attention layer of the Transformer network in (a) includes an input parameter weight matrix; the attention layer of the Transformer network in (b) includes an output parameter weight matrix; the attention layer of the Transformer network in (c) includes an input parameter weight matrix and an output parameter weight matrix.

[0089] In one implementation, the computer device extracts the weight matrices of the attention layers of all Transformer networks in the model to be compressed. In another implementation, the computer device extracts the weight matrices of the attention layers of some Transformer networks in the model to be compressed according to a weight matrix extraction strategy (such as a specified local area of the model to be compressed, or determining whether the weight matrix is a sparse matrix, etc.). It can be understood that the weight matrix extraction strategy can be dynamically adjusted based on performance requirements, compression ratio requirements, etc., and the present application does not limit this.

[0090] S202. Adjust the parameter distribution of the weight matrix to obtain an adjusted matrix.

[0091] The parameter density of the adjusted matrix in the target matrix area is higher than that of the weight matrix in the target matrix area. It should be noted that the processing results of the weight matrix and the adjusted matrix for the same input signal are the same; that is to say, the parameter distribution adjustment is performed on the premise of ensuring that the processing results of the attention layer (Transformer network) remain unchanged.

[0092] Figure 3b This is a schematic diagram of the processing result provided by the embodiment of the present application. As Figure 3b shown, the fact that the calculation results are the same means that: the output result A obtained by passing the input signal (input parameter) through the attention layer before adjustment in the Transformer network is the same as the output result B obtained by passing the input signal through the attention layer with the weight matrix adjusted in the Transformer network. The parameter distribution adjustment is used to adjust the position / distribution of the parameters in the weight matrix. After the parameter positions are adjusted, the parameter values can remain unchanged or can be changed. It can be seen that by using the feature of "the calculation results are the same", not only can the transparency of the impact of subsequent pruning operations on the model performance be improved, but also it provides convenience for subsequent model adjustment and evaluation. This method effectively eliminates the performance fluctuations caused by the change of the weight (matrix) during the model compression process, and lays a good foundation for the improvement of the inference speed and the optimization of computing resources. The adjustment methods can include but are not limited to: orthogonal transformation, permutation transformation. Taking orthogonal transformation as an example:

[0093] Orthogonal transformation is a linear transformation that keeps the vector length and angle unchanged, usually represented by an orthogonal matrix Q. The matrix Q that satisfies the orthogonality condition satisfies Q T Q = I; where I is the identity matrix. The basic properties of orthogonal transformation include: (1) The vector length remains unchanged: for any vector v, ||Qv|| = ||v|| (2) The angle remains unchanged: for any vectors v1, v2, <Qv1, Qv2> = <v1, v2>. These properties ensure that the characteristics and information content of the signal will not be changed when passing through the Transformer network.

[0094] If the attention layer of the Transformer network includes an input parameter weight matrix, the computer device obtains the orthogonal transformation matrix corresponding to the input parameter weight matrix. The input parameter weight matrix is adjusted by the orthogonal transformation matrix to obtain an adjusted input parameter weight matrix, and the transpose matrix of the orthogonal transformation matrix is configured as the output parameter matrix corresponding to the adjusted input parameter weight matrix. Figure 3c This is a schematic diagram of orthogonal transformation provided by an embodiment of the present application. As Figure 3c shown, the attention layer before adjustment includes an input parameter weight matrix, and after adjustment, the attention layer includes an adjusted input parameter weight matrix (obtained based on the input parameter weight matrix and the orthogonal transformation matrix) and an output parameter weight matrix (i.e., the transpose matrix of the orthogonal transformation matrix).

[0095] If the attention layer of the Transformer network includes an output parameter weight matrix, the computer device obtains the orthogonal transformation matrix corresponding to the output parameter weight matrix. The output parameter weight matrix is adjusted by the transpose matrix of the orthogonal transformation matrix to obtain an adjusted output parameter matrix, and the orthogonal transformation matrix is configured as the input parameter weight matrix corresponding to the adjusted output parameter matrix. Figure 3d This is another schematic diagram of orthogonal transformation provided by an embodiment of the present application. As Figure 3d shown, the attention layer before adjustment includes an output parameter weight matrix, and after adjustment, the attention layer includes an input parameter weight matrix (i.e., the orthogonal transformation matrix) and an adjusted output parameter weight matrix (obtained based on the output parameter weight matrix and the transpose matrix of the orthogonal transformation matrix).

[0096] If the attention layer of the Transformer network includes an input parameter weight matrix and an output parameter weight matrix, the computer device obtains the orthogonal transformation matrices corresponding to the input parameter weight matrix and the output parameter weight matrix. The input parameter weight matrix is adjusted by the orthogonal transformation matrix to obtain an adjusted input parameter weight matrix, and the output parameter weight matrix is adjusted by the transpose matrix of the orthogonal transformation matrix to obtain an adjusted output parameter matrix. Figure 3e This is yet another schematic diagram of orthogonal transformation provided by an embodiment of the present application. As Figure 3e shown, the attention layer before adjustment includes an input parameter weight matrix and an output parameter weight matrix, and after adjustment, the attention layer includes an adjusted input parameter weight matrix (obtained based on the input parameter weight matrix and the orthogonal transformation matrix) and an adjusted output parameter weight matrix (obtained based on the output parameter weight matrix and the transpose matrix of the orthogonal transformation matrix).

[0097] In one embodiment, the computer device calculates the covariance of the weight matrix to obtain a first covariance matrix, which is used to indicate the correlation between any two parameters in the weight matrix. The computer device determines the orthogonal transformation matrix corresponding to the weight matrix based on the correlation of each parameter in the weight matrix. In one embodiment, the computer device performs eigenvalue decomposition on the first covariance matrix to obtain a first eigenvector matrix and a first eigenvalue matrix; wherein, the first eigenvector matrix is used to indicate the change direction of each parameter in the weight matrix, and the first eigenvalue matrix is used to indicate the importance degree of each change direction. After obtaining the first eigenvector matrix and the first eigenvalue matrix, the computer device determines the i-th column eigenvector in the first eigenvector matrix as the orthogonal transformation matrix corresponding to the weight matrix; wherein, the eigenvalue corresponding to the i-th column eigenvector is greater than the eigenvalue threshold; for example, the i-th column eigenvector is the column eigenvector with the largest corresponding eigenvalue in the first eigenvector matrix. After determining the orthogonal transformation matrix corresponding to the weight matrix, the computer device performs matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain an adjusted matrix.

[0098] In another embodiment, the computer device generates at least one candidate orthogonal transformation matrix based on the weight matrix. Then, according to the parameter distribution law of the weight matrix, the computer device determines the screening strategy for the candidate orthogonal transformation matrix. In one embodiment, if there is a preset distribution law whose similarity to the parameter distribution law of the weight matrix is greater than the similarity threshold, the computer device determines the screening strategy associated with the preset distribution law as the screening strategy for the candidate orthogonal transformation matrix; correspondingly, if there is no preset distribution law whose similarity to the parameter distribution law of the weight matrix is greater than the similarity threshold, the computer device may configure the screening strategy for the candidate orthogonal transformation matrix as a random selection strategy. After determining the screening strategy for the candidate orthogonal transformation matrix, the computer device screens at least one candidate orthogonal transformation matrix according to the screening strategy to obtain the orthogonal transformation matrix corresponding to the weight matrix, and performs matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain an adjusted matrix.

[0099] S203. Perform dimensionality reduction processing on the adjusted matrix to obtain a compressed matrix.

[0100] Dimensionality reduction processing means reducing the dimension of the matrix; for example, reducing a 5*5 matrix to a 3*3 matrix. The purpose of dimensionality reduction processing is to remove redundant parameters in the adjusted matrix. Redundant parameters can be understood as parameters in the adjusted matrix whose contribution degree (the influence degree on the data processing process / effect of the model to be compressed) is less than the contribution degree threshold; the contribution degree threshold can be a preset value or a dynamic value. The removal methods of redundant parameters can include but are not limited to: pruning, low-rank factorization. Taking pruning as an example:

[0101] In one embodiment, the computer device performs parameter contribution analysis on the adjusted matrix to obtain the contribution degrees of the respective parameters in the adjusted matrix in the model to be compressed; among them, the parameter contribution analysis strategy can be preset or dynamically configured for different models (such as judging that there are differences in the used metrics and weights). After obtaining the contribution degrees of the respective parameters in the adjusted matrix in the model to be compressed, the computer device can perform parameter pruning on the adjusted matrix based on the contribution degrees of the respective parameters in the model to be compressed and the contribution degree threshold to obtain the compressed matrix; or it can perform parameter pruning on the adjusted matrix based on the contribution degrees of the respective parameters in the model to be compressed and a preset ratio to obtain the compressed matrix.

[0102] In another embodiment, the computer device can directly perform pruning on the adjusted matrix based on the threshold method to obtain the compressed matrix.

[0103] S204. Generate a compressed model corresponding to the model to be compressed based on the compressed matrix.

[0104] In one embodiment, the computer device outputs the model parameters corresponding to the model to be compressed based on the compressed matrix, and uses these model parameters to replace the parameters at the corresponding positions in the model to be compressed to obtain the compressed model; for example, assuming that the model parameters include the input parameter weight matrix and the output parameter weight matrix corresponding to the attention layer of the i-th Transformer network in the model to be compressed, then based on the input parameter weight matrix and the output parameter weight matrix, the original input parameter weight matrix and the original output parameter weight matrix in the attention layer of the i-th Transformer network in the model to be compressed are replaced, and the i-th Transformer network can be any one of the Transformer networks in the model to be compressed.

[0105] In the embodiments of the present application, the weight matrix included in the model to be compressed is obtained. The weight matrix is extracted from the attention layer in the neural network included in the model to be compressed. The parameter distribution of the weight matrix is adjusted so that the parameter density in the target matrix region of the adjusted matrix is higher than the parameter density in the target matrix region of the weight matrix. The adjusted matrix is subjected to dimensionality reduction processing to obtain the compressed matrix. A compressed model corresponding to the model to be compressed is generated based on the compressed matrix. It can be seen that by adjusting the parameter distribution of the weight matrix, the distribution of parameters can be made denser in the target matrix region and looser in the non-target matrix region, thereby reducing the performance loss of the model during the dimensionality reduction process and achieving high-quality compression of the model.

[0106] Please refer to Figure 4 , Figure 4The flowchart of another model compression method provided by an embodiment of this application, and this model compression method can be executed by a computer device; for example, executed by the Figure 1 terminal device 101 or server 102 shown in Figure 4 . As shown in

[0107] S401. Obtain the weight matrix of the model to be compressed.

[0108] For the specific implementation manner of step S401, reference can be made to the implementation manner of step S201 in Figure 2 , and details are not described herein again.

[0109] S402. Obtain the orthogonal transformation matrix corresponding to the weight matrix.

[0110] In one implementation manner, the computer device calculates the covariance of the weight matrix to obtain the first covariance matrix, and the first covariance matrix is used to indicate the correlation between any two parameters of the weight matrix; for example, in the covariance matrix, C i j is used to indicate the correlation between the i-th parameter and the j-th parameter in the weight matrix, and C i j . The larger the value, the greater the influence of the change of the i-th parameter on the j-th parameter, and if it is a preset value (such as 0), it means that the change of the i-th parameter does not affect the j-th parameter. Then, the computer device determines the orthogonal transformation matrix corresponding to the weight matrix based on the correlation between the parameters in the weight matrix.

[0111] In one embodiment, the computer device performs eigenvalue decomposition on the first covariance matrix to obtain the first eigenvector matrix and the first eigenvalue matrix; wherein, the first eigenvector matrix is used to indicate the change direction of each parameter in the weight matrix, and the first eigenvalue matrix is used to indicate the importance of each change direction. After obtaining the first eigenvector matrix and the first eigenvalue matrix, the computer device determines the i-th column eigenvector in the first eigenvector matrix as the orthogonal transformation matrix corresponding to the weight matrix, and the eigenvalue corresponding to the i-th column eigenvector is greater than the eigenvalue threshold.

[0112] In another implementation manner, if the Transformer network includes an input parameter matrix, the computer device determines at least one candidate orthogonal transformation matrix based on the input parameter matrix; if the Transformer network includes an output parameter matrix, the computer device determines at least one candidate orthogonal transformation matrix based on the output parameter matrix; if the Transformer network includes an input parameter matrix and an output parameter matrix, the computer device determines at least one candidate orthogonal transformation matrix based on the input parameter matrix or the output parameter matrix.

[0113] ​​After obtaining at least one candidate orthogonal transformation matrix, the computer device determines a screening strategy for the candidate orthogonal transformation matrix according to the parameter distribution law of the weight matrix. In one embodiment, if there is a preset distribution law whose similarity to the parameter distribution law of the weight matrix is greater than the similarity threshold, the computer device determines the screening strategy associated with the preset distribution law as the screening strategy for the candidate orthogonal transformation matrix; if there is no preset distribution law whose similarity to the parameter distribution law of the weight matrix is greater than the similarity threshold, the computer device configures the screening strategy for the candidate orthogonal transformation matrix as a random selection strategy. After determining the screening strategy for the candidate orthogonal transformation matrix, the computer device screens the candidate orthogonal transformation matrix according to the screening strategy of the orthogonal matrix to obtain the orthogonal transformation matrix corresponding to the weight matrix.

[0114] S403. Adjust the weight matrix through the orthogonal transformation matrix to obtain an adjusted weight matrix.

[0115] The computer device performs matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain an adjusted matrix.

[0116] In one implementation, the attention layer of the Transformer network includes an input parameter weight matrix. On the one hand, the computer device uses the orthogonal transformation matrix corresponding to the input parameter weight matrix as the left multiplication matrix and the input parameter weight matrix as the right multiplication matrix to calculate the adjusted input parameter weight matrix, which can be specifically expressed as:

[0117] W′ in =QW in

[0118] Where W′ in is the adjusted input parameter matrix, W in is the input parameter matrix before adjustment, and Q is the orthogonal transformation matrix corresponding to W in On the other hand, the computer device configures the transpose matrix of the orthogonal transformation matrix as the output parameter matrix corresponding to the adjusted input parameter weight matrix, that is, the output parameter matrix corresponding to W′ in is Q T .

[0119] In another implementation, the attention layer of the Transformer network includes an output parameter weight matrix. On the one hand, the computer device uses the output parameter weight matrix as the left multiplication matrix and the transpose matrix of the orthogonal transformation matrix corresponding to the output parameter weight matrix as the right multiplication matrix to calculate the adjusted output parameter matrix, which can be specifically expressed as:

[0120] W′ out =W out Q T

[0121] Where W′out is the adjusted output parameter matrix, W out is the output parameter matrix before adjustment, Q T is for W out The transpose matrix of the corresponding orthogonal transformation matrix. On the other hand, the computer device configures the orthogonal transformation matrix corresponding to the output parameter weight matrix as the input parameter matrix corresponding to the adjusted output parameter weight matrix, that is, W′ out The corresponding input parameter matrix is Q

[0122] In yet another embodiment, the attention layer of the Transformer network includes an input parameter weight matrix and an output parameter weight matrix. On the one hand, the computer device uses the orthogonal transformation matrix corresponding to the (input parameter / output parameter) weight matrix as the left multiplication matrix and the input parameter weight matrix as the right multiplication matrix to calculate the adjusted input parameter weight matrix. On the other hand, the computer device uses the output parameter weight matrix as the left multiplication matrix and the transpose matrix of the orthogonal transformation matrix corresponding to the (input parameter / output parameter) weight matrix as the right multiplication matrix to calculate the adjusted output parameter weight matrix. Specifically, it can be expressed as:

[0123] W′ in = QW in

[0124] W′ out = W out Q T

[0125] where, W′ in is the adjusted input parameter matrix, W in is the input parameter matrix before adjustment, W′ out is the adjusted output parameter matrix, W out is the output parameter matrix before adjustment, Q is for W in / W out The corresponding orthogonal transformation matrix, Q T is the transpose matrix of Q, W′ in and W′ out correspond to each other

[0126] Taking the attention layer including the input parameter weight matrix and the output parameter weight matrix as an example, the calculation process of the input signal X can be expressed as:

[0127] Y = W out (W in X)

[0128] where, Y is the output result of the input signal X in the attention layer before adjustment, W in is the input parameter matrix before adjustment, W out is the output parameter matrix before adjustment. The calculation process of the input signal X after orthogonal transformation can be expressed as:

[0129] Y' = W' out (W' in X) = W out Q T (QW in X) = W out W in X = Y

[0130] Among them, Y' is the output result of the input signal X in the adjusted attention layer, and W' in is the adjusted input parameter matrix, and W' out is the adjusted output parameter matrix. It can be seen from this equation that by performing an orthogonal transformation on the weight matrix in the attention layer, it can be ensured that the output result of the attention layer remains unchanged before and after the (weight) matrix adjustment, thereby ensuring the consistency of the model output.

[0131] S404. Perform parameter contribution analysis on the adjusted matrix to obtain the contribution degrees of each parameter in the adjusted matrix in the model to be compressed.

[0132] The contribution degree of any parameter in the model to be compressed is used to indicate the influence degree of this parameter on the data processing process / effect of the model to be compressed, and the contribution degree is proportional to the influence degree.

[0133] In one implementation, the computer device performs mean normalization (centralization) processing on the adjusted matrix to obtain the matrix after mean normalization (centralization) processing. The purpose of mean normalization (centralization) processing is to eliminate the offset between features, which can be specifically expressed as:

[0134]

[0135] Among them, X is the adjusted matrix, is the mean of each column in the adjusted matrix, is the matrix after mean normalization (centralization) processing.

[0136] Then the computer device calculates the second covariance matrix of the matrix after mean normalization processing, and its purpose is to describe the relationship between parameters. It can be specifically expressed as:

[0137]

[0138] Among them, C is the second covariance matrix, n is the number of samples, is the matrix after mean normalization (centralization) processing, is the transposed matrix of.

[0139] Perform eigenvalue decomposition on the second covariance matrix to obtain a second eigenvector matrix and a second eigenvalue matrix; wherein, the second eigenvector matrix and the second eigenvalue matrix are used to indicate the contribution degrees of the respective parameters in the adjusted matrix in the model to be compressed. Specifically, it can be expressed as:

[0140] C = VΛV T

[0141] wherein, V is the second eigenvector matrix (used to indicate the direction of the adjusted weight matrix in the new feature space), V T is the transpose matrix of V, and Λ is the second eigenvalue matrix (reflecting the variances of the principal components).

[0142] S405. Based on the contribution degrees of the respective parameters in the model to be compressed, perform parameter pruning on the adjusted matrix to obtain a compressed matrix.

[0143] In one implementation, the computer device performs parameter pruning on the adjusted matrix based on the contribution degrees of the respective parameters in the model to be compressed and a contribution degree threshold to obtain a compressed matrix. Specifically, the computer device (based on the second eigenvalue matrix) removes the eigenvectors in the second eigenvector matrix whose eigenvalues are less than the contribution degree threshold to obtain a compressed eigenvector matrix; specifically, it can be expressed as:

[0144] P = V k

[0145] wherein, P represents the compressed eigenvector matrix, and V k is obtained by removing the eigenvectors in the eigenvector matrix V whose eigenvalues are less than the contribution degree threshold. After obtaining the compressed eigenvector matrix, perform matrix multiplication calculation on the adjusted matrix and the compressed eigenvector matrix to obtain a compressed matrix, specifically, it can be expressed as:

[0146] W new = WP

[0147] wherein, W new is the compressed matrix, W is the adjusted matrix, and P is the compressed eigenvector matrix.

[0148] In another implementation, the computer device performs parameter pruning on the adjusted matrix based on the contribution degrees of the respective parameters in the model to be compressed and a preset ratio, to obtain a compressed matrix. Specifically, the computer device retains the top k eigenvectors in the second eigenvector matrix based on the eigenvalues of the respective eigenvectors, to obtain a compressed eigenvector matrix, or retains the top k rows / columns of eigenvectors based on the mean of the row / column eigenvectors; where k is determined based on the preset ratio and the number of eigenvectors in the second eigenvector matrix. After obtaining the compressed eigenvector matrix, calculate the product of the adjusted matrix and the compressed eigenvector matrix to obtain the compressed matrix. The specific process can refer to the previous implementation and will not be elaborated here.

[0149] It can be seen that through data dimensionality reduction and feature extraction, the key parameters (information) in the adjusted weight matrix can be effectively identified and retained, while the redundant parameters in the model are removed. It should be noted that since the parameter distribution of the weight matrix is adjusted in advance before pruning, the impact of the pruning process on the output quality of the Transformer network (model to be compressed) can be reduced (or even avoided).

[0150] Figure 5a This is a schematic diagram of the principle of weight matrix compression provided by an embodiment of this application. As Figure 5a shown, the gray blocks in the input parameter weight matrix and the output parameter weight matrix can be understood as important parameters, and the white blocks can be understood as redundant parameters. By performing orthogonal transformation on the input parameter weight matrix and the output parameter weight matrix, the parameter distribution in the input parameter weight matrix and the output parameter weight matrix can be adjusted (making the important parameters in one space of the weight matrix more dense, and the other space does not contain, or contains fewer important parameters). It should be noted that the orthogonal transformation matrix can be determined based on both the input parameter weight matrix and the output parameter weight matrix, or can be determined based on one of the input parameter weight matrix and the output parameter weight matrix; for example, if the input parameter weight matrix is a dense matrix and the output parameter weight matrix is a sparse matrix, the computer device determines the orthogonal transformation matrix based on the output parameter weight matrix. Further, after obtaining the adjusted input parameter matrix and the adjusted output parameter matrix (which are more convenient for structured pruning), the computer device performs pruning on them (removing the redundant parameters in the weight matrix) to obtain the compressed input parameter weight matrix and the compressed output parameter weight matrix.

[0151] It can be seen that the compressed matrix maintains a dense structure, can break through the computational bottleneck caused by the appearance of sparse matrices, and realizes the efficient operation of the model on existing hardware. This method is universal for different types of models and neural network layers and can be widely applied to neural networks of various sizes and structures.

[0152] Figure 5bA flow chart of matrix pruning provided by an embodiment of the present application. As Figure 5b shown, the pruning process includes: ① Obtain the adjusted (input parameter / output parameter) weight matrix. ② Perform mean normalization on the obtained weight matrix. ③ Calculate the covariance matrix based on the matrix obtained after mean normalization. ④ Perform eigenvalue decomposition on the covariance matrix to obtain the eigenvector matrix and the eigenvalue matrix. ⑤ Screen the eigenvector matrix according to the screening rule and the eigenvalue matrix (remove the redundant eigenvectors therein) to obtain the compressed eigenvector matrix (i.e., select the main components). ⑥ Calculate the compressed matrix based on the adjusted weight matrix and the compressed eigenvector matrix. Among them, the specific implementation manners of ① to ④ can refer to step S403 and step S404, and the specific implementation manners of ⑤ and ⑥ can refer to step S405, which will not be elaborated here.

[0153] It can be seen that in the above pruning process, by identifying and retaining key parameters, efficient pruning of the weight matrix is achieved.

[0154] S406. Generate a compressed model corresponding to the model to be compressed based on the compressed matrix.

[0155] The specific implementation manner of step S406 can refer to Figure 2 the implementation manner of step S204 in

[0156] Figure 5c A flow chart of a model compression architecture provided by an embodiment of the present application. As Figure 5c shown, first load the model to be compressed. The model to be compressed can be an initialized model or a trained model; for example, the model to be compressed can be a trained large language model. After loading the model to be compressed, obtain the weight matrix set of the model to be compressed (the main step of input preparation), and determine the pruning strategy (such as dense matrix pruning based on principal component analysis) and the pruning ratio (such as 10%). Then generate an orthogonal matrix based on the weight matrix set, and perform an orthogonal transformation on the corresponding weight matrix through the orthogonal matrix to obtain the adjusted weight matrix. Further, prune the adjusted weight matrix to obtain the compressed model, and then generate a compressed model corresponding to the model to be compressed. It can be seen that through this method, the computational complexity of the model to be compressed can be reduced, compression can be achieved, and the performance of the model can be maintained.

[0157] In one implementation, after generating the compressed model corresponding to the model to be compressed, the computer device can further verify the output stability of the compressed model. In one embodiment, the computer device processes M pieces of data to be processed by separately invoking the model to be compressed and the compressed model, where M is a positive integer. If the processing result of the j-th piece of data to be processed by the model to be compressed is inconsistent with the processing result of the j-th piece of data to be processed by the compressed model, the computer device can verify the validity of the j-th processing result of the compressed model, where j is a positive integer less than or equal to M. In one implementation, if the j-th processing result of the compressed model fails the validity verification and the j-th processing result of the model to be compressed passes the validity verification, the computer device determines that there is a performance loss.

[0158] Further, the computer device generates a stability verification result of the compressed model based on the number of identical processing results and the validity verification result of the i-th processing result. For example, assuming the quantity threshold is H, if the number of different processing results is G and G < H, the computer device determines that the compressed model passes the stability verification; if G > H and G - g < H (where g is the number of processing results that pass the validity verification among the different processing results obtained by the compressed model), the computer device determines that the compressed model passes the stability verification; if G > H and G - g ≥ H, the computer device determines that the compressed model fails the stability verification.

[0159] Furthermore, if the compressed model fails the output stability verification, the computer device can adjust the weight matrix using the updated parameter adjustment strategy and generate a new compressed model based on the adjustment result; or, the computer device can remove the redundant parameters in the adjusted matrix using the updated parameter removal strategy and generate a new compressed model based on the removal result; or, the computer device can perform lightweight fine-tuning processing on the compressed model to obtain a fine-tuned compressed model.

[0160] In another implementation, the computer device invokes the compressed model to process the data to be processed associated with the model to be compressed and obtains the processing result of the data to be processed. Specifically, the data to be processed can be news data, text interaction data (such as the text input by an object's question), business data of a certain industry (such as catering order data, travel capacity data), etc., and the model to be compressed can be a large language model, an artificial intelligence model, etc. It can be understood that the larger the scale of the model parameters, the more obvious the compression effect of the model compression method provided by this application. Figure 5d This is a schematic diagram of the compression effect provided by the embodiments of this application. As Figure 5dAs shown, the average processing time of the content of the compressed model obtained based on the model compression method provided in this application is significantly reduced, and the processing speed is greatly improved compared to the model before compression. Through the model compression method provided in this application, the computational complexity of the model to be compressed can be reduced, compression can be achieved, and the performance of the model can be maintained at the same time.

[0161] In practical applications, by using the model compression method provided in this application to compress the LLAMA-2 7B model (an open-source large language model), the number of parameters of this model can be reduced from 7 billion to 5.16 billion (the parameters are reduced by about 26%). In terms of storage, the storage requirements of this model can be effectively reduced. In terms of inference speed, the inference speed of the compressed LLAMA-2 7B model on a test Graphics Processing Unit (GPU) is increased by 39%, which is faster than the speed without pruning. This acceleration enables the compressed LLAMA-2 7B model to handle higher concurrent requests in practical applications and meet the requirements of fast response. In terms of model performance and accuracy, after compression (pruning), the compressed model can still maintain similar performance to the uncompressed model on many natural language tasks; for example, the performance of zero-shot drops by less than 10%. This result shows that after compression, important model parameters and input information are effectively retained, providing guarantee for the accuracy in practical applications. Generally speaking, the model compression method provided in this application utilizes the computational invariance of orthogonal transformation to change the parameter distribution and prune unimportant weights, so that the compressed model does not require large-scale recovery fine-tuning. This feature greatly simplifies the retraining process of the model and reduces the time and computational costs.

[0162] In the embodiments of this application, the weight matrix included in the model to be compressed is obtained. The weight matrix is extracted from the attention layer in the neural network included in the model to be compressed. The parameter distribution of the weight matrix is adjusted so that the parameter density of the adjusted matrix in the target matrix region is higher than the parameter density of the weight matrix in the target matrix region. The adjusted matrix is subjected to dimensionality reduction processing to obtain a compressed matrix. Based on the compressed matrix, a compressed model corresponding to the model to be compressed is generated. It can be seen that the model compression method provided in this application can efficiently compress models (such as large language models), while maintaining inference accuracy and improving computational efficiency, and solves problems such as performance loss, poor hardware adaptability, and complex pruning process existing in the compressed models in the prior art.

[0163] The method of the embodiments of this application is described in detail above. To facilitate better implementation of the above solutions of the embodiments of this application, correspondingly, the device of the embodiments of this application is provided below.

[0164] Please refer to Figure 6 , Figure 6Schematic structural diagram of a model compression device provided by an embodiment of the present application Figure 6 The shown model compression device can be mounted in a computer device, which can specifically be a terminal device or a server Figure 6 The shown model compression device can be used to execute some or all of the functions in the method embodiments described above Figure 2 and Figure 4 Please refer to Figure 6 The model compression device includes

[0165] An acquisition unit 601, configured to acquire a weight matrix of a model to be compressed, the model to be compressed includes a neural network, and the weight matrix is extracted from an attention layer in the neural network

[0166] A processing unit 602, configured to adjust the parameter distribution of the weight matrix to obtain an adjusted matrix, and the parameter density of the adjusted matrix in the target matrix region is higher than that of the weight matrix in the target matrix region

[0167] And configured to perform dimensionality reduction processing on the adjusted matrix to obtain a compressed matrix

[0168] And configured to generate a compressed model corresponding to the model to be compressed based on the compressed matrix

[0169] In one implementation, the processing unit 602 is configured to adjust the parameter distribution of the weight matrix to obtain an adjusted matrix, specifically configured to

[0170] Calculate the covariance of the weight matrix to obtain a first covariance matrix, and the first covariance matrix is used to indicate the correlation between any two parameters of the weight matrix

[0171] Determine an orthogonal transformation matrix corresponding to the weight matrix based on the correlation of each parameter in the weight matrix

[0172] Perform matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain an adjusted matrix

[0173] In one implementation, the processing unit 602 is configured to determine an orthogonal transformation matrix corresponding to the weight matrix based on the correlation of each parameter in the weight matrix, specifically configured to

[0174] Perform eigenvalue decomposition on the first covariance matrix to obtain a first eigenvector matrix and a first eigenvalue matrix, the first eigenvector matrix is used to indicate the change direction of each parameter in the weight matrix, and the first eigenvalue matrix is used to indicate the importance degree of each change direction

[0175] Determine the i-th column eigenvector in the first eigenvector matrix as the orthogonal transformation matrix corresponding to the weight matrix, where the eigenvalue corresponding to the i-th column eigenvector is greater than the eigenvalue threshold.

[0176] In one implementation, the processing unit 602 is configured to adjust the parameter distribution of the weight matrix to obtain an adjusted matrix, specifically:

[0177] Generate at least one candidate orthogonal transformation matrix based on the weight matrix;

[0178] Determine a screening strategy for the candidate orthogonal transformation matrix according to the parameter distribution law of the weight matrix;

[0179] Screen at least one candidate orthogonal transformation matrix according to the screening strategy to obtain an orthogonal transformation matrix corresponding to at least one weight matrix;

[0180] Perform matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain an adjusted matrix.

[0181] In one implementation, the processing unit 602 is configured to determine a screening strategy for the candidate orthogonal transformation matrix according to the parameter distribution law of the weight matrix, specifically:

[0182] If there is a preset distribution law whose similarity to the parameter distribution law of the weight matrix is greater than the similarity threshold, then determine the screening strategy associated with the preset distribution law as the screening strategy for the candidate orthogonal transformation matrix;

[0183] If there is no preset distribution law whose similarity to the parameter distribution law of the weight matrix is greater than the similarity threshold, then configure the screening strategy for the candidate orthogonal transformation matrix as a random selection strategy.

[0184] In one implementation, the weight matrix includes an input parameter weight matrix and an output parameter weight matrix; the processing unit 602 is configured to perform matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain an adjusted matrix, specifically:

[0185] Use the orthogonal transformation matrix as the left multiplication matrix and the input parameter weight matrix as the right multiplication matrix to calculate the adjusted input parameter weight matrix;

[0186] Use the transpose matrix of the orthogonal transformation matrix as the right multiplication matrix and the output parameter weight matrix as the left multiplication weight matrix to calculate the adjusted output parameter weight matrix.

[0187] In one implementation, the processing unit 602 is configured to perform dimensionality reduction processing on the adjusted matrix to obtain a compressed matrix, specifically:

[0188] Perform parameter contribution analysis on the adjusted matrix to obtain the contribution degrees of the respective parameters in the adjusted matrix to the model to be compressed;

[0189] Based on the contribution degrees of each parameter in the model to be compressed and the contribution degree threshold, parameter pruning is performed on the adjusted matrix to obtain the compressed matrix; or, based on the contribution degrees of each parameter in the model to be compressed and a preset ratio, parameter pruning is performed on the adjusted matrix to obtain the compressed matrix.

[0190] In one implementation, the processing unit 602 is configured to perform parameter contribution analysis on the adjusted matrix to obtain the contribution degrees of each parameter in the adjusted matrix in the model to be compressed, specifically:

[0191] Perform mean normalization processing on the adjusted matrix to obtain the matrix after mean normalization processing;

[0192] Calculate the second covariance matrix based on the matrix after mean normalization processing;

[0193] Perform eigenvalue decomposition on the second covariance matrix to obtain a second eigenvector matrix and a second eigenvalue matrix, where the second eigenvector matrix and the second eigenvalue matrix are used to indicate the contribution degrees of each parameter in the adjusted matrix in the model to be compressed.

[0194] In one implementation, the processing unit 602 is configured to perform parameter pruning on the adjusted matrix based on the contribution degrees of each parameter in the model to be compressed and the contribution degree threshold to obtain the compressed matrix, specifically:

[0195] Remove the eigenvectors in the eigenvector matrix whose eigenvalues are less than the contribution degree threshold to obtain the compressed eigenvector matrix;

[0196] Perform matrix multiplication calculation on the adjusted matrix and the compressed eigenvector matrix to obtain the compressed matrix.

[0197] In one implementation, the processing unit 602 is further configured to:

[0198] Verify the output stability of the compressed model;

[0199] If the compressed model fails the output stability verification, the weight matrix is adjusted using the updated parameter adjustment strategy, and a new compressed model is generated based on the adjustment result; or, the redundant parameters in the adjusted matrix are removed using the updated parameter removal strategy, and a new compressed model is generated based on the removal result; or, lightweight fine-tuning processing is performed on the compressed model.

[0200] In one implementation, the processing unit 602 is configured to verify the output stability of the compressed model, specifically:

[0201] Call the model to be compressed and the compression model respectively to process M pieces of data to be processed, where M is a positive integer;

[0202] If the processing results of the j-th piece of data to be processed are inconsistent, then perform validity verification on the j-th processing result of the compression model, where j is a positive integer less than or equal to M;

[0203] Generate a stability verification result of the compression model based on the number of identical processing results and the validity verification result of the j-th processing result.

[0204] According to an embodiment of the present application, Figure 2 and Figure 4 Some steps involved in the model compression method shown can be executed by each unit in the Figure 6 model compression device shown. For example, Figure 2 the step S201 shown in can be executed by the Figure 6 acquisition unit 601 shown, and the steps S202 - S204 can be executed by the Figure 6 processing unit 602 shown; Figure 4 the steps S401 and S402 shown in can be executed by the Figure 6 acquisition unit 601 shown, and the steps S403 - S406 can be executed by the Figure 6 processing unit 602 shown. Figure 6 Each unit in the model compression device shown can be separately or all combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In actual applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the model compression device can also include other units. In actual applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0205] According to another embodiment of the present application, it can be achieved by running a computer program (including program code) capable of executing the respective steps involved in the corresponding methods shown in Figure 2 and Figure 4 on a general computing device such as a computer device including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct as shown in Figure 6The model compression device shown in [description] is used to implement the model compression method of the embodiments of the present application. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.

[0206] Based on the same inventive concept, the principle of problem-solving and the beneficial effects of the model compression device provided in the embodiments of the present application are similar to those of the model compression method in the method embodiments of the present application. The principle and beneficial effects of the method implementation can be referred to. For the sake of brevity of description, they will not be elaborated here.

[0207] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a computer device provided in the embodiments of the present application. The computer device can be a terminal device or a server. As Figure 7 shown, the computer device at least includes a processor 701, a communication interface 702, and a memory 703. Among them, the processor 701, the communication interface 702, and the memory 703 can be connected through a bus or other means. Among them, the processor 701 (or the Central Processing Unit (CPU)) is the computing core and control core of the computer device. It can parse various instructions in the computer device and process various data of the computer device. For example, the CPU can be used to parse the power-on and power-off instructions sent by an object to the computer device and control the computer device to perform power-on and power-off operations. Another example is that the CPU can transmit various types of interaction data between the internal structures of the computer device, and so on. The communication interface 702 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), and can be controlled by the processor 701 to be used for sending and receiving data; the communication interface 702 can also be used for the transmission and interaction of internal data of the computer device. The memory 703 (Memory) is a memory device in the computer device, used to store programs and data. It can be understood that the memory 703 here can include both the built-in memory of the computer device and, of course, the extended memory supported by the computer device. The memory 703 provides a storage space, and the operating system of the computer device is stored in this storage space, which can include but is not limited to: Android system, iOS system, Windows Phone system, etc. The present application does not make any limitations in this regard.

[0208] The embodiments of the present application also provide a computer-readable storage medium (Memory). A computer-readable storage medium is a memory device in a computer device, used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device, and of course can also include the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the computer device. And, a computer program suitable for being loaded and executed by the processor 701 is also stored in this storage space. It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.

[0209] In one embodiment, the processor 701 performs the following operations by running the computer program in the memory 703:

[0210] Obtain the weight matrix of the model to be compressed. The model to be compressed includes a neural network, and the weight matrix is extracted from the attention layer in the neural network;

[0211] Adjust the parameter distribution of the weight matrix to obtain an adjusted matrix. The parameter density of the adjusted matrix in the target matrix region is higher than that of the weight matrix in the target matrix region;

[0212] Perform dimensionality reduction processing on the adjusted matrix to obtain a compressed matrix;

[0213] Generate a compressed model corresponding to the model to be compressed based on the compressed matrix.

[0214] As an alternative embodiment, the specific embodiment in which the processor 701 adjusts the parameter distribution of the weight matrix to obtain an adjusted matrix is:

[0215] Calculate the covariance of the weight matrix to obtain a first covariance matrix. The first covariance matrix is used to indicate the correlation between any two parameters of the weight matrix;

[0216] Determine the orthogonal transformation matrix corresponding to the weight matrix based on the correlation between the parameters in the weight matrix;

[0217] Perform matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain an adjusted matrix.

[0218] As an alternative embodiment, the specific embodiment in which the processor 701 determines the orthogonal transformation matrix corresponding to the weight matrix based on the correlation between the parameters in the weight matrix is:

[0219] Perform eigenvalue decomposition on the first covariance matrix to obtain a first eigenvector matrix and a first eigenvalue matrix. The first eigenvector matrix is used to indicate the change directions of the various parameters in the weight matrix, and the first eigenvalue matrix is used to indicate the importance degree of each change direction.

[0220] Determine the i-th column eigenvector in the first eigenvector matrix as the orthogonal transformation matrix corresponding to the weight matrix, where the eigenvalue corresponding to the i-th column eigenvector is greater than the eigenvalue threshold.

[0221] As an alternative embodiment, a specific embodiment in which the processor 701 adjusts the parameter distribution of the weight matrix to obtain an adjusted matrix is as follows:

[0222] Generate at least one candidate orthogonal transformation matrix based on the weight matrix;

[0223] Determine a screening strategy for the candidate orthogonal transformation matrices according to the parameter distribution law of the weight matrix;

[0224] Screen at least one candidate orthogonal transformation matrix according to the screening strategy to obtain an orthogonal transformation matrix corresponding to at least one weight matrix;

[0225] Perform matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain an adjusted matrix.

[0226] As an alternative embodiment, a specific embodiment in which the processor 701 determines a screening strategy for the candidate orthogonal transformation matrices according to the parameter distribution law of the weight matrix is as follows:

[0227] If there is a preset distribution law whose similarity to the parameter distribution law of the weight matrix is greater than the similarity threshold, then determine the screening strategy associated with the preset distribution law as the screening strategy for the candidate orthogonal transformation matrices;

[0228] If there is no preset distribution law whose similarity to the parameter distribution law of the weight matrix is greater than the similarity threshold, then configure the screening strategy for the candidate orthogonal transformation matrices as a random selection strategy.

[0229] As an alternative embodiment, the weight matrix includes an input parameter weight matrix and an output parameter weight matrix; a specific embodiment in which the processor 701 performs matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain an adjusted matrix is as follows:

[0230] Use the orthogonal transformation matrix as the left multiplication matrix and the input parameter weight matrix as the right multiplication matrix to calculate an adjusted input parameter weight matrix;

[0231] Use the transpose matrix of the orthogonal transformation matrix as the right multiplication matrix and the output parameter weight matrix as the left multiplication weight matrix to calculate an adjusted output parameter weight matrix.

[0232] As an alternative embodiment, a specific example of the processor 701 performing dimensionality reduction on the adjusted matrix to obtain a compressed matrix is as follows:

[0233] Perform parameter contribution analysis on the adjusted matrix to obtain the contribution degrees of the respective parameters in the adjusted matrix in the model to be compressed;

[0234] Based on the contribution degrees of the respective parameters in the model to be compressed and a contribution degree threshold, perform parameter pruning on the adjusted matrix to obtain a compressed matrix; or, based on the contribution degrees of the respective parameters in the model to be compressed and a preset ratio, perform parameter pruning on the adjusted matrix to obtain a compressed matrix.

[0235] As an alternative embodiment, a specific example of the processor 701 performing parameter contribution analysis on the adjusted matrix to obtain the contribution degrees of the respective parameters in the adjusted matrix in the model to be compressed is as follows:

[0236] Perform mean normalization on the adjusted matrix to obtain a matrix after mean normalization;

[0237] Calculate a second covariance matrix based on the matrix after mean normalization;

[0238] Perform eigenvalue decomposition on the second covariance matrix to obtain a second eigenvector matrix and a second eigenvalue matrix, where the second eigenvector matrix and the second eigenvalue matrix are used to indicate the contribution degrees of the respective parameters in the adjusted matrix in the model to be compressed.

[0239] As an alternative embodiment, a specific example of the processor 701 performing parameter pruning on the adjusted matrix based on the contribution degrees of the respective parameters in the model to be compressed and a contribution degree threshold to obtain a compressed matrix is as follows:

[0240] Remove the eigenvectors in the eigenvector matrix whose eigenvalues are less than the contribution degree threshold to obtain a compressed eigenvector matrix;

[0241] Perform matrix multiplication calculation on the adjusted matrix and the compressed eigenvector matrix to obtain a compressed matrix.

[0242] As an alternative embodiment, the component instances of the N components are cached in the software development kit; the processor 701 further performs the following operations by running the computer program in the memory 703:

[0243] Verify the output stability of the compressed model;

[0244] If the compressed model fails the output stability verification, the weight matrix is adjusted using the updated parameter adjustment strategy, and based on the adjustment result, a new compressed model is generated; or, the redundant parameters in the adjusted matrix are removed using the updated parameter removal strategy, and based on the removal result, a new compressed model is generated; or, the compressed model is subjected to lightweight fine-tuning.

[0245] As an alternative embodiment, a specific embodiment of the processor 701 for performing output stability verification on the compressed model is as follows:

[0246] The to-be-compressed model and the compressed model are respectively called to process M pieces of to-be-processed data, where M is a positive integer;

[0247] If the processing results of the j-th piece of to-be-processed data are inconsistent, the validity of the j-th processing result of the compressed model is verified, where j is a positive integer less than or equal to M;

[0248] Based on the number of the same processing results and the validity verification result of the j-th processing result, the stability verification result of the compressed model is generated.

[0249] Based on the same inventive concept, the principle of problem-solving and the beneficial effects of the computer device provided in the embodiments of the present application are similar to those of the model compression method in the method embodiments of the present application. The principle of problem-solving and the beneficial effects of the method can be referred to. For the sake of concise description, they will not be elaborated here.

[0250] The embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored, and the computer program is adapted to be loaded and executed by a processor to perform the model compression method in the above method embodiments.

[0251] The embodiments of the present application further provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above model compression method.

[0252] The steps in the method embodiments of the present application can be adjusted, combined, and deleted according to actual needs.

[0253] The modules in the device embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0254] In the embodiments of the present application, the "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0255] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The readable storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disc, etc.

[0256] The above disclosure is only a preferred embodiment of the present application. Of course, it cannot be used to limit the scope of rights of the present application. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the application.

Claims

1. A model compression method, characterized in that, The method includes: Obtaining a weight matrix of a model to be compressed, where the model to be compressed includes a neural network, and the weight matrix is extracted from an attention layer in the neural network; Adjusting the parameter distribution of the weight matrix to obtain an adjusted matrix, where the parameter density of the adjusted matrix in a target matrix region is higher than that of the weight matrix in the target matrix region; Performing dimensionality reduction processing on the adjusted matrix to obtain a compressed matrix; Generating a compressed model corresponding to the model to be compressed based on the compressed matrix.

2. The method according to claim 1, characterized in that, The adjusting the parameter distribution of the weight matrix to obtain an adjusted matrix includes: Calculating the covariance of the weight matrix to obtain a first covariance matrix, where the first covariance matrix is used to indicate the correlation between any two parameters of the weight matrix; Determining an orthogonal transformation matrix corresponding to the weight matrix based on the correlation between the parameters in the weight matrix; Performing matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain an adjusted matrix.

3. The method according to claim 2, wherein The determining an orthogonal transformation matrix corresponding to the weight matrix based on the correlation between the parameters in the weight matrix includes: Performing eigenvalue decomposition on the first covariance matrix to obtain a first eigenvector matrix and a first eigenvalue matrix, where the first eigenvector matrix is used to indicate the change directions of the parameters in the weight matrix, and the first eigenvalue matrix is used to indicate the importance degree of each change direction; Determining the i-th column eigenvector in the first eigenvector matrix as the orthogonal transformation matrix corresponding to the weight matrix, where the eigenvalue corresponding to the i-th column eigenvector is greater than an eigenvalue threshold.

4. The method according to claim 1, wherein The adjusting the parameter distribution of the weight matrix to obtain an adjusted matrix includes: Generating at least one candidate orthogonal transformation matrix based on the weight matrix; Determining a screening strategy for the candidate orthogonal transformation matrix according to the parameter distribution rule of the weight matrix; Screening the at least one candidate orthogonal transformation matrix according to the screening strategy to obtain the orthogonal transformation matrix corresponding to the at least one weight matrix; Performing matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain an adjusted matrix.

5. The method according to claim 4, wherein The determining a screening strategy for the candidate orthogonal transformation matrix according to the parameter distribution rule of the weight matrix includes: If there exists a preset distribution rule whose similarity to the parameter distribution rule of the weight matrix is greater than a similarity threshold, then determining the screening strategy associated with the preset distribution rule as the screening strategy for the candidate orthogonal transformation matrix; If there does not exist a preset distribution rule whose similarity to the parameter distribution rule of the weight matrix is greater than a similarity threshold, then configuring the screening strategy for the candidate orthogonal transformation matrix as a random selection strategy.

6. The method according to any one of claims 2 to 5, characterized in that The weight matrix includes an input parameter weight matrix and an output parameter weight matrix; the performing matrix multiplication on the orthogonal transformation matrix and the weight matrix to obtain an adjusted matrix includes: Taking the orthogonal transformation matrix as the left multiplication matrix and the input parameter weight matrix as the right multiplication matrix, and calculating to obtain an adjusted input parameter weight matrix; Using the transpose matrix of the orthogonal transformation matrix as the right multiplication matrix and the output parameter weight matrix as the left multiplication weight matrix, calculate the adjusted output parameter weight matrix.

7. The method according to claim 1, wherein The dimensionality reduction process on the adjusted matrix to obtain the compressed matrix includes: Conduct a parameter contribution analysis on the adjusted matrix to obtain the contribution degrees of each parameter in the adjusted matrix in the model to be compressed; Based on the contribution degrees of each parameter in the model to be compressed and the contribution degree threshold, perform parameter pruning on the adjusted matrix to obtain the compressed matrix; or, based on the contribution degrees of each parameter in the model to be compressed and a preset ratio, perform parameter pruning on the adjusted matrix to obtain the compressed matrix.

8. The method according to claim 7, characterized in that, The conduct of a parameter contribution analysis on the adjusted matrix to obtain the contribution degrees of each parameter in the adjusted matrix in the model to be compressed includes: Perform mean normalization on the adjusted matrix to obtain the matrix after mean normalization; Calculate the second covariance matrix based on the matrix after mean normalization; Perform eigenvalue decomposition on the second covariance matrix to obtain the second eigenvector matrix and the second eigenvalue matrix, where the second eigenvector matrix and the second eigenvalue matrix are used to indicate the contribution degrees of each parameter in the adjusted matrix in the model to be compressed.

9. The method according to claim 8, wherein The perform of parameter pruning on the adjusted matrix based on the contribution degrees of each parameter in the model to be compressed and the contribution degree threshold to obtain the compressed matrix includes: Remove the eigenvectors in the eigenvector matrix with eigenvalues less than the contribution degree threshold to obtain the compressed eigenvector matrix; Perform matrix multiplication calculation on the adjusted matrix and the compressed eigenvector matrix to obtain the compressed matrix.

10. The method according to claim 1, characterized in that The method further includes: Verify the output stability of the compressed model; If the compressed model fails the output stability verification, then adjust the weight matrix using the updated parameter adjustment strategy, and generate a new compressed model based on the adjustment result; or, use the updated parameter removal strategy to remove the redundant parameters in the adjusted matrix, and generate a new compressed model based on the removal result; or, perform lightweight fine-tuning on the compressed model.

11. The method according to claim 10, wherein The verification of the output stability of the compressed model includes: Call the model to be compressed and the compressed model respectively to process M pieces of data to be processed, where M is a positive integer; If the processing results of the j-th piece of data to be processed are inconsistent, then verify the validity of the j-th processing result of the compressed model, where j is a positive integer less than or equal to M; Generate the stability verification result of the compressed model based on the number of identical processing results and the validity verification result of the j-th processing result.

12. A model compression device, characterized in that, The model compression device includes: An acquisition unit, configured to acquire the weight matrix of the model to be compressed, where the model to be compressed includes a neural network, and the weight matrix is extracted from the attention layer in the neural network; A processing unit for adjusting the feature distribution of the weight matrix to obtain an adjusted matrix, wherein the feature density of the adjusted matrix in the target matrix region is higher than that of the weight matrix in the target matrix region; and for performing dimensionality reduction processing on the adjusted matrix to obtain a compressed matrix, wherein the dimension of the compressed matrix is smaller than that of the weight matrix; and for generating a compressed model corresponding to the model to be compressed based on the compressed matrix.

13. A computer device, characterized in that, Comprising: A memory in which a computer program is stored; A processor for loading the computer program to implement the model compression method according to any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the model compression method according to any one of claims 1-11.

15. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the model compression method according to any one of claims 1-11.