Model processing method and device, program product and electronic equipment

By decomposing the weight matrix of the large language model low-rank matrix, generating a weight reduction matrix and cacheing intermediate results, the problem of KV Cache occupies too much video memory, improving the inference efficiency and maintaining the model accuracy.

CN119940537APending Publication Date: 2025-05-06CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411997648.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Due to the rapid growth of KV Cache in inference tasks, the large language model has over-used GPU video memory, which affects the inference efficiency.

Method used

By decomposing the low-rank matrix of the model, two weight reduction weight matrices are obtained. These matrices are used to reduce the input information dimensionally, and intermediate matrices are generated, and cached for later use.

Benefits of technology

It significantly reduces the space occupied by KV Cache in GPU video memory, improves the efficiency of inference tasks, and reduces the video memory requirements while ensuring model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940537A_ABST
    Figure CN119940537A_ABST
Patent Text Reader

Abstract

The invention provides a model processing method and device, a program product and electronic equipment, and relates to the technical field of computers. The method comprises the following steps: receiving a to-be-processed request; the to-be-processed request comprises input information; performing dimension reduction processing on the input information by using a first dimension reduction weight matrix of the model to obtain a first intermediate matrix; performing dimension reduction processing on the input information by using a second dimension reduction weight matrix of the model to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrixes with a unique corresponding relation of the model; and caching the first intermediate matrix and the second intermediate matrix so as to process the received new to-be-processed request. According to the method and the device, the input information is subjected to the dimension reduction processing, and the information subjected to the dimension reduction processing is cached, so that the video memory occupation of the KV Cache is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a model processing method, device, program product and electronic device. Background Art

[0002] The reasoning task of a large language model is a process in which one unit is based on the token obtained by another unit, that is, token by token. Specifically, for each reasoning request, the generation of each token depends on all previous inputs and generated tokens. Therefore, in the reasoning task of a large language model, KV Cache is generally used to cache the intermediate results of previously generated tokens to significantly reduce repeated calculations and reduce task latency.

[0003] However, KV Cache is a space-for-time strategy. The reduced inference latency comes at the expense of occupying more resources, such as more GPU (Graphics Processing Unit) memory resources. Specifically, as the input sequence length and generation length increase, the KV Cache that needs to be cached will grow rapidly, occupying a large amount of GPU memory.

[0004] In the related technology, a KV dimensionality reduction method for KV Cache quantization is provided, which quantizes the cached intermediate results into int8, int4, int2, etc. to reduce the GPU memory usage, but this method may cause serious loss of accuracy of large language models. Summary of the invention

[0005] The present disclosure provides a model processing method, a model processing device, a computer program product and an electronic device for reducing GPU memory occupancy.

[0006] According to a first aspect of the present disclosure, a model processing method is provided, the method comprising:

[0007] Receiving a request to be processed; the request to be processed includes input information;

[0008] The input information is subjected to dimensionality reduction processing using a first dimensionality reduction weight matrix of the model to obtain a first intermediate matrix; and the input information is subjected to dimensionality reduction processing using a second dimensionality reduction weight matrix of the model to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model that have a unique corresponding relationship;

[0009] The first intermediate matrix and the second intermediate matrix are cached to process the received new request to be processed.

[0010] In a possible implementation, before receiving the request to be processed, the method further includes:

[0011] Determining a rank matrix when performing low-rank matrix decomposition on the model;

[0012] According to the rank matrix and the singular value decomposition method, the first weight matrix of the model is decomposed to obtain the first dimension reduction weight matrix of the model.

[0013] In a possible implementation, decomposing the first weight matrix of the model according to the rank matrix and the singular value decomposition method to obtain the first dimension reduction weight matrix of the model includes:

[0014] According to the singular value decomposition method, decomposing the first weight matrix into a first orthogonal matrix, a rank matrix and a first transposed matrix;

[0015] A first dimensionality reduction weight matrix of the model is obtained according to the first orthogonal matrix and the rank matrix.

[0016] In a possible implementation, the method further includes:

[0017] Determining a rank matrix when performing low-rank matrix decomposition on the model;

[0018] According to the rank matrix and the singular value decomposition method, the second weight matrix of the model is decomposed to obtain the second dimension-reduced weight matrix of the model.

[0019] In a possible implementation, the second weight matrix of the model is decomposed according to the rank matrix and the singular value decomposition method to obtain a first dimension reduction weight matrix of the model, including:

[0020] According to the singular value decomposition method, decomposing the second weight matrix into a second orthogonal matrix, a rank matrix and a second transposed matrix;

[0021] A second dimensionality reduction weight matrix of the model is obtained according to the second orthogonal matrix and the rank matrix.

[0022] In a possible implementation, the method further includes:

[0023] Obtaining a first dimension-upgraded weight matrix of the model according to the first transposed matrix and the rank matrix, and fusing the first dimension-upgraded weight matrix with the third weight matrix to obtain first fusion information;

[0024] According to the second transposed matrix and the rank matrix, a second dimension-upgraded weight matrix of the model is obtained, and the second dimension-upgraded weight matrix and the fourth weight matrix are fused to obtain second fusion information;

[0025] The first fused information and the second fused information are input into the model to respond to the pending request.

[0026] In a possible implementation, determining a rank matrix when performing low-rank matrix decomposition on the model includes:

[0027] According to a plurality of rank thresholds of each network layer in each network layer of the model, reversely determine the discard information of the initial rank matrix;

[0028] The initial rank matrix is ​​processed according to the discard information to determine the rank matrix when performing low-rank matrix decomposition on the model.

[0029] According to a second aspect of the present disclosure, a model processing device is provided, the device comprising:

[0030] A receiving unit, configured to receive a request to be processed; the request to be processed includes input information;

[0031] An obtaining unit is used to perform dimensionality reduction processing on the input information using a first dimensionality reduction weight matrix of the model to obtain a first intermediate matrix; and to perform dimensionality reduction processing on the input information using a second dimensionality reduction weight matrix of the model to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model that have a unique corresponding relationship;

[0032] A response unit is used to cache the first intermediate matrix and the second intermediate matrix to process a received new request to be processed.

[0033] In a possible implementation manner, the device further includes a processing unit, configured to:

[0034] Determining a rank matrix when performing low-rank matrix decomposition on the model;

[0035] According to the rank matrix and the singular value decomposition method, the first weight matrix of the model is decomposed to obtain the first dimension reduction weight matrix of the model.

[0036] In a possible implementation manner, the device further includes a processing unit, configured to:

[0037] According to the singular value decomposition method, decomposing the first weight matrix into a first orthogonal matrix, a rank matrix and a first transposed matrix;

[0038] A first dimensionality reduction weight matrix of the model is obtained according to the first orthogonal matrix and the rank matrix.

[0039] In a possible implementation manner, the device further includes a processing unit, configured to:

[0040] Determining a rank matrix when performing low-rank matrix decomposition on the model;

[0041] According to the rank matrix and the singular value decomposition method, the second weight matrix of the model is decomposed to obtain the second dimension-reduced weight matrix of the model.

[0042] In a possible implementation manner, the device further includes a processing unit, configured to:

[0043] According to the singular value decomposition method, decomposing the second weight matrix into a second orthogonal matrix, a rank matrix and a second transposed matrix;

[0044] A second dimensionality reduction weight matrix of the model is obtained according to the second orthogonal matrix and the rank matrix.

[0045] In a possible implementation manner, the device further includes a processing unit, configured to:

[0046] Obtaining a first dimension-upgraded weight matrix of the model according to the first transposed matrix and the rank matrix, and fusing the first dimension-upgraded weight matrix with the third weight matrix to obtain first fusion information;

[0047] According to the second transposed matrix and the rank matrix, a second dimension-upgraded weight matrix of the model is obtained, and the second dimension-upgraded weight matrix and the fourth weight matrix are fused to obtain second fusion information;

[0048] The first fused information and the second fused information are input into the model to respond to the pending request.

[0049] In a possible implementation manner, the device further includes a processing unit, configured to:

[0050] According to a plurality of rank thresholds of each network layer in each network layer of the model, reversely determine the discard information of the initial rank matrix;

[0051] The initial rank matrix is ​​processed according to the discard information to determine the rank matrix when performing low-rank matrix decomposition on the model.

[0052] According to a third aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the method of the first aspect and possible implementations thereof are implemented.

[0053] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the method of the above-mentioned first aspect and its possible implementation methods by executing the executable instructions.

[0054] The technical solution disclosed in this disclosure has the following beneficial effects:

[0055] In an embodiment of the present disclosure, a pending request can be received; the pending request includes input information; the input information is subjected to dimensionality reduction processing using the first dimensionality reduction weight matrix of the model to obtain a first intermediate matrix; and the input information is subjected to dimensionality reduction processing using the second dimensionality reduction weight matrix of the model to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model with a unique correspondence; further, the first intermediate matrix and the second intermediate matrix can be cached to process the received new pending request. It can be seen that in the embodiment of the present disclosure, the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix obtained by low-rank matrix decomposition are used to obtain the intermediate results after dimensionality reduction, namely the first intermediate matrix and the second intermediate matrix, and the first intermediate matrix and the second intermediate matrix are cached. Since the storage space occupied by the intermediate result after dimensionality reduction is much smaller than the original storage, a substantial compression of the KV Cache is achieved, which significantly reduces the video memory occupancy.

[0056] Other features and advantages of the present disclosure will be described in the following description, and partly become apparent from the description, or be understood by practicing the present disclosure. The purpose and other advantages of the present disclosure can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the embodiments of the present disclosure will be briefly introduced below. Obviously, the drawings introduced below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0058] Figure 1 A schematic diagram of an application scenario in this exemplary embodiment is shown;

[0059] Figure 2 A schematic flow chart showing a model processing method in this exemplary embodiment;

[0060] Figure 3 A schematic diagram showing a process of model processing in this exemplary embodiment;

[0061] Figure 4 A schematic diagram showing the structure of a model processing device in this exemplary embodiment is shown;

[0062] Figure 5A schematic structural diagram of an electronic device in this exemplary embodiment is shown. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure. In the absence of conflict, the embodiments in the present disclosure and the features in the embodiments can be arbitrarily combined with each other. In addition, although the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in an order different from that here.

[0064] The term "comprising" and any variations thereof in the specification and claims of the present disclosure are intended to cover non-exclusive protection. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices.

[0065] In the embodiments of the present disclosure, one or more, "many" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or plural.

[0066] It should be noted that the terms "first", "second", "third", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order, sequence, size, and priority. For example, the first pending request and the second pending request in the embodiment of the present disclosure are only used to distinguish different pending requests. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0067] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, which are schematic diagrams of the present disclosure and are not necessarily drawn to scale. Some of the block diagrams shown in the accompanying drawings may be functional entities, which do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in hardware modules or integrated circuits, or in networks, processors or microcontrollers. The implementation method can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein. The features, structures or characteristics described in the present disclosure may be combined in one or more implementations in any suitable manner. In the following description, many specific details are provided to provide a full description of the implementation methods of the present disclosure. However, those skilled in the art should be aware that one or more specific details may be omitted when implementing the technical solution of the present disclosure, or one or more specific details may be replaced by other methods, components, devices, steps, etc.

[0068] It should be noted that in the embodiments of the present disclosure, some existing solutions in the industry such as software, components, models, etc. may be mentioned, which should be considered as exemplary, and their purpose is only to illustrate the feasibility of implementing the technical solution of the present disclosure, but it does not mean that the applicant has or will necessarily use the solution. In the technical solution of the present disclosure, the collection, dissemination, and use of data are in compliance with relevant national laws and regulations.

[0069] In the embodiments of the present disclosure, in order to better understand the solutions provided by the present disclosure, some key names involved are introduced below, as follows:

[0070] KV Cache: It is an optimization method for large model reasoning. It stores the previously calculated key and value vectors in the attention mechanism. When the model needs to process new input, it can directly obtain the previous results from the cache without having to recalculate the processed parts, which greatly reduces the amount of calculation and improves the processing speed.

[0071] Token: refers to dividing text into smaller units. These units can be words, numbers, symbols, or a combination of them.

[0072] Graphics Processing Unit (GPU): A specialized processor designed to accelerate computer graphics rendering. It provides powerful parallel computing capabilities for both model training and reasoning, and thus plays an important role in training and deploying large-scale neural network models.

[0073] Inference: In the field of artificial intelligence and natural language processing, inference refers to applying a trained model to new, unseen data to make predictions or classifications based on the training completed.

[0074] In the related art, the reasoning task of a large language model is a process in which one unit obtains tokens based on what another unit obtains, i.e., token by token. Specifically, for each reasoning request, the generation of each token depends on all previous inputs and generated tokens. Therefore, in the reasoning task of a large language model, KV Cache is generally used to cache the intermediate results of previously generated tokens, so as to significantly reduce repeated calculations and reduce task latency.

[0075] However, KV Cache is a space-for-time strategy. The reduced inference latency takes up more resources, such as more GPU (Graphics Processing Unit) memory resources. Specifically, as the input sequence length and generation length increase, the KV Cache that needs to be cached will grow rapidly, occupying a large amount of GPU memory.

[0076] In the related technology, a KV dimensionality reduction method for KV Cache quantization is provided, which quantizes the cached intermediate results into int8, int4, int2, etc. to reduce the GPU memory usage, but this method may cause serious loss of accuracy of large language models.

[0077] In view of this, the embodiment of the present disclosure provides a model processing method, through which a request to be processed can be received; the request to be processed includes input information; the input information is subjected to dimensionality reduction processing using the first dimensionality reduction weight matrix of the model to obtain a first intermediate matrix; and the input information is subjected to dimensionality reduction processing using the second dimensionality reduction weight matrix of the model to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model with a unique corresponding relationship; further, the first intermediate matrix and the second intermediate matrix can be cached to process the received new request to be processed. It can be seen that the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix obtained by low-rank matrix decomposition in the embodiment of the present disclosure obtain the intermediate results after dimensionality reduction, namely the first intermediate matrix and the second intermediate matrix, and the first intermediate matrix and the second intermediate matrix are cached. Since the storage space occupied by the intermediate result after dimensionality reduction is much smaller than the original storage, a substantial compression of the KV Cache is achieved, which significantly reduces the video memory occupancy.

[0078] In order to better understand the technical solutions provided by the embodiments of the present disclosure, the following briefly introduces the application scenarios to which the technical solutions provided by the embodiments of the present disclosure are applicable. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present disclosure and are not intended to limit the present disclosure. In specific implementation, the technical solutions provided by the embodiments of the present disclosure can be flexibly applied according to actual needs.

[0079] In the embodiments of the present disclosure, the model processing technology can be applied to various business scenarios that require the processing of models that perform KV Cache compression, which is not limited in the embodiments of the present disclosure. Among them, the model is, for example, various mainstream large language models, such as Llama2-7b, Qwen, etc., and the model processing technology provided by the embodiments of the present disclosure does not require retraining the model, so that the existing model can be easily compressed and optimized, thereby improving the practicality and flexibility of the model.

[0080] See also Figure 1 As shown, Figure 1 This is an application scenario to which the technical solution of the embodiment of the present disclosure can be applied. In the scenario diagram, multiple terminal devices 101 and electronic devices 102 are included. Among them, the terminal device 101-1, the terminal device 101-2, ..., the terminal device 101-n can be used by different users.

[0081] The terminal device 101 and the electronic device 102 , as well as each terminal device 101 , may be directly or indirectly connected to each other through one or more networks 103 .

[0082] In an embodiment of the present disclosure, a user can trigger a task processing request to a large language model based on a terminal device 101, and then the terminal device 101 sends the task processing request to an electronic device 102, so that the electronic device 102 can receive the request to be processed; the request to be processed includes input information; the input information is subjected to dimensionality reduction processing using a first dimensionality reduction weight matrix of the model to obtain a first intermediate matrix; and, the input information is subjected to dimensionality reduction processing using a second dimensionality reduction weight matrix of the model to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model that have a unique correspondence; further, based on the first intermediate matrix and the second intermediate matrix, the information pre-stored in the model can be replaced, so that the first intermediate matrix and the second intermediate matrix are input into the model to obtain response information of the request to be processed.

[0083] In the disclosed embodiment, Figure 1 The terminal devices 101 may be mobile phones, tablet computers (PADs), personal computers (PCs), smart TVs, smart watches, smart speakers, smart car devices, and wearable devices, but are not limited thereto. These devices may have the function of supporting triggering task processing requests.

[0084] as well as, Figure 1 The electronic device 102 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server or cloud server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as big data and artificial intelligence platforms, but is not limited to these.

[0085] Of course, the method provided in the embodiment of the present disclosure is not limited to Figure 1 The application scenarios shown can also be used in other possible application scenarios, such as application scenarios where the model processing method is implemented only by electronic devices, and the embodiments of the present disclosure are not limited thereto.

[0086] To further illustrate the technical solution provided by the embodiments of the present disclosure, this is described in detail below in conjunction with the accompanying drawings and specific implementation methods. Although the embodiments of the present disclosure provide method operation steps as shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on routine or no creative labor. In the steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiments of the present disclosure. The method may be executed in the order of the method shown in the embodiments or drawings or in parallel during the actual processing process or when the device is executed.

[0087] See also Figure 2 , Figure 2 FIG. 1 is a flow chart of a model processing method in an embodiment of the present disclosure. The flow of the method may be executed by an electronic device, for example, and the electronic device may be Figure 1 The electronic device 102 in the embodiment is executed, and the specific implementation process of the method is as follows:

[0088] Step 201: receiving a request to be processed; the request to be processed includes input information;

[0089] Step 202: using the first dimensionality reduction weight matrix of the model to perform dimensionality reduction processing on the input information to obtain a first intermediate matrix; and using the second dimensionality reduction weight matrix of the model to perform dimensionality reduction processing on the input information to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model that have a unique corresponding relationship;

[0090] Step 203: Cache the first intermediate matrix and the second intermediate matrix to process the received new request to be processed.

[0091] In an embodiment of the present disclosure, a request to be processed can be received; the request to be processed includes input information; the input information is subjected to dimensionality reduction processing using the first dimensionality reduction weight matrix of the model to obtain a first intermediate matrix; and the input information is subjected to dimensionality reduction processing using the second dimensionality reduction weight matrix of the model to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model that have a unique corresponding relationship; further, the information pre-stored in the model can be replaced based on the first intermediate matrix and the second intermediate matrix, so that the first intermediate matrix and the second intermediate matrix are input into the model to obtain response information of the request to be processed. It can be seen that in the embodiment of the present disclosure, the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix obtained by low-rank matrix decomposition replace the K and V that originally need to be cached with the intermediate results after dimensionality reduction, namely the first intermediate matrix and the second intermediate matrix. Since the storage space occupied by the intermediate results after dimensionality reduction is much smaller than the original storage, a substantial compression of the KV Cache is achieved, which significantly reduces the video memory occupancy.

[0092] In the embodiment of the present disclosure, before the model provides reasoning services, KV Cache compression of low-rank matrix decomposition can be performed on the trained model. In the following text, taking the model as a large language model as an example, the scheme of KV Cache compression of low-rank matrix decomposition for the trained model is introduced.

[0093] In the embodiments of the present disclosure, see Figure 3 , Figure 3 A schematic diagram of a KV Cache compression scheme for low-rank matrix decomposition of a trained model provided by an embodiment of the present disclosure. In a standard transformer, the QKV-based attention mechanism algorithm can be referenced. Figure 3 The formula shown. Figure 3 In the present invention, New Cache can be understood as the KV cache provided in the present invention, and Origal Cache can be understood as the KV cache provided in the related art.

[0094] In the embodiment of the present disclosure, the two weight matrices in the standard large language model can be decomposed, and the two weight matrices are respectively a first weight matrix and a second weight matrix. Among them, the first weight matrix is, for example, Figure 3 W k , the second weight matrix is, for example, Figure 3 W v Each weight matrix can be decomposed into a reduced-dimensional matrix and an increased-dimensional matrix.

[0095] In the embodiment of the present disclosure, the matrix method can be decomposed into three matrices using SVD singular value decomposition. The specific decomposition can be determined by referring to the following formula 1:

[0096] A=U∑V T Formula 1.

[0097] Optionally, the first weight matrix can be decomposed into a first orthogonal matrix, a rank matrix and a first transposed matrix according to a singular value decomposition method, and then a first dimension reduction weight matrix of the model is obtained according to the first orthogonal matrix and the rank matrix. And, a first dimension increase weight matrix of the model is obtained according to the first transposed matrix and the rank matrix.

[0098] For the first weight matrix W k , see the following formula 2:

[0099]

[0100] It can be seen that W k The first dimension reduction matrix A can be obtained k And the first dimension-raising matrix B k .

[0101] Optionally, the second weight matrix can be decomposed into a second orthogonal matrix, a rank matrix and a second transposed matrix according to a singular value decomposition method, and then a second dimension reduction weight matrix of the model is obtained according to the second orthogonal matrix and the rank matrix. And, a second dimension increase weight matrix of the model is obtained according to the second transposed matrix and the rank matrix.

[0102] For the second weight matrix W v , see the following formula 3:

[0103]

[0104] It can be seen that W v The second dimension reduction matrix A can be obtained v and the second dimension-raising matrix B v .

[0105] In the embodiment of the present disclosure, considering that in a standard transformer, K and V are obtained by multiplying the input information X by its corresponding weights, as shown in the following formulas 4 and 5:

[0106] K=XW k Formula 4.

[0107] V=XW v Formula five.

[0108] Therefore, K and V can be obtained by combining the above formulas 2 to 5, as shown in the following formulas 6 and 7:

[0109] K=XW k =X*A k *B k Formula six.

[0110] V=XW v =X*A v *B v Formula seven.

[0111] Furthermore, we can infer that the compressed intermediate result h k and h v , as shown in the following formula eight and formula nine:

[0112] h k =X*A k Formula eight.

[0113] h v =X*A v Formula nine.

[0114] So far, the compressed intermediate result is obtained, namely Figure 3 h in k and h v .

[0115] In the embodiment of the present disclosure, Σ in the above formula is a matrix whose rank is sorted from large to small. Since the larger the rank value, the greater the impact on the model accuracy, and the smaller the rank, the smaller the impact on the model accuracy, it can be considered to discard part of the rank.

[0116] Optionally, the discarding information of the initial rank matrix can be determined by reverse deduction based on multiple rank thresholds of each network layer in each network layer of the model, and then the initial rank matrix is ​​processed according to the discarding information to determine the rank matrix when performing low-rank matrix decomposition on the model.

[0117] That is to say, the size of the rank is directly related to the accuracy of the model, and the large language model of the transformer architecture is a multi-layer model. If each network layer adopts a uniform rank size, it may cause excessive accuracy loss to the layer. Therefore, the embodiment of the present disclosure provides a method for determining the rank size, which is: setting a threshold for each rank selection, and making the approximation result of the matrix of the first k selected ranks and the original matrix not less than the threshold to ensure the accuracy of the model.

[0118] For example, use calibration data, such as data in wikitext2, as model input, perform forward propagation, and calculate the rank of the original model layer; then discard a portion of the ranks, such as one percent, and calculate logits. If the loss does not exceed the threshold corresponding to the network layer, the threshold is recorded. Continue to discard a portion of the ranks and calculate the loss of logits and the original model again. If the loss exceeds the threshold, stop discarding ranks and select the previous value of the current rank that does not exceed the threshold as the number of ranks retained at the end of the layer. And, other network layers can continue to operate in the above manner until the ranks of all network layers are selected.

[0119] In this way, by selecting the rank in the rank matrix, the intermediate cache of the model reasoning process can be greatly reduced while ensuring acceptable model accuracy, so as to achieve the purpose of greatly improving the reasoning efficiency of large models.

[0120] In the disclosed embodiment, considering that the introduction of SVD brings additional computational overhead, the first dimension-raising weight matrix of the model can be obtained according to the first transposed matrix and the rank matrix, and the first dimension-raising weight matrix and the third weight matrix are fused to obtain first fusion information; the second dimension-raising weight matrix of the model can be obtained according to the second transposed matrix and the rank matrix, and the second dimension-raising weight matrix and the fourth weight matrix are fused to obtain second fusion information; the first fusion information and the second fusion information are input into the model to respond to the request to be processed. In this way, the additional computational overhead can be offset.

[0121] In the disclosed embodiment, the large language model can be processed in the aforementioned manner to obtain a processed large language model. For example, after the llama2-7b model is processed in the aforementioned manner, a llama2-7b-low-rank model can be obtained. For another example, after the qwen model is processed in the aforementioned manner, a processed qwen model can be obtained.

[0122] In the disclosed embodiment, after introducing the compression processing of the large language model, the specific process of the large language model reasoning part after the compression processing is introduced below.

[0123] In step 201, a request to be processed is received, wherein the request is, for example, an application programming interface (API) request.

[0124] In step 202, a first intermediate matrix is ​​obtained based on the input information carried by the request to be processed and the first dimensionality reduction weight matrix of the model; and a second intermediate matrix is ​​obtained based on the input information carried by the request to be processed and the second dimensionality reduction weight matrix of the model; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on the weight matrix of the model.

[0125] In the embodiment of the present disclosure, the input information X can be substituted into the above formula 8 and formula 9, that is, according to the input information X and the dimension reduction matrix A of K k Multiply to get the intermediate hidden state h of K k , and the dimension reduction matrix A of V v Multiplied to the intermediate hidden state h of V v .

[0126] Step 203: Cache the first intermediate matrix and the second intermediate matrix to process the received new request to be processed.

[0127] In the embodiment of the present disclosure, h k and h v Cache, thus based on h k and h v Process the next calculation.

[0128] It can be seen that in the embodiment of the present disclosure, through low-rank matrix decomposition, compared with the previously cached unprocessed K and V, the intermediate result h after dimensionality reduction is now cached. k and h v , significantly reducing the memory usage of KV Cache, and the method of dynamically selecting the rank size ensures that the accuracy of the model will not be significantly affected while compressing the KV Cache. In this way, by setting a suitable threshold, the performance of the model during reasoning can be guaranteed to be close to that of the uncompressed version of the model during reasoning.

[0129] In addition, by reducing the memory usage and solving the memory bandwidth throughput bottleneck in the inference process, the present disclosure significantly improves the inference efficiency of large language models. Especially in resource-constrained environments, such as edge devices and mobile devices, this method can greatly improve the throughput of the model, improve concurrency, and processing capabilities in various scenarios such as long texts and multi-round conversations.

[0130] The exemplary embodiment of the present disclosure also provides a model processing device. Figure 4 As shown, the model processing device 400 includes the following program modules:

[0131] The receiving unit 401 is used to receive a request to be processed; the request to be processed includes input information;

[0132] The obtaining unit 402 is used to perform dimensionality reduction processing on the input information using the first dimensionality reduction weight matrix of the model to obtain a first intermediate matrix; and to perform dimensionality reduction processing on the input information using the second dimensionality reduction weight matrix of the model to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model that have a unique corresponding relationship;

[0133] The responding unit 403 is configured to cache the first intermediate matrix and the second intermediate matrix to process a received new request to be processed.

[0134] In a possible implementation manner, the device further includes a processing unit, configured to:

[0135] Determining a rank matrix when performing low-rank matrix decomposition on the model;

[0136] According to the rank matrix and the singular value decomposition method, the first weight matrix of the model is decomposed to obtain the first dimension reduction weight matrix of the model.

[0137] In a possible implementation manner, the device further includes a processing unit, configured to:

[0138] According to the singular value decomposition method, decomposing the first weight matrix into a first orthogonal matrix, a rank matrix and a first transposed matrix;

[0139] A first dimensionality reduction weight matrix of the model is obtained according to the first orthogonal matrix and the rank matrix.

[0140] In a possible implementation manner, the device further includes a processing unit, configured to:

[0141] Determining a rank matrix when performing low-rank matrix decomposition on the model;

[0142] According to the rank matrix and the singular value decomposition method, the second weight matrix of the model is decomposed to obtain the second dimension-reduced weight matrix of the model.

[0143] In a possible implementation manner, the device further includes a processing unit, configured to:

[0144] According to the singular value decomposition method, decomposing the second weight matrix into a second orthogonal matrix, a rank matrix and a second transposed matrix;

[0145] A second dimensionality reduction weight matrix of the model is obtained according to the second orthogonal matrix and the rank matrix.

[0146] In a possible implementation manner, the device further includes a processing unit, configured to:

[0147] Obtaining a first dimension-upgraded weight matrix of the model according to the first transposed matrix and the rank matrix, and fusing the first dimension-upgraded weight matrix with the third weight matrix to obtain first fusion information;

[0148] According to the second transposed matrix and the rank matrix, a second dimension-upgraded weight matrix of the model is obtained, and the second dimension-upgraded weight matrix and the fourth weight matrix are fused to obtain second fusion information;

[0149] The first fused information and the second fused information are input into the model to respond to the pending request.

[0150] In a possible implementation manner, the device further includes a processing unit, configured to:

[0151] According to a plurality of rank thresholds of each network layer in each network layer of the model, reversely determine the discard information of the initial rank matrix;

[0152] The initial rank matrix is ​​processed according to the discard information to determine the rank matrix when performing low-rank matrix decomposition on the model.

[0153] The specific details of each part of the above-mentioned device have been described in detail in the implementation method of the method part. The undisclosed details can be found in the implementation method of the method part, so they will not be repeated here.

[0154] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the exemplary embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.

[0155] The exemplary embodiments of the present disclosure also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the above-mentioned model processing method is implemented.

[0156] In one embodiment, the computer program product may be a tangible product containing a computer program, such as a computer-readable storage medium storing a computer program. The readable storage medium may be a storage medium based on electrical, magnetic, optical, electromagnetic, infrared, or other signals, including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory (Flash), mechanical hard disk (HDD), solid-state drive (SSD), and the like. Exemplarily, the computer program product may be implemented as a non-volatile storage medium storing a computer program, such as a read-only memory, a NAND flash memory, and the like.

[0157] In one embodiment, the computer program product may be an intangible product including a computer program. Exemplarily, the computer program product may be implemented as a virtual digital product, such as a digital file storing an executable file, an installation package, etc. of the computer program.

[0158] The code of the computer program can be written in one or more programming languages. Programming languages ​​such as C language, Java, C++, etc. The program code can be executed entirely on the user computing device, or partially on the user computing device, or as a separate software package, or partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, such as a local area network (LAN), a wide area network (WAN), etc., or can be connected to an external computing device (e.g., an Internet connection provided by an operator).

[0159] The computer program can be carried or transmitted by electrical, magnetic, optical, electromagnetic, infrared and other signals. The electronic device can convert the signal carrying the computer program into a digital signal, and then run the computer program. When the computer program is running on the electronic device, its code is used to make the electronic device execute (more specifically, the processor of the electronic device can be executed) the method steps of various exemplary embodiments of the present disclosure, such as the above-mentioned model processing method, which includes the following steps: Step 201: Receive a pending request; the pending request includes input information; Step 202: Use the first dimensionality reduction weight matrix of the model to reduce the dimension of the input information to obtain a first intermediate matrix; and, use the second dimensionality reduction weight matrix of the model to reduce the dimension of the input information to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model with a unique corresponding relationship; Step 203: Cache the first intermediate matrix and the second intermediate matrix to process the received new pending request.

[0160] By implementing the above method steps through a computer program, a request to be processed can be received; the request to be processed includes input information; the input information is subjected to dimensionality reduction processing using the first dimensionality reduction weight matrix of the model to obtain a first intermediate matrix; and the input information is subjected to dimensionality reduction processing using the second dimensionality reduction weight matrix of the model to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model with a unique corresponding relationship; further, the first intermediate matrix and the second intermediate matrix can be cached to process the received new request to be processed. It can be seen that in the embodiment of the present disclosure, the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix obtained by low-rank matrix decomposition are used to obtain the intermediate results after dimensionality reduction, namely the first intermediate matrix and the second intermediate matrix, and the first intermediate matrix and the second intermediate matrix are cached. Since the storage space occupied by the intermediate result after dimensionality reduction is much smaller than the original storage, a substantial compression of the KV Cache is achieved, which significantly reduces the video memory occupancy.

[0161] The exemplary embodiments of the present disclosure also provide an electronic device, which may include a processor and a memory. The memory stores executable instructions of the processor, such as a computer program. The processor executes the method steps of various exemplary embodiments of the present disclosure by executing the executable instructions.

[0162] Reference below Figure 5 , the electronic device is exemplarily described in the form of a general-purpose computing device. It should be understood that Figure 5 The electronic device 500 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0163] like Figure 5 As shown, the electronic device 500 may include: a processor 510 , a memory 520 , a bus 530 , an I / O (input / output) interface 540 , a network adapter 550 , and a display 580 .

[0164] The memory 520 may include a volatile memory, such as a RAM 521, a cache unit 522, and may also include a non-volatile memory, such as a ROM 523. The memory 520 may also include one or more program modules 524, such program modules 524 include but are not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or a combination thereof may include the implementation of a network environment. For example, the program module 524 may include each module in the above-mentioned device.

[0165] The processor 510 may include one or more processing units. For example, the processor 510 may include an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor and / or an NPU (Neural-Network Processing Unit), etc.

[0166] The processor 510 can be used to execute executable instructions stored in the memory 520, such as the above-mentioned model processing method, which includes the following steps: Step 201: Receive a request to be processed; the request to be processed includes input information; Step 202: Use the first dimensionality reduction weight matrix of the model to reduce the dimension of the input information to obtain a first intermediate matrix; and, use the second dimensionality reduction weight matrix of the model to reduce the dimension of the input information to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model with a unique correspondence; Step 203: Cache the first intermediate matrix and the second intermediate matrix to process the received new request to be processed.

[0167] By executing the above method steps through the processor 510, a request to be processed can be received; the request to be processed includes input information; the input information is subjected to dimensionality reduction processing using the first dimensionality reduction weight matrix of the model to obtain a first intermediate matrix; and the input information is subjected to dimensionality reduction processing using the second dimensionality reduction weight matrix of the model to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model with a unique corresponding relationship; further, the first intermediate matrix and the second intermediate matrix can be cached to process the received new request to be processed. It can be seen that in the embodiment of the present disclosure, the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix obtained by low-rank matrix decomposition obtain the intermediate results after dimensionality reduction, namely the first intermediate matrix and the second intermediate matrix, and the first intermediate matrix and the second intermediate matrix are cached. Since the storage space occupied by the intermediate result after dimensionality reduction is much smaller than the original storage, a substantial compression of the KV Cache is achieved, which significantly reduces the video memory occupancy.

[0168] The bus 530 is used to realize the connection between different components of the electronic device 500, and may include a data bus, an address bus, and a control bus.

[0169] The electronic device 500 can communicate with one or more external devices 600 (eg, a keyboard, a mouse, an external controller, etc.) through the I / O interface 540 .

[0170] The electronic device 500 can communicate with one or more networks through the network adapter 550. For example, the network adapter 550 can provide mobile communication solutions such as 3G / 4G / 5G, or provide wireless communication solutions such as wireless LAN, Bluetooth, near field communication, etc. The network adapter 550 can communicate with other modules of the electronic device 500 through the bus 530.

[0171] The electronic device 500 can display various types of information and the like through the display 580 .

[0172] although Figure 5 Not shown, other hardware and / or software modules may also be provided in the electronic device 500, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0173] As can be seen from the above, the technical solution of the present disclosure can be implemented as a method, an apparatus, a system, a computer program product, a storage medium, an electronic device, etc. Those skilled in the art can understand that various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software, such as being respectively referred to as a "circuit", "module" or "system".

[0174] It should be understood that the present disclosure is not limited to the specific method steps or structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from the scope thereof. Those skilled in the art will easily think of other embodiments based on the specific embodiments provided by the present disclosure. Therefore, the specific embodiments provided by the present disclosure are only exemplary, and the scope and spirit of the present disclosure are indicated by the claims, and any variations, uses or adaptive changes of the present disclosure should be covered, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the technical field that are not disclosed in the present disclosure.

Claims

1. A model processing method, characterized in that: The method comprises: Receiving a request to be processed; the request to be processed includes input information; The input information is subjected to dimensionality reduction processing using a first dimensionality reduction weight matrix of the model to obtain a first intermediate matrix; and the input information is subjected to dimensionality reduction processing using a second dimensionality reduction weight matrix of the model to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model that have a unique corresponding relationship; The first intermediate matrix and the second intermediate matrix are cached to process the received new request to be processed.

2. The method according to claim 1, characterized in that Before receiving the request to be processed, the method further includes: Determining a rank matrix when performing low-rank matrix decomposition on the model; According to the rank matrix and the singular value decomposition method, the first weight matrix of the model is decomposed to obtain the first dimension reduction weight matrix of the model.

3. The method according to claim 2, characterized in that Decomposing the first weight matrix of the model according to the rank matrix and the singular value decomposition method to obtain the first dimension reduction weight matrix of the model includes: Decomposing the first weight matrix into a first orthogonal matrix, a rank matrix and a first transposed matrix; A first dimensionality reduction weight matrix of the model is obtained according to the first orthogonal matrix and the rank matrix.

4. The method according to claim 1, characterized in that: Before receiving the request to be processed, the method further includes: Determining a rank matrix when performing low-rank matrix decomposition on the model; According to the rank matrix and the singular value decomposition method, the second weight matrix of the model is decomposed to obtain the second dimension-reduced weight matrix of the model.

5. The method according to claim 4, characterized in that Decomposing the second weight matrix of the model according to the rank matrix and the singular value decomposition method to obtain the first dimension reduction weight matrix of the model includes: According to the singular value decomposition method, decomposing the second weight matrix into a second orthogonal matrix, a rank matrix and a second transposed matrix; A second dimensionality reduction weight matrix of the model is obtained according to the second orthogonal matrix and the rank matrix.

6. The method according to claim 3 or 5, characterized in that: The method further comprises: Obtaining a first dimension-upgraded weight matrix of the model according to the first transposed matrix and the rank matrix, and fusing the first dimension-upgraded weight matrix with the third weight matrix to obtain first fusion information; According to the second transposed matrix and the rank matrix, a second dimension-upgraded weight matrix of the model is obtained, and the second dimension-upgraded weight matrix and the fourth weight matrix are fused to obtain second fusion information; The first fused information and the second fused information are input into the model to respond to the pending request.

7. The method according to any one of claims 1 to 5, characterized in that: Determining a rank matrix when performing low-rank matrix decomposition on the model includes: According to a plurality of rank thresholds of each network layer in each network layer of the model, reversely determine the discard information of the initial rank matrix; The initial rank matrix is ​​processed according to the discard information to determine the rank matrix when performing low-rank matrix decomposition on the model.

8. A model processing device, characterized in that: The device comprises: A receiving unit, configured to receive a request to be processed; the request to be processed includes input information; An obtaining unit is used to perform dimensionality reduction processing on the input information using a first dimensionality reduction weight matrix of the model to obtain a first intermediate matrix; and to perform dimensionality reduction processing on the input information using a second dimensionality reduction weight matrix of the model to obtain a second intermediate matrix; the first dimensionality reduction weight matrix and the second dimensionality reduction weight matrix are obtained by performing low-rank matrix decomposition on two different weight matrices of the model that have a unique corresponding relationship; A response unit is used to cache the first intermediate matrix and the second intermediate matrix to process a received new request to be processed.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: processor; A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 7 by executing the executable instructions.

Citation Information

Cited By

  • Large language model weight inverse quantization reasoning device and method

    CN121413783A