Large model reasoning method and device based on quantitative perception fine tuning and medium

Through the large-model inference method of quantitative perception fine-tuning, the problem of large-model parameters large-scale and training fine-tuning resources is solved, and the large-model deployment inference under limited hardware resources is realized, which accelerates the inference speed and maintains model performance.

CN119962665AActive Publication Date: 2025-05-09AISINO CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411830901.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-09
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

The existing large models are difficult to achieve privatized deployment due to large parameter scale and training and fine-tuning. Especially in the process of digital transformation of government and enterprises, insufficient hardware resources lead to deployment difficulties.

Method used

The large model inference method based on quantization perception fine-tuning is adopted. By structurally transforming the original parameter matrix of the large model, horizontal vectors, vertical vectors and low-bit fixed matrices are determined, parameter quantization fine-tuning is pre-trained, and deployment parameter matrix is ​​obtained to reduce memory consumption during training and deployment stages.

Benefits of technology

The deployment inference of large models under limited hardware resources is realized, which accelerates the inference speed, avoids the drawbacks of human adjustment in the pruning method, and maintains the performance of the model in language modeling and small sample situation learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962665A_ABST
    Figure CN119962665A_ABST
Patent Text Reader

Abstract

The invention discloses a large model reasoning method and device based on quantitative perception fine tuning and a medium. The method comprises the following steps: performing structure transformation on an original parameter matrix of a large model, and determining a transverse vectorization vector, a longitudinal vectorization vector and a low-bit fixed matrix corresponding to the original parameter matrix of each channel; performing parameter quantization fine tuning pre-training on the large model layer by layer based on the transverse vectorization vector, the longitudinal vectorization vector and the low-bit fixed matrix to obtain a transverse vectorization vector value and a longitudinal vectorization vector value of each channel of the large model; determining a deployment parameter matrix of deployment reasoning of the large model according to the transverse vectorization vector value and the longitudinal vectorization vector value of each channel of the large model and the low-bit fixed matrix; and performing reasoning analysis on the input data by adopting the large model deployed with the deployment parameter matrix to obtain a reasoning result of the input data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large model accelerated reasoning technology, and more specifically, to a large model reasoning method, device and medium based on quantization-aware fine-tuning. Background Art

[0002] With the rapid development of artificial intelligence technology, especially the continuous breakthroughs in the field of natural language processing (NLP), generative AI technology has become an important force in promoting industry innovation and change. Traditional generative models have made significant progress in text generation, but are often limited by the huge model size and computing power requirements. In particular, considering that LLMs have billions or even trillions of parameters, the hardware requirements for training and deployment are extremely high, making it difficult to carry out private deployment and implementation in the process of digital transformation of government and enterprises. From a security perspective, government and enterprise customers often find it difficult to carry out large model training applications through the SaaS service model, resulting in a vicious cycle in which the better the model effect, the larger the scale, the more difficult it is to deploy applications.

[0003] There are currently three methods to address the problems of large model parameter scale, high resource consumption for training and fine-tuning, and difficulty in private deployment: quantization, pruning, and distillation. Quantization is a favorable method to compress and accelerate neural networks by discretizing parameters into low-bit integers (such as converting Float16 to int8) while maintaining a shared high-precision scale within each parameter group (such as a channel or layer). However, quantization requires very low update efficiency for all parameters during the training phase. The large model pruning method effectively reduces the number of model parameters and the amount of calculation by accurately identifying and pruning parameters or connections that contribute less to performance, allowing large models to achieve the same effect while using fewer parameters and connection layers. However, the pruning process may cause slight performance loss, requiring accurate evaluation and fine-tuning after pruning, which often requires a large amount of human resources. Large model distillation uses link layer merging or parameter synchronization to achieve the goal of fast training or fine-tuning with fewer connection layers or multiple connection layers with the same parameters, which often loses a certain degree of accuracy. Summary of the invention

[0004] In view of the deficiencies in the prior art, the present invention provides a large model inference method, device and medium based on quantization-aware fine-tuning.

[0005] According to one aspect of the present invention, a large model inference method based on quantization-aware fine-tuning is provided, comprising:

[0006] The original parameter matrix of the large model is restructured to determine the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix corresponding to the original parameter matrix of each channel;

[0007] Based on the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix, the large model is pre-trained for parameter quantization fine-tuning layer by layer to obtain the horizontal vectorization vector value and the vertical vectorization vector value of each channel of the large model;

[0008] Determine the deployment parameter matrix of the large model deployment reasoning according to the horizontal vectorization vector value, the vertical vectorization vector value and the low-bit fixed matrix of each channel of the large model;

[0009] A large model with a deployed parameter matrix is ​​used to perform reasoning analysis on the input data to obtain the reasoning results of the input data.

[0010] Optionally, the original parameter matrix of the large model is structurally restructured to determine the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix corresponding to the original parameter matrix of each channel, including:

[0011] The high-precision original parameter matrix W of each channel of the large model is decomposed horizontally and vertically to decompose the high-precision horizontal vectorization vector T0, the high-precision vertical vectorization vector S0 and the low-bit fixed matrix W0', where the original parameter matrix W of each channel is ≈ S0*W0'.

[0012] Optionally, the bit width of parameter quantization fine-tuning pre-training is 4.

[0013] Optionally, the large model is pre-trained for parameter quantization fine-tuning layer by layer based on the horizontal vectorization vector, the vertical vectorization vector, and the low-bit fixed matrix to obtain the horizontal vectorization vector value and the vertical vectorization vector value of each channel of the large model, including:

[0014] Based on the longitudinal vectorization vector and low-bit fixed matrix of the first-layer channel, the first-layer channel of the large model is quantized and fine-tuned for pre-training to obtain the first-layer quantization result;

[0015] Based on the longitudinal vectorization vector and low-bit fixed matrix of the next-layer channel, the next-layer channel of the large model is quantized and fine-tuned for pre-training according to the first-layer quantization result and the horizontal vectorization vector to obtain the next quantization result;

[0016] Iteratively use the quantization results of the previous channel of the large model to perform quantization fine-tuning pre-training on the next channel until the last channel of the large model, and obtain the horizontal vectorization vector value and the vertical vectorization vector value of each channel of the large model.

[0017] Optionally, the quantization expression of the channel of this layer of the large model quantization fine-tuning pre-training is:

[0018]

[0019] The quantization expression of the next channel of the large model quantization fine-tuning pre-training is:

[0020] W″0=(T0+ΔT)*W

[0021] Where W is the quantization result of the channel in this layer; W″0 is the quantization result of the channel in the next layer of the channel in this layer; S0 is the vertical quantization vector; T0 is the horizontal quantization vector; ΔS represents the gradient update of S0 obtained by fine-tuning training; ΔT represents the gradient update of T0 obtained by fine-tuning training; z0 represents an adjustment parameter; clamp represents a function that limits the value to [a, b].

[0022] According to another aspect of the present invention, a large model inference device based on quantization-aware fine-tuning is provided, comprising:

[0023] The structural transformation module is used to structurally transform the original parameter matrix of the large model and determine the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix corresponding to the original parameter matrix of each channel;

[0024] A pre-training module is used to perform parameter quantization fine-tuning pre-training on the large model layer by layer based on the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix, and obtain the horizontal vectorization vector value and the vertical vectorization vector value of each channel of the large model;

[0025] A determination module, used to determine a deployment parameter matrix for deployment reasoning of a large model according to a horizontal vectorization vector value, a vertical vectorization vector value, and a low-bit fixed matrix of each channel of the large model;

[0026] The reasoning module is used to perform reasoning analysis on the input data using a large model deployed with a deployment parameter matrix to obtain the reasoning results of the input data.

[0027] Optionally, the structural transformation module includes:

[0028] The decomposition submodule is used to decompose the high-precision original parameter matrix W of each channel of the large model horizontally and vertically, decomposing it into a high-precision horizontal vectorization vector T0, a high-precision vertical vectorization vector S0 and a low-bit fixed matrix W0', where the original parameter matrix of each channel is W≈S0*W0'.

[0029] Optionally, the bit width of parameter quantization fine-tuning pre-training is 4.

[0030] Optionally, pre-training modules include:

[0031] The first pre-training submodule is used to perform quantization fine-tuning pre-training on the first-layer channel of the large model based on the longitudinal vectorization vector and the low-bit fixed matrix of the first-layer channel to obtain the first-layer quantization result;

[0032] The second pre-training submodule is used to perform quantization fine-tuning pre-training on the next layer of channels of the large model based on the longitudinal vectorization vector and the low-bit fixed matrix of the next layer of channels, according to the first layer quantization result and the horizontal vectorization vector, to obtain the next quantization result;

[0033] The three pre-training sub-modules are used to iteratively perform quantization fine-tuning pre-training on the next layer of channels through the quantization results of the previous layer of channels of the large model until the last layer of channels of the large model, and obtain the horizontal vectorization vector value and the vertical vectorization vector value of each channel of the large model.

[0034] Optionally, the quantization expression of the channel of this layer of the large model quantization fine-tuning pre-training is:

[0035]

[0036] The quantization expression of the next channel of the large model quantization fine-tuning pre-training is:

[0037] W0=(T0+ΔT)*W

[0038] Where W is the quantization result of the channel in this layer; W″0 is the quantization result of the channel in the next layer of the channel in this layer; S0 is the vertical quantization vector; T0 is the horizontal quantization vector; ΔS represents the gradient update of S0 obtained by fine-tuning training; ΔT represents the gradient update of T0 obtained by fine-tuning training; z0 represents an adjustment parameter; clamp represents a function that limits the value to [a, b].

[0039] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and the computer program is used to execute the method described in any one of the above aspects of the present invention.

[0040] According to another aspect of the present invention, an electronic device is provided, comprising: a processor; a memory for storing instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method described in any one of the above aspects of the present invention.

[0041] The large model reasoning acceleration method based on quantization-aware fine-tuning provided by the present invention has the following beneficial effects:

[0042] 1. This method inherits the advantages of large quantized models and distillation acceleration, reduces memory consumption in the training and deployment stages, and enables ultra-large-scale language models to be deployed and inferred with very few hardware resources, which can be quickly applied in digital transformation projects of government and enterprises.

[0043] 2. At the same time, due to the reduction in model size brought by quantization, memory consumption is reduced during the deployment phase. By reducing the number of memory accesses, the inference speed during the deployment phase is improved.

[0044] 3. Under low precision, the model can basically maintain its capabilities in language modeling, learning and understanding of a small number of sample situations, and achieve performance comparable to that of high-precision models. This avoids the drawback of fine-tuning and then investing a lot of manpower in adjustments like pruning methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] A more complete understanding of exemplary embodiments of the present invention may be obtained by referring to the following drawings:

[0046] Figure 1 It is a flowchart of a large model inference method based on quantization-aware fine-tuning provided by an exemplary embodiment of the present invention;

[0047] Figure 2 is another flowchart of a large model inference method based on quantization-aware fine-tuning provided by an exemplary embodiment of the present invention;

[0048] Figure 3 is a schematic diagram of an original parameter matrix of a large model provided by an exemplary embodiment of the present invention;

[0049] Figure 4 It is a schematic diagram of structural transformation in a large model reasoning method based on quantization-aware fine-tuning provided by an exemplary embodiment of the present invention;

[0050] Figure 5 is a schematic diagram of quantization fine-tuning of a large model inference method based on quantization-aware fine-tuning provided by an exemplary embodiment of the present invention;

[0051] Figure 6 It is a deployment reasoning diagram of a large model reasoning method based on quantization-aware fine-tuning provided by an exemplary embodiment of the present invention;

[0052] Figure 7 It is a structural schematic diagram of a large model inference device based on quantization-aware fine-tuning provided by an exemplary embodiment of the present invention;

[0053] Figure 8 This is a structure of an electronic device provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0054] Below, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described here.

[0055] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present invention unless specifically stated otherwise.

[0056] Those skilled in the art can understand that the terms "first" and "second" in the embodiments of the present invention are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate the necessary logical order between them.

[0057] It should also be understood that, in the embodiments of the present invention, “plurality” may refer to two or more than two, and “at least one” may refer to one, two or more than two.

[0058] It should also be understood that any component, data or structure mentioned in the embodiments of the present invention can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.

[0059] In addition, the term "and / or" in the present invention is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present invention generally indicates that the associated objects before and after are in an "or" relationship.

[0060] It should also be understood that the description of the various embodiments of the present invention focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced to each other, and for the sake of brevity, they will not be described one by one.

[0061] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0062] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.

[0063] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0064] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0065] Embodiments of the present invention can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments including any of the above systems, etc.

[0066] Electronic devices such as terminal devices, computer systems, servers, etc. can be described in the general context of computer system executable instructions (such as program modules) executed by computer systems. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.

[0067] Exemplary Methods

[0068] Figure 1 is a flow chart of a large model inference method based on quantization-aware fine-tuning provided by an exemplary embodiment of the present invention. This embodiment can be applied to electronic devices, such as Figure 1 As shown, the large model inference method 100 based on quantization-aware fine-tuning includes the following steps:

[0069] Step 101, structurally transforming the original parameter matrix of the large model, and determining the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix corresponding to the original parameter matrix of each channel;

[0070] Step 102, based on the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix, the large model is pre-trained for parameter quantization fine-tuning layer by layer to obtain the horizontal vectorization vector value and the vertical vectorization vector value of each channel of the large model;

[0071] Step 103, determining a deployment parameter matrix for deployment reasoning of the large model according to the horizontal vectorization vector value, the vertical vectorization vector value and the low-bit fixed matrix of each channel of the large model;

[0072] Step 104: Use a large model deployed with a deployment parameter matrix to perform reasoning analysis on the input data to obtain reasoning results of the input data.

[0073] Specifically, the purpose of the present invention is to provide an inference acceleration method based on quantization-aware fine-tuning. Under the structure of traditional quantization compression, combined with the idea of ​​model distillation, by decomposing the parameter matrix of each fully connected layer into a low-bit fixed matrix and a quantization vector, only the quantization vector is fine-tuned in the inference task, while the fixed matrix remains unchanged. For the quantized large model, only the quantization vector is updated, fewer trainable parameters are used, as well as efficient task-specific parameter storage and fast switching, which reduces the use of DRAM during training and deployment, and the inference acceleration achieved due to reduced memory access during deployment. It meets the problem of difficult privatization deployment due to insufficient hardware resources during the digital transformation of government and enterprises, and meets the actual application needs.

[0074] The overall solution of the present invention can be divided into three stages: structural transformation, quantitative fine-tuning, and deployment reasoning.

[0075] The structural transformation aims to modify the original fine-tuning training structure, including two steps: pre-quantization and post-quantization. First, the parameter matrix is ​​converted into a quantized vector and a fixed matrix through pre-quantization. The fixed matrix parameters are discretized into low-bit integers, while maintaining a shared high-precision scale within each parameter group (such as a channel or layer) to compress and accelerate the neural network. Then, the perceptual quantization vector part of each quantization fine-tuning process is post-quantized. For specific examples, see Figure 2 shown.

[0076] The original parameter matrix is ​​the training parameter matrix of the model itself, represented by W0, such as Figure 3 shown.

[0077] Furthermore, if Figure 4 As shown, the structural transformation process decomposes the original parameter matrix into a horizontal quantization vector T0 (also called a horizontal distillation vector) and a vertical quantization vector S0. The decomposed original parameter matrix becomes a low-bit fixed matrix W0', and the parameters of the subsequent fixed matrix no longer change. The original matrix is ​​the product of the vertical quantization vector and the low-bit fixed matrix, that is, W0≈S0*W0'. After obtaining the approximate original parameter matrix W0 through the low-bit fixed matrix W0', the original parameter matrix W0 between each parameter channel is converted through the horizontal quantization vector T0. In this way, the original huge amount of training parameters (k channels * a rows * b columns) can be changed to the current small amount of parameters (horizontal quantization vector * vertical quantization vector) to achieve the same effect.

[0078] Furthermore, in the quantization fine-tuning stage, for the pre-trained original parameter matrix W0 of each fully connected layer, given the bit width b (set to 4 in the present invention), the pre-trained weight matrix W can be expressed as follows (the weight matrix is ​​a general formula. In actual calculation, the matrix W of the first channel is first calculated, and then the weight matrix W of the remaining channels is calculated based on W and T0):

[0079]

[0080] W″0=T0*W=s0*T0*W

[0081] Among them, W0" is the matrix of other channels, · represents the product of two matrices, represents the rounding function, clamp represents the function that limits the value to [a, b], that is, the value in this patent is within [0, 15], and the scale and zero point of each channel, z0 represents an adjustment parameter, when z0 is 0, it means no adjustment, to simplify the calculation, this patent uses 0. In addition, S0 represents the longitudinal vector, T0 represents the transverse vector, and the final weight matrix is ​​obtained by multiplying the variable transverse and longitudinal matrices with the fixed vector matrix, such as Figure 5 shown.

[0082] W'0 is the integer quantized index of W0, which is applied to each fully connected layer in the pre-trained large language model. We then fine-tune only Sx and Tx (outside the clamp function in the formula), while sharing W0' across all downstream tasks (all changes are in S x and T x Therefore, the quantized pre-trained weight W0 is adapted to the downstream task as follows:

[0083]

[0084] W″0=(T0+ΔT)*W

[0085] Among them, ΔS represents the gradient update of S0 obtained through fine-tuning training, and ΔT represents the gradient update of T0 obtained through fine-tuning training. This method is a memory-efficient fine-tuning method designed for quantizing large language models. It combines the idea of ​​large model distillation and only updates the quantization vector S0 and the distillation vector T0. Since W0' is a fixed matrix, it is frozen and shared with all training channels. When it is necessary to switch to different downstream tasks, the values ​​of S0 and T0 can be quickly replaced. Therefore, S0 mainly updates the weight matrix of the current channel, while T0 updates the cross-channel weight matrix. While ensuring that the fixed matrix remains unchanged and there is only one, only the weights of the horizontal and vertical dimensions need to be updated to affect the weight matrix parameters of all channels.

[0086] Furthermore, if Figure 6As shown in the figure, during the deployment and inference stage, the values ​​of the horizontal and vertical vectors Sx and Tx are adjusted to express the difference between different channels. For the fine-tuning task of each channel, only the horizontal and vertical quantization vectors are fine-tuned (that is, only the values ​​of Tx and Sx are updated), while the integer fixed matrix remains unchanged. For the quantized large language model, the quantization perception ability is retained while reducing the amount of calculation for each layer of channels. At the same time, the idea of ​​distillation compression is adopted to reduce the amount of calculation for updating cross-layer vector matrices, using fewer trainable parameters, and efficient task-specific parameter storage and fast switching. At the same time, the benefits of quantization are reflected in deployment and inference, which reduces the use of DRAM during training and deployment (from the original N vector matrices of Float16 or higher precision to a fixed 1 vector matrix of int4), and the reasoning is accelerated due to the reduction of memory access during deployment.

[0087] The key technical points of the present invention are:

[0088] 1. The two general ideas of quantization and distillation are introduced, which convert the update of the traditional high-precision feature (vector) matrix into the update of the vertical quantization vector and the horizontal distillation vector, greatly reducing the resource consumption during deployment inference and fine-tuning.

[0089] 2. This method inherits the advantages of large quantized models and distillation acceleration. It reduces memory consumption during training and deployment, only updates the quantized vector, and keeps the fixed matrix frozen, thus reducing memory usage during training.

[0090] 3. At the same time, due to the reduction in model size brought by quantization, memory consumption is reduced during the deployment phase. By reducing the number of memory accesses, the inference speed during the deployment phase is improved.

[0091] 4. At low precision, the model’s capabilities in language modeling, small sample context learning, and understanding are maintained, achieving performance comparable to or better than that of the full-precision model.

[0092] The large model reasoning acceleration method based on quantization-aware fine-tuning provided by the present invention has the following advantages:

[0093] 1. This method inherits the advantages of large quantized models and distillation acceleration, reduces memory consumption in the training and deployment stages, and enables ultra-large-scale language models to be deployed and inferred with very few hardware resources, which can be quickly applied in digital transformation projects of government and enterprises.

[0094] 2. At the same time, due to the reduction in model size brought by quantization, memory consumption is reduced during the deployment phase. By reducing the number of memory accesses, the inference speed during the deployment phase is improved.

[0095] 3. Under low precision, the model can basically maintain its capabilities in language modeling, learning and understanding of a small number of sample situations, and achieve performance comparable to that of high-precision models. This avoids the drawback of fine-tuning and then investing a lot of manpower in adjustments like pruning methods.

[0096] Exemplary Devices

[0097] Figure 7 Schematic diagram of a large model inference device based on quantization-aware fine-tuning provided by an exemplary embodiment of the present invention. Figure 7 As shown, the apparatus 700 includes:

[0098] The structural transformation module 710 is used to perform structural transformation on the original parameter matrix of the large model, and determine the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix corresponding to the original parameter matrix of each channel;

[0099] A pre-training module 720 is used to perform parameter quantization fine-tuning pre-training on the large model layer by layer based on the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix, and obtain the horizontal vectorization vector value and the vertical vectorization vector value of each channel of the large model;

[0100] A determination module 730 is used to determine a deployment parameter matrix for deployment reasoning of the large model according to the horizontal vectorization vector value, the vertical vectorization vector value and the low-bit fixed matrix of each channel of the large model;

[0101] The reasoning module 740 is used to perform reasoning analysis on the input data using a large model deployed with a deployment parameter matrix to obtain the reasoning results of the input data.

[0102] Optionally, the structural transformation module 710 includes:

[0103] The decomposition submodule is used to decompose the high-precision original parameter matrix W of each channel of the large model horizontally and vertically, decomposing it into a high-precision horizontal vectorization vector T0, a high-precision vertical vectorization vector S0 and a low-bit fixed matrix W0', where the original parameter matrix of each channel is W≈S0*W0'.

[0104] Optionally, the bit width of parameter quantization fine-tuning pre-training is 4.

[0105] Optionally, the pre-training module 720 includes:

[0106] The first pre-training submodule is used to perform quantization fine-tuning pre-training on the first-layer channel of the large model based on the longitudinal vectorization vector and the low-bit fixed matrix of the first-layer channel to obtain the first-layer quantization result;

[0107] The second pre-training submodule is used to perform quantization fine-tuning pre-training on the next layer of channels of the large model based on the longitudinal vectorization vector and the low-bit fixed matrix of the next layer of channels, according to the first layer quantization result and the horizontal vectorization vector, to obtain the next quantization result;

[0108] The three pre-training sub-modules are used to iteratively perform quantization fine-tuning pre-training on the next layer of channels through the quantization results of the previous layer of channels of the large model until the last layer of channels of the large model, and obtain the horizontal vectorization vector value and the vertical vectorization vector value of each channel of the large model.

[0109] Optionally, the quantization expression of the channel of this layer of the large model quantization fine-tuning pre-training is:

[0110]

[0111] The quantization expression of the next channel of the large model quantization fine-tuning pre-training is:

[0112] W″0=(T0+ΔT)*W

[0113] Where W is the quantization result of the channel in this layer; W″0 is the quantization result of the channel in the next layer of the channel in this layer; S0 is the vertical quantization vector; T0 is the horizontal quantization vector; ΔS represents the gradient update of S0 obtained by fine-tuning training; ΔT represents the gradient update of T0 obtained by fine-tuning training; z0 represents an adjustment parameter; clamp represents a function that limits the value to [a, b].

[0114] Exemplary Electronic Devices

[0115] Figure 8 This is a structure of an electronic device provided by an exemplary embodiment of the present invention. Figure 8 As shown, the electronic device 80 includes one or more processors 81 and a memory 82 .

[0116] The processor 81 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0117] The memory 82 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 81 may run the program instructions to implement the methods of the software programs of the various embodiments of the present invention described above and / or other desired functions. In one example, the electronic device may also include: an input device 83 and an output device 84, which are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0118] In addition, the input device 83 may also include, for example, a keyboard, a mouse, etc.

[0119] The output device 84 can output various information to the outside, and can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto.

[0120] Of course, to simplify, Figure 8 Only some of the components related to the present invention in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application conditions.

[0121] Exemplary computer program products and computer-readable storage media

[0122] In addition to the above-mentioned methods and devices, an embodiment of the present invention may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the method according to various embodiments of the present invention described in the above-mentioned "Exemplary Method" section of this specification.

[0123] The computer program product may be written in any combination of one or more programming languages ​​to write program code for performing the operations of the embodiments of the present invention, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0124] In addition, an embodiment of the present invention may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enable the processor to execute the steps of the method according to various embodiments of the present invention described in the above “Exemplary Method” section of this specification.

[0125] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, system or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0126] The basic principle of the present invention is described above in conjunction with specific embodiments. However, it should be pointed out that the advantages, strengths, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. must be possessed by each embodiment of the present invention. In addition, the specific details disclosed above are only for the purpose of illustration and facilitation of understanding, rather than limitation, and the above details do not limit the present invention to being implemented by adopting the above specific details.

[0127] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0128] The block diagrams of the devices, systems, equipment, and systems involved in the present invention are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagram. As will be appreciated by those skilled in the art, these devices, systems, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open words, referring to "including but not limited to", and can be used interchangeably with them. The words "or" and "and" used here refer to the words "and / or" and can be used interchangeably with them, unless the context clearly indicates otherwise. The word "such as" used here refers to the phrase "such as but not limited to", and can be used interchangeably with it.

[0129] The method and system of the present invention may be implemented in many ways. For example, the method and system of the present invention may be implemented by software, hardware, firmware or any combination of software, hardware, firmware. The above order of steps for the method is only for illustration, and the steps of the method of the present invention are not limited to the order specifically described above, unless otherwise specifically stated. In addition, in some embodiments, the present invention may also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present invention. Thus, the present invention also covers a recording medium storing a program for executing the method according to the present invention.

[0130] It should also be noted that in the system, device and method of the present invention, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent schemes of the present invention. The above description of the disclosed aspects is provided to enable any technician in the field to make or use the present invention. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined here can be applied to other aspects without departing from the scope of the present invention. Therefore, the present invention is not intended to be limited to the aspects shown here, but in accordance with the widest range consistent with the principles and novel features disclosed here.

[0131] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present invention to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.

Claims

1. A large model inference method based on quantization-aware fine-tuning, characterized in that: include: The original parameter matrix of the large model is structurally transformed to determine the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix corresponding to the original parameter matrix of each channel; Based on the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix, the large model is pre-trained for parameter quantization fine-tuning layer by layer to obtain the horizontal vectorization vector value and the vertical vectorization vector value of each channel of the large model; Determine a deployment parameter matrix for deployment reasoning of the large model according to the horizontal vectorization vector value, the vertical vectorization vector value, and the low-bit fixed matrix of each channel of the large model; A large model deployed with the deployment parameter matrix is ​​used to perform inference analysis on the input data to obtain an inference result of the input data.

2. The method according to claim 1, characterized in that The original parameter matrix of the large model is structurally transformed to determine the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix corresponding to the original parameter matrix of each channel, including: The high-precision original parameter matrix W of each channel of the large model is decomposed horizontally and vertically to decompose the high-precision horizontal vectorization vector T0, the high-precision vertical vectorization vector S0 and the low-bit fixed matrix W0', where the original parameter matrix W of each channel ≈ S0*W0'.

3. The method according to claim 1, characterized in that The bit width of parameter quantization fine-tuning pre-training is 4.

4. The method according to claim 3, characterized in that: Based on the horizontal vectorization vector, the vertical vectorization vector, and the low-bit fixed matrix, the large model is pre-trained for parameter quantization fine-tuning layer by layer to obtain horizontal vectorization vector values ​​and vertical vectorization vector values ​​of each channel of the large model, including: Based on the longitudinal vectorization vector of the first-layer channel and the low-bit fixed matrix, the first-layer channel of the large model is quantized and fine-tuned for pre-training to obtain a first-layer quantization result; Based on the longitudinal vectorization vector of the next layer of channels and the low-bit fixed matrix, the next layer of channels of the large model are quantized and fine-tuned for pre-training according to the first layer quantization result and the transverse vectorization vector to obtain the next quantization result; Iteratively perform quantization fine-tuning pre-training on the next layer of channels through the quantization results of the previous layer of channels of the large model until the last layer of channels of the large model, and obtain the horizontal vectorization vector value and the vertical vectorization vector value of each channel of the large model.

5. The method according to claim 4, characterized in that The quantization expression of the large model quantization fine-tuning pre-training channel of this layer is: The quantization expression of the next channel of the large model quantization fine-tuning pre-training is: W"0=(T0+ΔT)*W Where W is the quantization result of the channel in this layer; W”0 is the quantization result of the next layer of channels in this layer; S0 is the vertical quantization vector; T0 is the horizontal quantization vector; ΔS represents the gradient update of S0 obtained by fine-tuning training; ΔT represents the gradient update of T0 obtained by fine-tuning training; z0 represents an adjustment parameter; clamp represents a function that limits the value to [a, b].

6. A large model inference device based on quantization-aware fine-tuning, characterized in that: include: A structural transformation module is used to structurally transform the original parameter matrix of the large model and determine the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix corresponding to the original parameter matrix of each channel; A pre-training module, used to perform parameter quantization fine-tuning pre-training on the large model layer by layer based on the horizontal vectorization vector, the vertical vectorization vector and the low-bit fixed matrix, and obtain the horizontal vectorization vector value and the vertical vectorization vector value of each channel of the large model; A determination module, used to determine a deployment parameter matrix of the large model deployment reasoning according to the horizontal vectorization vector value, the vertical vectorization vector value and the low-bit fixed matrix of each channel of the large model; The reasoning module is used to use a large model deployed with the deployment parameter matrix to perform reasoning analysis on the input data to obtain the reasoning results of the input data.

7. The device according to claim 6, characterized in that Structural transformation module, including: The decomposition submodule is used to decompose the high-precision original parameter matrix W of each channel of the large model horizontally and vertically, and decompose it into the high-precision horizontal vectorization vector T0, the high-precision vertical vectorization vector S0 and the low-bit fixed matrix W0', wherein the original parameter matrix W of each channel ≈ S0*W0'.

8. The device according to claim 6, characterized in that The bit width of parameter quantization fine-tuning pre-training is 4.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to execute the method according to any one of claims 1 to 5.

10. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Transform large model reasoning method and device, computer equipment and storage medium

    CN116992965A

  • Prediction method and device based on low-rank quantization large model, electronic equipment, storage medium and computer program product

    CN118886453A