Deployment methods, electronic devices, and readable storage media for pre-trained language models
By generating sparse transformation matrices through the subsampling random Hadamard transform algorithm and combining it with the Hessian matrix adaptive rounding operation, the problems of high memory consumption and low quantization efficiency in the deployment of pre-trained language models are solved, achieving more efficient deployment and inference.
Patent Information
- Application Number
- CN202511152269.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-08-18
AI Technical Summary
Existing technologies suffer from high memory consumption, large inference latency, and low efficiency in the quantization process when deploying pre-trained language models. In particular, the quantization methods for incoherent processing result in large matrix operations, leading to low deployment efficiency.
A sparse transformation matrix is generated using the subsampled random Hadamard transform algorithm. The initial weight matrix is then incoherently processed. The orthogonal matrix is replaced by the sparse transformation matrix, and adaptive rounding is performed in conjunction with the Hessian matrix to generate the quantized weight matrix and deploy the target language model.
It significantly reduces the computational cost of matrix operations, improves the quantization and deployment efficiency of pre-trained language models, reduces GPU memory usage, and increases the inference speed of the model.
Smart Images

Figure CN120725074B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of pre-trained language model technology, and more particularly to a method for deploying a pre-trained language model, an electronic device, and a readable storage medium. Background Technology
[0002] In related technologies, deploying large-scale pre-trained language models (LLMs) in industrial or other fields and using them for inference presents challenges such as high memory consumption and large inference latency. Model quantization can compress model parameters from high-bit floating-point representation to low-bit width. For example, compressing 16-bit floating-point (BF16) to the bit width of INT8 / INT4 can significantly reduce memory consumption and improve decoding speed. However, using low-bit representation for the weights after quantization of the pre-trained language model can lead to numerical error accumulation and distribution shift, resulting in decreased model accuracy.
[0003] Quantization with Incoherence Processing (QuIP) can measure the consistency of activation values through inconsistency and perform targeted quantization, thereby improving the accuracy of the quantized model. However, QuIP involves a large number of matrix operations, resulting in a long quantization process for pre-trained language models and ultimately low efficiency in deploying them. Summary of the Invention
[0004] This application provides a method for deploying a pre-trained language model, an electronic device, and a readable storage medium to at least address the problem of low efficiency in deploying pre-trained language models in related technologies.
[0005] This application provides a method for deploying a pre-trained language model, the method comprising:
[0006] Determine the initial weight matrix and Hessian matrix of the pre-trained language model to be quantized; wherein, the initial weight matrix is used to characterize the model information topology of the pre-trained language model, and the Hessian matrix is used to characterize the low-sensitivity weights in the initial weight matrix.
[0007] Generating sparse transformation matrices based on the subsampled random Hadamard transform algorithm;
[0008] The initial weight matrix is incoherently processed by a sparse transformation matrix to obtain the optimized weight matrix.
[0009] An adaptive rounding operation is performed based on the optimized weight matrix and Hessian matrix to obtain the quantized weight matrix.
[0010] The quantized target language model is determined based on the quantization weight matrix, and the target language model is sent to the device to be deployed so as to deploy the target language model on the device.
[0011] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described methods for deploying a pre-trained language model when executing the computer program.
[0012] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described methods for deploying a pre-trained language model.
[0013] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described methods for deploying a stored pre-trained language model.
[0014] Through this application, because during the quantization process of the pre-trained language model, when performing incoherent processing on the initial weight matrix, the sparse transformation matrix generated by the subsampled random Hadamard transform algorithm replaces the original orthogonal matrix transformation, the computational cost of matrix operations can be reduced from... Reduced to , where m and n are the row and column dimensions of the matrix, respectively. Therefore, this significantly reduces the computational cost of matrix operations, solves the technical problems of long quantization processes and low quantization efficiency in pre-trained language models, and achieves the technical effect of improving the efficiency of deploying pre-trained language models. Attached Figure Description
[0015] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 The diagram illustrates the architecture of a pre-trained language model according to some embodiments of this application;
[0017] Figure 2 The diagram illustrates the architecture of the encoder module of a pre-trained language model according to some embodiments of this application.
[0018] Figure 3 A schematic diagram of the architecture of the feedforward network activation function of a pre-trained language model according to some embodiments of this application is shown;
[0019] Figure 4A flowchart illustrating a method for deploying a pre-trained language model according to some embodiments of this application is shown;
[0020] Figure 5 A schematic diagram of a butterfly matrix according to some embodiments of this application is shown;
[0021] Figure 6 The diagram shows a structural block diagram of a deployment apparatus for a pre-trained language model according to some embodiments of this application;
[0022] Figure 7 Structural block diagrams of electronic devices according to some embodiments of this application are shown. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0024] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0025] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] The specific application environment architecture or specific hardware architecture on which the deployment method of the combined pre-trained language model depends is described here.
[0027] Current pre-trained language models suffer from high memory consumption and large inference latency. Therefore, quantization is needed to compress model parameters from high-order floating-point to low-order bit width to reduce memory consumption and improve decoding speed. However, quantization methods such as QuIP involve a large number of matrix operations, resulting in a long quantization process and ultimately low efficiency in deploying pre-trained language models for inference.
[0028] To address the technical problem of low efficiency in quantization processing of stored pre-trained language models in the aforementioned related technologies, this application proposes a method for deploying stored pre-trained language models. The following is a further detailed description of this application in conjunction with the accompanying drawings.
[0029] In some embodiments of this application, exemplarily, Figure 1 The following are schematic diagrams illustrating the architecture of pre-trained language models according to some embodiments of this application, such as... Figure 1 As shown, a pre-trained language model includes input, output, and multiple decoder blocks. For example, Figure 1 The pre-trained language model shown is the Llama2 13B inference network model.
[0030] For example, Figure 2 The following are schematic diagrams illustrating the architecture of the encoder module of a pre-trained language model according to some embodiments of this application, such as... Figure 2 As shown, the encoder module includes input, output, softmax activation function, word vector representation, root mean square normalization (RMSNorm), multi-head attention mechanism (MHA), and feedforward network SwishGLU activation function.
[0031] The process begins by processing the input data using a tokenizer, followed by embedding to create vectorized data. This data then undergoes a series of processing steps, including RMSNorm, MHA, RMSNorm, and FeedForwardSwishGLU. Finally, after 40 layers of data processing, the data is decoded using RMSNorm, Embedding, and Softmax to extract the characters.
[0032] The formula for RMSNorm is:
[0033] ;
[0034] ;
[0035] in, x For input, For feature dimension, The hyperparameters of RMSNorm are determined after the model is trained. The value, Dimensions and same. This represents element-wise multiplication, i.e., multiplying element by element. for x The estimated value.
[0036] The Softmax activation function is:
[0037] ;
[0038] in, z = ( z 1. z 2…… z i , z j Let be the input vector, where is z i Specific elements, z j To iterate over the elements, e The base of the natural constant is... K This represents the total number of vector dimensions.
[0039] The formula for multi-head attention mechanism (MHA) is:
[0040] ;
[0041] ;
[0042] in, They are respectively The weights of each attention head, These are the weights of the linear layer.
[0043] Figure 3 The following is a schematic diagram of the architecture of the feedforward network activation function of a pre-trained language model according to some embodiments of this application, such as... Figure 3 As shown, FeedForwardSwishGLU includes input, output, up-line linear layer, down-line linear layer, gate linear layer, and self-gated activation function (Swish).
[0044] The data output from RMSNorm serves as the input to FeedForwardSwishGLU, passes through Gate_Linear→Swish and Up_Linear respectively, and then undergoes Hadamard Product processing before finally passing through Down_Linear as the output and being sent to the next unit.
[0045] The Swish function is as follows:
[0046] ;
[0047] in, x For input, To learn good parameters, σ (•) represents the sigmoid function. e is the base of the natural constant.
[0048] The Swish function can still generate non-zero gradients when the input is negative, avoiding the "neuron death" problem of traditional activation functions (such as ReLU activation function) when the input is negative.
[0049] For example, Gate_Linear, Up_Linear, and Down_Linear are all multilayer perceptron (MLP) networks.
[0050] Taking the execution of an inference task after quantizing the aforementioned pre-trained language model as an example, the deployment method of the pre-trained language model proposed in this application embodiment is illustrated as follows:
[0051] In some embodiments of this application, a method for deploying a pre-trained language model is proposed. Figure 4 Flowcharts illustrating deployment methods of pre-trained language models according to some embodiments of this application are shown, such as... Figure 4 As shown, the method includes:
[0052] Step 402: Determine the initial weight matrix and Hessian matrix of the initial pre-trained language model to be quantized; wherein, the initial weight matrix is used to characterize the model information topology of the pre-trained language model, and the Hessian matrix is used to characterize the low-sensitivity weights in the initial weight matrix.
[0053] In this embodiment of the application, exemplarily, the initial pre-trained language model to be quantized is as follows: Figure 1 The pre-trained language model is shown. The initial weight matrix is set to W, and the Hessian matrix is set to H1.
[0054] Step 404: Generate a sparse transformation matrix based on the subsampled random Hadamard transform algorithm.
[0055] In this embodiment, during the quantization of the weight matrix, a similarity transformation is performed on the original weights. This similarity transformation does not change the eigenvalues and trace of the matrix, but each matrix similarity transformation generates a... and With the matrix dimension m andn As the number of elements increases, the processing time for incoherence processing becomes longer, which is the main reason for the low quantization efficiency of pre-trained language models.
[0056] To address this issue, considering the special characteristics of randomized orthogonal matrices, this application generates a sparse transformation matrix R using the Subsampled Randomized Hadamard Transform (SRHT) algorithm, which is then used to replace the randomized orthogonal matrix. Since the proportion of non-zero elements in the sparse transformation matrix is low, typically less than 10%, the computational cost during matrix operations can be reduced to less than one-tenth that of a dense matrix. Furthermore, because the sparse transformation matrix only stores the positions and values of non-zero elements, it significantly reduces memory usage.
[0057] Step 406: Perform incoherent processing on the initial weight matrix using the sparse transformation matrix to obtain the processed optimized weight matrix.
[0058] In this embodiment, the initial weight matrix is incoherently processed by a sparse transformation matrix to obtain an optimized weight matrix. The optimized weight matrix satisfies sub-Gaussianity, which can reduce outliers and thus reduce the difficulty of subsequent quantization.
[0059] Step 408: Perform adaptive rounding operation based on the optimized weight matrix and the Hessian matrix to obtain the quantized weight matrix.
[0060] In this embodiment, during the adaptive rounding step, the optimized weight matrix is derived from the output of the incoherent processing, which is a more uniformly distributed floating-point weight matrix. The Hessian matrix provides second-order derivative information to guide the rounding direction and minimize quantization error.
[0061] The quantized weight matrix can be obtained by mapping the floating-point weight matrix to integers.
[0062] Step 410: Determine the quantized target language model based on the quantization weight matrix, and send the target language model to the device to be deployed so as to deploy the target language model on the device to be deployed.
[0063] In this embodiment of the application, by performing weight quantization on each transformer layer in the pre-trained language model in sequence in the manner described above, and saving the quantized weight matrix, the quantized target language model can be obtained.
[0064] For example, when deploying the quantized target language model to the device to be deployed, the quantized target language model is sent to the device to be deployed and loaded into the Vectorized Large Language Model Inference / Serving System (VLLM) of the device to be deployed. The quantized model inference result can be obtained by accessing it through the corresponding port.
[0065] For example, the deployment command is:
[0066] python3 -m vllm.entrypoints.openai.api_server --model {llama2 13Bquantized model path} --swap-space 16 --disable-log-requests --port 8081
[0067] At this point, the quantized model inference results can be obtained by accessing port 8081.
[0068] For example, the device to be deployed includes a computing subsystem, a storage subsystem, and a network subsystem. The computing subsystem may be a server or a terminal.
[0069] For example, a server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services such as cloud servers, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0070] For example, the terminal can be an electronic device such as a mobile phone, tablet computer, vehicle terminal (vehicle system), wearable device, personal computer (PC), unmanned reservation terminal, smart speaker, etc.
[0071] For example, the computing subsystem includes an artificial intelligence (AI) acceleration cluster. The storage subsystem can employ a multi-tiered, hierarchical design. The networking subsystem enables low-latency interconnection between multiple computing nodes within the AI acceleration cluster.
[0072] For example, the computing subsystem includes a graphics processing unit (GPU) matrix, the storage subsystem includes GPU memory, and the network subsystem includes an intra-node interconnect architecture.
[0073] For example, the computing subsystem includes a cluster of Tensor Processing Units (TPUs), the storage subsystem includes a Non-Volatile Memory Express Solid State Drive (NVMe SSD), and the network subsystem includes an intra-node interconnect architecture.
[0074] For example, in scenarios where the target language model is deployed on terminal devices such as mobile phones and wearable devices, the target language model receives the user's voice audio as data input, and performs vectorization processing on the voice audio to obtain the model input. This model input is fed into the target language model, where it performs inference through multiple neural network layers, and finally outputs artificial speech, realizing natural dialogue with the user and achieving artificial intelligence dialogue.
[0075] For example, in a scenario where the target language model is deployed in an industrial environment, the target language model receives workpiece images captured by an image sensor as data input, and performs vectorization processing on the workpiece images to obtain the model input. This model input is fed into the target language model, where inference is performed through multiple neural network layers, ultimately outputting the quality inspection results of the workpiece, thus achieving automated quality inspection.
[0076] During the quantization process of the pre-trained language model, when performing incoherent processing on the initial weight matrix, the sparse transformation matrix generated by the subsampled random Hadamard transform algorithm replaces the original orthogonal matrix transformation, thus reducing the computational cost of matrix operations. Reduced to , where m and n are the row and column dimensions of the matrix, respectively. Therefore, this significantly reduces the computational cost of matrix operations, solves the technical problems of long quantization processes and low quantization efficiency in pre-trained language models, and achieves the technical effect of improving the efficiency of deploying pre-trained language models.
[0077] In some embodiments of this application, a sparse transformation matrix is generated based on the subsampled random Hadamard transform algorithm, including: generating a diagonal matrix and a random sampling matrix; wherein the diagonal elements of the diagonal matrix are 1 or -1, and the other elements of the diagonal matrix are 0; the random sampling matrix is obtained by randomly sampling multiple basis vectors from a standard basis vector set; the input sequence is grouped by a fast Walsh transform to obtain a butterfly matrix, and a Hadamard matrix is generated by a butterfly algorithm; and a sparse transformation matrix is generated based on the diagonal matrix, the random sampling matrix, and the Hadamard matrix.
[0078] In this embodiment, the Fast Walsh-Hadamard Transform (FWHT) is applied to the operation of the SRHT matrix, by transforming the input sequence... W ( u Grouping matrix operations can reduce the computational load and improve efficiency.
[0079] For example, the specific relationship of FWHT is as follows:
[0080] ;
[0081] in, For the Walsh-Hadamard transform results in the index The value at that location, The transformation result of the even-indexed subsequence of the input sequence. The transformation result of the odd-indexed subsequence of the input sequence. N The total length of the input sequence.
[0082] Taking a matrix with dimensions of 8×8 as an example, then the following conditions are met:
[0083] ;
[0084] in, H n Let the Hadamard matrix be of dimension n×n. I n It is an identity matrix of dimension n×n, with diagonal elements being 1 and off-diagonal elements being 0;
[0085] , , ;
[0086] Then we can get:
[0087] ;
[0088] make:
[0089]
[0090] Then we can get the following Figure 5 The butterfly matrix shown.
[0091] For example, the tools required for the subsampled random Hadamard transform algorithm include a diagonal matrix, a random sampling matrix, and a Hadamard matrix. Specifically, a random diagonal matrix D is constructed, with elements following a Bernoulli distribution (equal probability ±1), whose diagonal elements are between {-1, 1}, to break data correlation.
[0092] Incoherent processing achieves "decorrelatedness" of the weight distribution through orthogonal transformation. Generating the Hadamard matrix H using the Fast Walsh Transform reduces the complexity of the Hadamard matrix, thereby improving the efficiency of matrix operations.
[0093] After obtaining the diagonal matrix D, the random sampling matrix S, and the Hadamard matrix H, the sparse transformation matrix R can be generated by the subsampled random Hadamard transform algorithm.
[0094] This embodiment generates the Hadamard matrix using the Fast Walsh Transform. Through the aforementioned butterfly algorithm, the computational complexity of matrix operations can be reduced from... Reduced to This can reduce the complexity of matrix operations and the computation time required for matrix operations, thereby accelerating the entire quantization process and improving the quantization efficiency of pre-trained language models.
[0095] In some embodiments of this application, a sparse transformation matrix is generated based on a diagonal matrix, a random sampling matrix, and a Hadamard matrix, including:
[0096] The sparse transformation matrix is generated using the following formula:
[0097] ;
[0098] Where R is the sparse transformation matrix, d is the column dimension of the initial weight matrix, r is the dimension of the sparse transformation matrix, D is the diagonal matrix, H is the Hadamard matrix, and S is the random sampling matrix.
[0099] In the embodiments of this application, for SRHT definition (For matrix) ,but For matrix ,but The matrix is shown in the formula above.
[0100] in, It is a diagonal matrix whose diagonal elements are in between, It is an orthogonal Walsh-Hadamard matrix.
[0101] The Walsh-Hadamard matrix is recursively defined as follows:
[0102] ;
[0103] in, , .
[0104] For a random sampling matrix, ,in .
[0105] in, Q For the process After calculation .
[0106] This application generates sparse transformation matrices using a subsampled random Hadamard transform algorithm. By replacing random orthogonal matrices with sparse transformation matrices, the computational load and memory usage during matrix operations can be significantly reduced.
[0107] In some embodiments of this application, the initial weight matrix is incoherently processed by a sparse transformation matrix to obtain the optimized weight matrix, including: performing a similarity transformation on the initial weight matrix by a sparse transformation matrix to obtain the optimized weight matrix.
[0108] In this embodiment, the initial weight matrix is incoherently processed by a sparse transformation matrix, which can optimize the weight distribution, reduce outliers, and thus improve the subsequent quantization effect.
[0109] For example, the "decorrelated" weight distribution is achieved through orthogonal transformation. A fast transformation is performed on each column vector v of the initial weight matrix, and v is split into even / odd index sub-vectors through recursive divide-and-conquer, followed by butterfly operations to output the frequency domain representation under the Hadamard basis, which is the optimized weight matrix.
[0110] This application uses a sparse transformation matrix to perform a similarity transformation on the initial weight matrix, which can significantly reduce the complexity of the weight matrix and thus reduce the computational cost of matrix operations in the subsequent quantization process.
[0111] In some embodiments of this application, before performing incoherent processing on the initial weight matrix using a sparse transformation matrix, the method further includes: padding the initial weight matrix to ensure that the dimension of the initial weight matrix satisfies 2. k k is a positive integer.
[0112] In this embodiment of the application, if the dimension n of the initial weight matrix does not satisfy 2 kThen, fill the initial weight matrix with zeros until the dimension d of the filled initial weight matrix is 2. k For example, if the initial weight matrix has a dimension n=1000, then the initial weight matrix is padded with zeros, and the dimension d of the padded initial weight matrix is 1024.
[0113] By addressing the dimension that does not satisfy 2 k The initial weight matrix is padded to ensure its dimension is 2. k This meets the requirements of the subsequent butterfly algorithm, thus enabling computational acceleration of quantization processing through the butterfly algorithm.
[0114] In some embodiments of this application, an adaptive rounding operation is performed based on the optimized weight matrix and the Hessian matrix to obtain the quantized weight matrix, including: determining a proxy Hessian matrix of the optimized weight matrix based on the optimized weight matrix and the Hessian matrix; performing iterative processing based on the proxy Hessian matrix and the proxy loss function to obtain the target input activation value; and determining the quantized weight matrix based on the target input activation value, the optimized weight matrix, and the sparse transformation matrix.
[0115] In the embodiments of this application, the purpose of adaptive rounding operation is to minimize quantization error by dynamically adjusting the rounding direction (up / down).
[0116] In the incoherent processing step, the optimized weight matrix and sparse transformation matrix are obtained. Through Hessian matrix analysis, the proxy Hessian matrix of the optimized weight matrix is determined, thereby minimizing the proxy loss.
[0117] After obtaining the surrogate Hessian matrix, the initial input activation value is used as input for iterative processing according to the selected surrogate loss function. The input activation value is continuously iterated until the loss function converges, and the target input activation value that minimizes the quantization error is obtained.
[0118] After obtaining the target input activation value, the final quantization weight matrix is obtained by combining the optimized weight matrix and the sparse transformation matrix.
[0119] This application iterates over the input activation values using a surrogate loss function to obtain convergent rounding variables, thereby finding the target input activation values that minimize quantization error and improving the inference accuracy of the quantized target language model.
[0120] In some embodiments of this application, the proxy loss function is:
[0121] ;
[0122] in, For the proxy loss function, Let W be the quantized weight matrix, H be the surrogate Hessian matrix, Ex[·] be the expected value operation, x be the input activation value, and tr(·) be the trace of the matrix.
[0123] In this embodiment of the application, the proxy Hessian matrix H can be decomposed using LDL decomposition as follows:
[0124] ;
[0125] in, It is a unit upper triangular matrix, and It is an identity matrix. Where, the matrix... It has the following properties The proof is as follows:
[0126] make Feature decomposition into By observing the non-correlation, we can obtain the following:
[0127] ;
[0128] in, Let k be a column vector with 1s in the k-th row and 0s in the other rows. for The eigenvalues in the i-th row and i-th column, Incoherence (defined as the degree of incoherence for a matrix) Characteristic decomposition For each and ,if ,but (where n is the degree of incoherence), and n is the number of rows and columns of matrix H.
[0129] Then, ;
[0130] in ,in , where is the covariance of the input activation values.
[0131] Assumption Then we have:
[0132] ;
[0133] Therefore, we can obtain .
[0134] This application improves the inference accuracy of the quantized target language model by reasonably setting the proxy loss function and iterating the input activation values through the proxy loss function.
[0135] In some embodiments of this application, determining the quantization weight matrix based on the target input activation value, the optimized weight matrix, and the sparse transformation matrix includes: generating an integer weight matrix based on the target input activation value, the optimized weight matrix, and the sparse transformation matrix; and performing an inverse sparse transformation on the integer weight matrix to obtain the quantization weight matrix.
[0136] In this embodiment, the target input activation value is the optimal rounding variable. Let the target input activation value be v, the optimized weight matrix be W', and the sparse transformation matrix be R, then the integer weight matrix Q can be generated:
[0137] ;
[0138] Next, an inverse sparse transformation is performed on the integer weight matrix Q to restore its original dimensions, thus obtaining the quantization matrix weights W: W = RQR T .
[0139] This application embodiment replaces the random orthogonal matrix with a sparse transformation matrix during the quantization process of the weight matrix, enabling the quantization weight matrix to be stored in a sparse storage format, thereby reducing the GPU memory usage.
[0140] In some embodiments of this application, adaptive rounding operations are performed based on the optimized weight matrix and the Hessian matrix, including: performing adaptive rounding operations based on the optimized weight matrix, the Hessian matrix, a rounding algorithm, and a random rounding algorithm.
[0141] In this embodiment of the application, the QuIP quantization algorithm includes:
[0142] Required: , , , , , ;
[0143] 1: ;
[0144] 2: ;
[0145] 3: for ;
[0146] 4: return .
[0147] in, That is, the rounding algorithm. That is, the random rounding algorithm.
[0148] The first step of the QuIP quantization algorithm is preprocessing and adaptive scaling, which is achieved by calling sub-algorithms ( Complete the weight distribution optimization.
[0149] The second step is Hessian matrix decomposition, which provides a theoretical guarantee for quantization error control.
[0150] The third step is column vector quantization, which processes each column vector of the initial weight matrix independently. This includes error compensation, quantization function application, and range constraints. The quantization function employs the rounding algorithm and random rounding algorithm described above.
[0151] The fourth step is the inverse transformation and output, which is achieved by calling a sub-algorithm ( The original spatial structure is restored and the final quantization weight matrix is output.
[0152] The embodiments of this application use rounding and random rounding algorithms as quantization algorithms, which have the advantages of good determinism, low hardware overhead and better theoretical boundaries.
[0153] In some embodiments of this application, determining the initial weight matrix and Hessian matrix of the initial pre-trained language model to be quantized includes: running the initial pre-trained language model to be quantized on a calibration dataset to obtain the Hessian matrix; and extracting the initial weight matrix.
[0154] In this embodiment, the calibration dataset is exemplarily a C4-validation dataset. The key to the calibration dataset is selecting samples that represent the model's application scenario. These samples are used to calibrate the model parameters during the weight generation process to reduce performance loss. Using the calibration dataset as input to the network yields the Hessian matrix of the initial pre-trained language model.
[0155] This application uses the Hessian matrix for adaptive rounding, which enables diagonal approximate weighted quantization error, thereby reducing the inference error of the quantized target language model and improving the inference accuracy of the quantized target language model.
[0156] In some embodiments of this application, extracting the initial weight matrix includes: loading an initial pre-trained language model; wherein the initial pre-trained language model includes multiple transformer layers; and extracting the initial weight matrix through the application programming interface of the model framework of the initial pre-trained language model.
[0157] In this embodiment, when extracting the initial weight matrix of the pre-trained language model, the trained model file, such as .pth, .ckpt, .h5, etc., is read first. Then, the initial weight matrix is extracted through the application programming interface (API) of the model framework of the pre-trained language model.
[0158] For example, for a PyTorch model, you can use the `model.state_dict()` command to return an ordered dictionary containing all layer weight names and their corresponding tensors. Then, you can use the `model.named_parameters()` command to return a generator that iterates through and outputs layer names and weight tensors.
[0159] This application utilizes a model framework based on an initial pre-trained language model to obtain the initial weight matrix via the model framework's API, thereby reducing the amount of operations required to obtain the weight matrix and improving overall quantization efficiency.
[0160] In some embodiments of this application, the device to be deployed includes a language model inference and service framework, which includes a front-end interface and an inference engine. After sending the target language model to the device to be deployed, the method further includes: deploying the target language model in the language model inference and service framework; receiving the data to be inferred through the front-end interface, and sending the data to be inferred to the target language model through the inference engine to obtain the inference result corresponding to the data to be inferred.
[0161] In this embodiment, the device to be deployed includes a language model inference and service framework, namely vLLM. The language model inference and service framework includes a front-end interface (API Server) and an inference engine (LLMEngine). The front-end interface can provide a multi-compatible program interface, supporting multiple input formats such as text and multimodal.
[0162] After receiving the user-inputted data to be inferred, the front-end interface forwards the data to the inference engine, which then coordinates internal components to process the inference request. For example, the inference engine converts the user's prompt into a request object, containing the segmented prompt token. The prompt is added to a scheduling queue and distributed to worker nodes. The worker nodes invoke the deployed target language model to perform the actual inference and output the corresponding inference result.
[0163] This application deploys the quantized target language model through language model inference and service framework, and performs inference tasks, which can achieve low latency and high throughput inference, while providing "out-of-the-box" API services to improve the deployment capability and efficiency of the target language model.
[0164] Embodiments of this application also provide a deployment apparatus for a pre-trained language model. Figure 6 The following is a structural block diagram of a deployment apparatus for a pre-trained language model according to some embodiments of this application, such as... Figure 6 As shown, the deployment device 600 includes:
[0165] The determination module 602 is used to determine the initial weight matrix and Hessian matrix of the pre-trained language model to be quantized; wherein, the initial weight matrix is used to represent the model information topology of the pre-trained language model, and the Hessian matrix is used to represent the low-sensitivity weights in the initial weight matrix; the generation module 604 is used to generate a sparse transformation matrix based on the subsampled random Hadamard transform algorithm; the processing module 606 is used to perform incoherent processing on the initial weight matrix through the sparse transformation matrix to obtain the processed optimized weight matrix; the quantization module 608 is used to perform adaptive rounding operation based on the optimized weight matrix and the Hessian matrix to obtain the quantized weight matrix; the determination module 602 is also used to determine the quantized target language model according to the quantized weight matrix, and send the target language model to the device to be deployed so as to deploy the target language model on the device to be deployed.
[0166] During the quantization process of the pre-trained language model, when performing incoherent processing on the initial weight matrix, the sparse transformation matrix generated by the subsampled random Hadamard transform algorithm replaces the original orthogonal matrix transformation, thus reducing the computational cost of matrix operations. Reduced to , where m and n are the row and column dimensions of the matrix, respectively. Therefore, this significantly reduces the computational cost of matrix operations, solving the technical problems of long quantization processes and low quantization efficiency in pre-trained language models, thus achieving the technical effect of improving the quantization efficiency of pre-trained language models.
[0167] For a description of the features in the embodiment corresponding to the deployment device of the pre-trained language model, please refer to the relevant description in the embodiment corresponding to the deployment method of the pre-trained language model, which will not be repeated here.
[0168] Embodiments of this application also provide an electronic device. Figure 7 Structural block diagrams of electronic devices according to some embodiments of this application are shown, such as Figure 7 As shown, the electronic device 700 includes a memory 702 and a processor 704. The memory 702 stores a computer program, and the processor 704 is configured to run the computer program to perform the steps in any of the above-described pre-trained language model deployment method embodiments.
[0169] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the deployment method for pre-trained language models at runtime.
[0170] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0171] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in the deployment method embodiments of any of the pre-trained language models described above.
[0172] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described pre-trained language model deployment method embodiments.
[0173] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0174] The foregoing has provided a detailed description of the deployment method, apparatus, and electronic device for pre-trained language models provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for deploying a pre-trained language model, characterized in that, The method includes: Determine the initial weight matrix and Hessian matrix of the initial pre-trained language model to be quantized; wherein the initial weight matrix is used to characterize the model information topology of the pre-trained language model, and the Hessian matrix is used to characterize the low-sensitivity weights in the initial weight matrix; Generating sparse transformation matrices based on the subsampled random Hadamard transform algorithm; The initial weight matrix is incoherently processed by the sparse transformation matrix to obtain the optimized weight matrix. An adaptive rounding operation is performed based on the optimized weight matrix and the Hessian matrix to obtain the quantized weight matrix. The quantized target language model is determined based on the quantization weight matrix, and the target language model is sent to the device to be deployed so as to deploy the target language model on the device to be deployed.
2. The deployment method according to claim 1, characterized in that, The subsampled random Hadamard transform algorithm for generating sparse transform matrices includes: Generate a diagonal matrix and a random sampling matrix; wherein the diagonal elements of the diagonal matrix are 1 or -1, and the other elements of the diagonal matrix are 0; the random sampling matrix is obtained by randomly sampling multiple basis vectors from a standard basis vector set; The input sequence is grouped using the Fast Walsh Transform to obtain the butterfly matrix, and the Hadamard matrix is generated using the butterfly algorithm. The sparse transformation matrix is generated based on the diagonal matrix, the random sampling matrix, and the Hadamard matrix.
3. The deployment method according to claim 2, characterized in that, The step of generating the sparse transformation matrix based on the diagonal matrix, the random sampling matrix, and the Hadamard matrix includes: The sparse transformation matrix is generated using the following formula: ; Where R is the sparse transformation matrix, d is the column dimension of the initial weight matrix, r is the dimension of the sparse transformation matrix, D is the diagonal matrix, H is the Hadamard matrix, and S is the random sampling matrix.
4. The deployment method according to any one of claims 1 to 3, characterized in that, The step of performing incoherent processing on the initial weight matrix using the sparse transformation matrix to obtain the processed optimized weight matrix includes: The optimized weight matrix is obtained by performing a similarity transformation on the initial weight matrix using the sparse transformation matrix.
5. The deployment method according to claim 4, characterized in that, Before performing incoherent processing on the initial weight matrix using the sparse transformation matrix, the method further includes: The initial weight matrix is padded to ensure that its dimension is 2. k k is a positive integer.
6. The deployment method according to any one of claims 1 to 3, characterized in that, The step of performing adaptive rounding based on the optimized weight matrix and the Hessian matrix to obtain the quantized weight matrix includes: Based on the optimized weight matrix and the Hessian matrix, determine the proxy Hessian matrix of the optimized weight matrix; The target input activation value is obtained by iterative processing based on the proxy Hessian matrix and proxy loss function. The quantization weight matrix is determined based on the target input activation value, the optimized weight matrix, and the sparse transformation matrix.
7. The deployment method according to claim 6, characterized in that, The proxy loss function is: ; in, Let be the proxy loss function. Let W be the quantization weight matrix, H be the initial weight matrix, H be the surrogate Hessian matrix, Ex[·] be the expected value operation, x be the input activation value, and tr(·) be the trace of the matrix.
8. The deployment method according to claim 6, characterized in that, The step of determining the quantization weight matrix based on the target input activation value, the optimized weight matrix, and the sparse transformation matrix includes: An integer weight matrix is generated based on the target input activation value, the optimized weight matrix, and the sparse transformation matrix. The integer weight matrix is subjected to an inverse sparse transformation to obtain the quantized weight matrix.
9. The deployment method according to any one of claims 1 to 3, characterized in that, The adaptive rounding operation based on the optimized weight matrix and the Hessian matrix includes: The adaptive rounding operation is performed based on the optimized weight matrix, the Hessian matrix, the rounding algorithm, and the random rounding algorithm.
10. The deployment method according to any one of claims 1 to 3, characterized in that, The determination of the initial weight matrix and Hessian matrix of the initial pre-trained language model to be quantized includes: The Hessian matrix is obtained by running the initial pre-trained language model to be quantized on the calibration dataset; and Extract the initial weight matrix.
11. The deployment method according to claim 10, characterized in that, The extraction of the initial weight matrix includes: Load the initial pre-trained language model; wherein the initial pre-trained language model includes multiple transformer layers; The initial weight matrix is extracted through the application programming interface of the model framework of the initial pre-trained language model.
12. The deployment method according to any one of claims 1 to 3, characterized in that, The device to be deployed includes a language model inference and service framework, which includes a front-end interface and an inference engine. After sending the target language model to the device to be deployed, the method further includes: Deploy the target language model within the language model inference and service framework; The front-end interface receives the data to be inferred, and the inference engine sends the data to be inferred to the target language model to obtain the inference result corresponding to the data to be inferred.
13. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the deployment method as described in any one of claims 1 to 12.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the deployment method as described in any one of claims 1 to 12.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the deployment method as described in any one of claims 1 to 12.
Citation Information
Patent Citations
Sparse training method of pre-training language model and deep language computing system
CN115222039A
Quantization method, reasoning method and equipment of industry large model and storage medium
CN120181146A