Iterative quantitative perceptual training method for large language model end-side deployment

Through the iterative quantization perception training method, dynamic allocation of weight quantization ratios and selective preservation of original parameters is solved, which solves the problems of large hardware resources consumption and poor quantization effects of generative large language models, and realizes efficient end-side deployment and accurate decoding.

CN120579587AActive Publication Date: 2025-09-02TIANJIN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511081704.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-09-02
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

In the process of parameter quantization of generative large language models, the prior art has high hardware resource consumption, poor quantization effect and cannot be directly deployed on end-side devices with low hardware performance, resulting in high deployment costs and risk of user privacy leakage.

Method used

Iterative quantization perception training method is adopted to dynamically allocate weights and quantize the proportions through the proportional scheduler, and selectively retain the original parameters using the Boolean mask matrix, combining sparse calculation optimization and quantized parameter update strategies to achieve end-side deployment.

Benefits of technology

It reduces the performance loss of end-side devices due to low-precision calculations, maintains the decoding accuracy of the generative large language model on the end-side, improves inference throughput, and avoids cross-platform compatibility issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579587A_ABST
    Figure CN120579587A_ABST
Patent Text Reader

Abstract

The invention provides an iterative quantitative perceptual training method for large language model end-side deployment, which can be applied to the technical field of large language models. The method comprises the following steps: dynamically distributing a weight quantification proportion according to a training stage through a proportion scheduler, avoiding excessive compression of key parameters, and reducing performance loss of end-side equipment caused by low-precision calculation; original parameters are selectively reserved through a Boolean type mask matrix, quantization errors of key weights in a self-attention layer are reduced, and the accuracy of end-side decoding of the generative large language model is maintained; through a sparse parameter structure generated by a mask matrix, sparse calculation optimization of a hardware accelerator can be triggered, and the reasoning throughput is improved; besides, according to the multi-stage quantization parameter updating strategy provided by the invention, the quantization granularity is allowed to be adjusted for different hardware, and the cross-platform compatibility problem caused by traditional one-time quantization is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models, and in particular to an iterative quantization-aware training method for terminal-side deployment of large language models. Background Art

[0002] With the development of large-scale generative language models, the number of model parameters has exploded. These parameters are stored using FP16 floating-point numbers (half-precision floating-point numbers), meaning each parameter occupies at least two bytes. Fully deploying large-scale generative language models requires high-performance server-side hardware, such as high bandwidth, large amounts of high-performance memory, and a large number of high-performance graphics cards. This significantly increases the deployment and application costs of large-scale generative language models. Furthermore, due to the high hardware performance required by large-scale generative language models, pre-trained large-scale generative language models cannot be directly deployed on the client side, where hardware performance is significantly lower than that of the server. Therefore, users typically interact with large-scale generative language models deployed on the server over the network, which reduces user convenience, especially when network traffic is congested. Furthermore, server-side deployment of models poses the risk of user privacy leakage. Therefore, large-scale generative language models deployed on the server typically require parameter quantization technology to reduce the storage space occupied by model parameters. The quantized large-scale generative language models are then deployed on the client side. However, existing parameter quantization technologies still require a large amount of hardware resources in the process of parameter quantization of generative large language models, and existing parameter quantization technologies mainly focus on post-training quantization technology. The advantages of this technology are small training volume and fast quantization speed, but it also has problems such as poor quantization effect. Summary of the Invention

[0003] In view of the above problems, the present invention provides an iterative quantization-aware training method for end-side deployment of a large language model, which is used to solve at least one of the above technical problems.

[0004] According to a first aspect of the present invention, an iterative quantization-aware training method for on-device deployment of a large language model is provided, comprising:

[0005] Randomly select text data samples used in the current training phase from the text training dataset, perform initial quantization on all original parameters in the target large language model deployed on the server in the current training phase, and obtain the initial quantization parameter matrix;

[0006] The scale scheduler is used to determine the weight quantization ratio of the current training stage, and the weight mask function is used to generate the Boolean mask matrix of the current training stage;

[0007] Using the weight quantization ratio and the Boolean mask matrix, some parameters in the initial quantization parameter matrix are replaced with the original parameters of the target large language model to obtain the quantization parameter matrix;

[0008] The target large language model with a quantized parameter matrix encodes the text data samples of the current training stage based on the self-attention mechanism and decodes them based on the cross-attention mechanism to obtain the text data processing results, and uses the text data processing results to obtain the loss value of the current training stage;

[0009] Update the parameters of the target large language model in the current training phase using the quantization parameter matrix, the loss value of the current training phase, and the preset learning rate;

[0010] Repeat the operations of the current training stage for each training stage until the preset training conditions are met, obtain the parameter-quantized large language model, and deploy the parameter-quantized target large language model to the client, where the parameter-quantized large language model is used to process the client's text data.

[0011] According to an embodiment of the present invention, the above-mentioned iterative quantization-aware training method for device-side deployment of a large language model further includes:

[0012] The original parameters of the pre-trained large language model are limited to a certain value range through parameter truncation to obtain the target large language model.

[0013] The original parameters in the pre-trained large language model are initially scored through the parameter scoring operation to obtain the initial score matrix, and the weight quantization ratio used in the first round of training is initialized.

[0014] According to an embodiment of the present invention, the original parameters of the target large language model and the original parameters of the pre-trained large language model have the same floating point precision;

[0015] The initial score matrix has the same shape as the parameter matrix of the pre-trained large language model and is used to screen the quantized model parameters during the parameter quantization process.

[0016] According to an embodiment of the present invention, the above-mentioned use of the proportional scheduler to determine the weight quantization ratio of the current training stage includes:

[0017] Based on the total number of stages used in training, the proportional scheduler is used to determine the weight quantization ratio of the current training stage in a linearly increasing manner.

[0018] According to an embodiment of the present invention, the Boolean mask matrix and the initial quantization parameter matrix have the same shape;

[0019] Among them, according to the weight quantization ratio of the current training stage, some elements in the Boolean mask matrix are set to preset values.

[0020] According to an embodiment of the present invention, the above-mentioned method of generating a Boolean mask matrix in the current training phase using a weight mask function includes:

[0021] Create or update the score matrix for the current training phase using the preset calibration data, the quantization parameter matrix from the previous training phase, and the original parameters of the target large language model;

[0022] The score matrix and weight quantization ratio of the current training stage are processed using a preset selection function to obtain a Boolean mask matrix of the current training stage.

[0023] According to an embodiment of the present invention, the step of creating or updating the score matrix of the current training phase using the preset calibration data, the quantization parameter matrix of the previous training phase, and the original parameters of the target large language model includes:

[0024] Calculate the local quantization loss for the current training phase using the preset calibration data, the quantization parameter matrix from the previous training phase, and the original parameters of the target large language model;

[0025] Calculate the second-order partial derivative of the original weight of the local quantization loss in the current training stage to obtain the Hessian matrix of the current training stage;

[0026] The Hessian matrix of the current training stage, the quantization parameter matrix of the previous training stage, and the original parameters of the target large language model are processed to obtain the score matrix of the current training stage.

[0027] According to an embodiment of the present invention, the above-mentioned updating of parameters of the target large language model in the current training phase using the quantization parameter matrix, the loss value of the current training phase, and the preset learning rate includes:

[0028] Calculate the loss value of the current training phase using the quantization parameter matrix, the original parameters of the target large language model, and the text data samples used in the current training phase;

[0029] Derivative the loss value of the current training stage to obtain the gradient value of the current training stage;

[0030] The parameters of the target large language model in the current training phase are updated through gradient descent operation using the preset learning rate and gradient value.

[0031] According to an embodiment of the present invention, the target large language model includes a multilingual machine translation model, and the text training dataset includes a corpus with an aligned relationship between a source language and a target language;

[0032] Among them, the clients include mobile communication devices and embodied intelligent robots.

[0033] According to an embodiment of the present invention, the target large language model includes an intelligent question-answering model, and the text training dataset includes a structured knowledge base with question-text answer labels and a multimodal dataset with image-question-text answer triple labels;

[0034] Among them, clients include smart home terminals, smart medical terminals, customer service terminals and smart education terminals.

[0035] The iterative quantization-aware training method for end-side deployment of large language models provided by the present invention dynamically allocates weight quantization ratios according to the training stage through a proportional scheduler, retains more original model performance during quantization, avoids the serious damage to the basic performance of the model caused by traditional static quantization, and reduces the performance loss of end-side devices due to low-precision calculations; selectively retains the original parameters through a Boolean mask matrix, reduces the quantization error of key weights in the self-attention layer, and maintains the accuracy of end-side decoding of the generative large language model; the sparse parameter structure generated by the mask matrix can trigger sparse computing optimization of the hardware accelerator, thereby improving inference throughput; in addition, the quantization parameter update strategy provided by the present invention allows the quantization granularity to be adjusted for different hardware, avoiding cross-platform compatibility issues caused by traditional one-time quantization. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0037] Figure 1 This is a diagram illustrating an application scenario of an iterative quantization-aware training method for device-side deployment of a large language model according to an embodiment of the present invention;

[0038] Figure 2 is a flowchart of an iterative quantization-aware training method for device-side deployment of a large language model according to an embodiment of the present invention;

[0039] Figure 3 1 is a structural block diagram of an iterative quantized perceptual training device for large language model terminal deployment according to an embodiment of the present invention;

[0040] Figure 4 1 is a block diagram of an electronic device suitable for implementing an iterative quantization-aware training method for end-side deployment of a large language model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0041] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.

[0042] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0043] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0044] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0045] With the development of large-scale generative language models, the number of model parameters has exploded. For example, the DeepSeek-R1 model has 671 billion parameters. A full deployment of this model requires at least 1248GB of memory. Deploying this model on a server equipped with eight NVIDIA A100 80GB graphics cards would require two such servers. These model parameters are stored as FP16 floating-point numbers, with each parameter occupying 16 bits (2 bytes). Quantization reduces the model's resource usage by reducing the parameter bit width to a lower precision (8-bit, 4-bit, etc.). However, this reduction in precision inevitably compromises model performance. Research on parameter quantization aims to alleviate these issues.

[0046] Currently popular quantization methods primarily focus on post-training quantization. This technique has the advantages of requiring minimal training and achieving rapid quantization, but also suffers from a significant disadvantage: poor quantization performance. Existing research has demonstrated that quantization-aware training can achieve superior performance compared to post-training quantization under equivalent settings. However, in low-bit quantization scenarios, direct quantization-aware training is suboptimal, as it can cause the model to lose most of its underlying capabilities and fail to fully utilize them. Furthermore, existing solutions still require significant hardware resources during parameter quantization.

[0047] In order to solve at least one of the problems in the prior art, the present invention provides a low-hard-cost and efficient parameter quantization method for the terminal-side deployment of a large generative language model.

[0048] Figure 1 This is an application scenario diagram of an iterative quantization-aware training method for terminal-side deployment of a large language model according to an embodiment of the present invention.

[0049] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a large language model and end-device inference acceleration. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0050] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0051] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0052] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.

[0053] It should be noted that the iterative quantization-aware training method for large language model end-side deployment provided by the embodiment of the present invention can generally be executed by the server 105. Accordingly, the iterative quantization-aware training device for large language model end-side deployment provided by the embodiment of the present invention can generally be set in the server 105. The iterative quantization-aware training method for large language model end-side deployment provided by the embodiment of the present invention can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the iterative quantization-aware training device for large language model end-side deployment provided by the embodiment of the present invention can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0054] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0055] The following will be based on Figure 1 The scene described by Figure 2 The iterative quantization-aware training method for end-side deployment of a large language model in the disclosed embodiment is described in detail.

[0056] Figure 2 This is a flowchart of an iterative quantization-aware training method for terminal-side deployment of a large language model according to an embodiment of the present invention.

[0057] like Figure 2 As shown, the iterative quantization-aware training method for terminal-side deployment of a large language model in this embodiment includes operations S210 to S260.

[0058] In operation S210, text data samples used in the current training phase are randomly selected from the text training data set, and all original parameters in the target large language model deployed on the server are initially quantized in the current training phase to obtain an initial quantization parameter matrix.

[0059] According to an embodiment of the present invention, the above-mentioned target large language model includes a multilingual machine translation model, and the above-mentioned text training data set includes a corpus with an alignment relationship between the source language and the target language; wherein the client includes a mobile communication device and an embodied intelligent robot.

[0060] According to an embodiment of the present invention, the above-mentioned target large language model includes an intelligent question-answering model, and the above-mentioned text training data set includes a structured knowledge base with question-text answer labels and a multimodal data set with image-question-text answer triple labels; wherein, the client includes smart home terminals, smart medical terminals, customer service terminals and smart education terminals.

[0061] In operation S220 , a weight quantization ratio of the current training phase is determined using a ratio scheduler, and a Boolean mask matrix of the current training phase is generated using a weight mask function.

[0062] In operation S230 , some parameters in the initial quantization parameter matrix are replaced with original parameters of the target large language model using the weighted quantization ratio and the Boolean mask matrix to obtain a quantization parameter matrix.

[0063] In operation S240, the target large language model with the quantization parameter matrix encodes the text data samples of the current training stage based on the self-attention mechanism and decodes them based on the cross-attention mechanism to obtain the text data processing result, and the text data processing result is used to obtain the loss value of the current training stage.

[0064] In operation S250 , parameters of the target large language model in the current training phase are updated using the quantization parameter matrix, the loss value in the current training phase, and a preset learning rate.

[0065] In operation S260, the operations of the current training stage are repeated for each training stage until the preset training conditions are met, and a large language model with quantized parameters is obtained. The target large language model with quantized parameters is deployed to the client, wherein the large language model with quantized parameters is used to process the text data of the client.

[0066] The iterative quantization-aware training method for end-side deployment of large language models provided by the present invention dynamically allocates weight quantization ratios according to the training stage through a proportional scheduler, avoiding excessive compression of key parameters by traditional static quantization and reducing performance losses caused by low-precision calculations on end-side devices; selectively retaining the original parameters through a Boolean mask matrix reduces the quantization error of key weights in the self-attention layer, and maintains the accuracy of end-side decoding of the generative large language model; the sparse parameter structure generated by the mask matrix can trigger sparse computing optimization of the hardware accelerator and improve inference throughput; in addition, the quantization parameter update strategy provided by the present invention allows the quantization granularity to be adjusted for different hardware, avoiding cross-platform compatibility issues caused by traditional one-time quantization.

[0067] According to an embodiment of the present invention, the above-mentioned iterative quantization-aware training method for terminal-side deployment of a large language model also includes: limiting the value range of the original parameters of the pre-trained large language model through a parameter truncation operation to obtain a target large language model; performing an initial scoring on the original parameters in the pre-trained large language model through a parameter scoring operation to obtain an initial score matrix, and initializing the weight quantization ratio used in the first round of training.

[0068] The above-mentioned pre-trained large language models, for example, the LLaMA model, deepseek model, GPT model, Claude model, ChatGLM-6B model, etc.

[0069] The above embodiment limits the parameters to a hardware-friendly range through truncation operations, avoids quantization overflow caused by extreme values, and improves the computing stability of the mobile chip.

[0070] According to an embodiment of the present invention, the original parameters of the above-mentioned target large language model have the same floating-point precision as the original parameters of the pre-trained large language model; wherein the initial score matrix has the same shape as the parameter matrix of the pre-trained large language model and is used to screen the quantized model parameters during the parameter quantization process.

[0071] According to an embodiment of the present invention, the above-mentioned use of a proportional scheduler to determine the weight quantization ratio of the current training stage includes: based on the total number of stages used in training, using a proportional scheduler to determine the weight quantization ratio of the current training stage in a linearly increasing manner.

[0072] The proportional scheduler is described in detail below through specific implementation methods.

[0073] The proportional scheduler is responsible for determining the proportion of weights to be quantized at each step ,The present invention fully considers the hardware cost problem in the parameter quantization process, and therefore directly chooses a simple scheduling method: linear increase. That is, It increases linearly from 0 to 1 within half of the total number of training steps, as shown in formula (1):

[0074] (1)

[0075] Where t is the number of training steps, , T is the total number of training steps.

[0076] According to an embodiment of the present invention, the Boolean mask matrix and the initial quantization parameter matrix have the same shape; wherein, according to the weight quantization ratio of the current training stage, some elements in the Boolean mask matrix are set to preset values.

[0077] According to an embodiment of the present invention, the above-mentioned generation of a Boolean mask matrix for the current training stage using a weight mask function includes: creating or updating a score matrix for the current training stage using preset calibration data, a quantization parameter matrix for a previous training stage, and original parameters of a target large language model; and processing the score matrix and weight quantization ratio for the current training stage using a preset selection function to obtain a Boolean mask matrix for the current training stage.

[0078] According to an embodiment of the present invention, the above-mentioned creation or updating of the score matrix of the current training stage using the preset calibration data, the quantization parameter matrix of the previous training stage, and the original parameters of the target large language model includes: calculating the local quantization loss of the current training stage using the preset calibration data, the quantization parameter matrix of the previous training stage, and the original parameters of the target large language model; calculating the second-order partial derivative of the original weight of the local quantization loss of the current training stage to obtain the Hessian matrix of the current training stage; processing the Hessian matrix of the current training stage, the quantization parameter matrix of the previous training stage, and the original parameters of the target large language model to obtain the score matrix of the current training stage.

[0079] The above embodiment uses the second-order derivative analysis of local quantization loss to accurately identify the parameters that have the greatest impact on the model, avoiding the blindness of traditional uniform quantization. It combines historical quantization parameters with the current Hessian matrix to achieve dynamic evaluation of parameter importance and improve the stability of end-side deployment.

[0080] The weight mask function is further described in detail below through a specific implementation method.

[0081] The overall calculation method of the weight mask function is shown in formula (2):

[0082] (2)

[0083] in is the unquantized FP16 weight, Generate a Score matrices of the same shape, The function generates a Boolean matrix with shape and Same, among them The element of is 1, The element of is 0.

[0084] The calculation method is shown in formulas (3) and (4):

[0085] (3)

[0086] (4)

[0087] in, is the unquantized FP16 weight, is the quantized weight, is the input from the calibration data, Is L relative to The Hessian matrix of is the loss when calculating the Hessian matrix, It is the weight of the unquantized FP16 Rank Elements of the column, is the first Rank Elements of the column, is the first weight after quantization Rank Elements of the column, is the first inverse Hessian matrix Rank Elements of a column.

[0088] The Sel function selects the top function here, which is the descending sort The front The elements of are set to 1 and the rest are set to 0. The weight mask function described above is denoted as max_to_min.

[0089] Each module here is pluggable. The present invention also includes two weight mask functions: min_to_max: Sel function is replaced by taking the minimum value, that is, after ascending sorting The front The elements of are set to 1, and the rest are 0; random: The elements in are sampled from a uniform distribution between 0 and 1, and the Sel function is the top function.

[0090] According to an embodiment of the present invention, the above-mentioned use of the quantization parameter matrix, the loss value of the current training stage and the preset learning rate to update the parameters of the target large language model in the current training stage includes: using the quantization parameter matrix, the original parameters of the target large language model and the text data samples used in the current training stage to calculate the loss value of the current training stage; derivatizing the loss value of the current training stage to obtain the gradient value of the current training stage; and using the preset learning rate and gradient value to update the parameters of the target large language model in the current training stage through a gradient descent operation.

[0091] The gradient calculation in the above embodiment considers both quantization parameters and original parameters, avoiding the gradient mismatch problem of traditional PTQ (post-training quantization) and improving the stability of on-device model fine-tuning. The learning rate is used to dynamically control the parameter update amplitude to prevent numerical overflow caused by gradient explosion on low-precision devices.

[0092] The iterative parameter quantization perceptual training method provided by the present invention is further described in detail below through specific implementation methods.

[0093] The iterative parameter quantization-aware training framework provided by this invention is a quantization-aware training framework based on gradient descent. Its core innovations include a scale scheduler and a weight mask function. First, given a model saved in FP16 format and a batch of training data, a training loop is started. In each training step, a scale scheduler is first used to determine the weight quantization scale in the range of 0 to 1. ,only The weights of will be quantized at this step. Then use the weight mask function to generate a Boolean mask matrix with the same shape as the weight, which contains The elements of the mask matrix are 1, and the remaining elements are 0. The weights of 1 in the mask matrix will be quantized, and the weights of 0 in the mask matrix will not be quantized. Then forward propagation is performed and the loss is calculated according to the loss function. That is to say, in the forward propagation process, only The quantization error is introduced into the weights of FP16. Then the FP16 weights are back-propagated to update the original FP16 weights.

[0094] The advantages of the above method provided by the present invention are further verified by specific experiments below.

[0095] This paper mainly experiments on the LLaMA (Large Language Model Meta AI) series of models to verify the superiority of this method, and compares two post-training methods, GPTQ and AWQ, and two quantization-aware training methods, LLM-QAT and BitDistiller. Table 1 shows the results of IQAT on the LLaMA-3.1-8B model. The training dataset is a subset of the FineWeb dataset. "Ours" indicates the method provided by the present invention. Table 1 shows that, regardless of the loss function corresponding to the LLM-QAT quantization-aware training method or the BitDistiller quantization-aware training method, the method provided by the present invention achieves good technical results in terms of PPL (perplexity) and AVG metrics compared to the two post-training methods, GPTQ and AWQ. PPL is evaluated on the WikiText-2 dataset, and AVG is the average of the model's results on HellaSwag, Winogrande, PIQA, ARC-c, GSM8k, and GSM8k-COT. "+random" in Table 1 indicates random weight selection based on score, "+max_to_min" in Table 1 indicates weight selection from highest to lowest score, and "+min_to_max" in Table 1 indicates weight selection from lowest to highest score. "avg" in Table 1 indicates the average of "+random", "+max_to_min", and "+min_to_max". Table 2 shows the results on the ‌LLaMA3.2-3B model, where Ours indicates that the method provided by the present invention is adopted. The avg, +random, +max_to_min, +min_to_max and AVG in Table 2 have the same meanings as avg, +random, +max_to_min, +min_to_max and AVG in Table 1. It can be seen from Table 2 that whether the loss function corresponding to the LLM-QAT quantization-aware training method or the loss function corresponding to the BitDistiller quantization-aware training method is adopted, the method provided by the present invention has achieved good technical effects on the PPL (perplexity) and AVG indicators compared with the two post-training methods of GPTQ and AWQ. Table 3 shows the results on the conversation model. MMLU is a dataset of test questions in various subjects, MT-Bench is a dataset used to evaluate the ability of large models to respond to conversations in real scenarios, and Ours indicates that the method provided by the present invention is used. 8B represents the ‌LLaMA-3.1-8B-instruct model, and 3B represents the ‌LLaMA-3.2-3B model. The average improvement percentage indicates the average improvement of the method provided by the present invention (Ours) on the three evaluation indicators of PPL, MMLU, and MT-Bench compared with the method corresponding to the previous row.

[0096] Table 1: IQAT results on the LLaMA-3.1-8B model

[0097]

[0098] Table 2: IQAT results on the LLaMA-3.2-3B model

[0099]

[0100] Table 3: Results on the ‌LLaMA-3.1-8B-instruct model and the ‌LLaMA-3.2-3B-instruct model

[0101]

[0102] Based on the above-mentioned iterative quantization-aware training method for large language model terminal deployment, the present invention also provides an iterative quantization-aware training device for large language model terminal deployment. Figure 3 The device is described in detail.

[0103] Figure 3 1 is a structural block diagram of an iterative quantized perceptual training device for terminal-side deployment of a large language model according to an embodiment of the present invention.

[0104] like Figure 3 As shown, the iterative quantization-aware training device 300 for end-side deployment of a large language model in this embodiment includes a parameter initial quantization module 310, a quantization scale and mask matrix acquisition module 320, a quantization parameter matrix acquisition module 330, a model training module 340, a parameter update module 350, and an iterative training and deployment module 360.

[0105] Initial parameter quantization module 310 is configured to randomly select text data samples used in the current training phase from the text training dataset and perform initial quantization on all original parameters of the target large language model deployed on the server for the current training phase, thereby obtaining an initial quantization parameter matrix. In one embodiment, initial parameter quantization module 310 can be used to perform operation S210 described above and will not be further described here.

[0106] Quantization ratio and mask matrix acquisition module 320 is configured to determine the weight quantization ratio for the current training phase using the ratio scheduler and generate a Boolean mask matrix for the current training phase using the weight mask function. In one embodiment, quantization ratio and mask matrix acquisition module 320 can be used to perform operation S220 described above and will not be further described here.

[0107] The quantization parameter matrix acquisition module 330 is configured to replace some parameters in the initial quantization parameter matrix with the original parameters of the target large language model using the weighted quantization ratio and the Boolean mask matrix, thereby obtaining the quantization parameter matrix. In one embodiment, the quantization parameter matrix acquisition module 330 can be configured to perform operation S230 described above, and will not be further described here.

[0108] Model training module 340 is configured to encode the text data samples in the current training phase using a self-attention mechanism and decode them using a cross-attention mechanism using a target large language model with a quantized parameter matrix, thereby obtaining a text data processing result, and using the text data processing result to obtain a loss value for the current training phase. In one embodiment, model training module 340 can be used to perform operation S240 described above, which will not be further described here.

[0109] The parameter update module 350 is configured to update the parameters of the target large language model for the current training phase using the quantization parameter matrix, the loss value for the current training phase, and a preset learning rate. In one embodiment, the parameter update module 350 may be configured to perform operation S250 described above and will not be further described herein.

[0110] Iterative training and deployment module 360 ​​is configured to repeat the operations of the current training phase for each training phase until preset training conditions are met, thereby obtaining a large language model with quantized parameters and deploying the target large language model with quantized parameters to the client. The large language model with quantized parameters is used to process text data from the client. In one embodiment, iterative training and deployment module 360 ​​can be configured to perform operation S260 described above and will not be further described here.

[0111] According to an embodiment of the present invention, any multiple modules among the initial parameter quantization module 310, the quantization scale and mask matrix acquisition module 320, the quantization parameter matrix acquisition module 330, the model training module 340, the parameter update module 350, and the iterative training and deployment module 360 ​​may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to an embodiment of the present invention, at least one of the initial parameter quantization module 310, the quantization scale and mask matrix acquisition module 320, the quantization parameter matrix acquisition module 330, the model training module 340, the parameter update module 350, and the iterative training and deployment module 360 ​​may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or implemented in any one of the three implementation methods, or any appropriate combination of any of them. Alternatively, at least one of the parameter initial quantization module 310, the quantization scale and mask matrix acquisition module 320, the quantization parameter matrix acquisition module 330, the model training module 340, the parameter update module 350, and the iterative training and deployment module 360 ​​can be at least partially implemented as a computer program module, which can perform corresponding functions when executed.

[0112] Figure 4 1 is a block diagram of an electronic device suitable for implementing an iterative quantization-aware training method for end-side deployment of a large language model according to an embodiment of the present invention.

[0113] like Figure 4 As shown, an electronic device 400 according to an embodiment of the present invention includes a processor 401, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 402 or a program loaded from a storage unit 408 into a random access memory (RAM) 403. Processor 401 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. Processor 401 may also include onboard memory for caching purposes. Processor 401 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0114] Various programs and data required for the operation of the electronic device 400 are stored in the RAM 403. The processor 401, ROM 402, and RAM 403 are connected to each other via a bus 404. The processor 401 executes the programs in the ROM 402 and / or RAM 403 to perform various operations according to the method flow of the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than the ROM 402 and RAM 403. The processor 401 may also execute the programs stored in the one or more memories to perform various operations according to the method flow of the embodiment of the present invention.

[0115] According to an embodiment of the present invention, electronic device 400 may further include an input / output (I / O) interface 405, which is also connected to bus 404. Electronic device 400 may also include one or more of the following components connected to I / O interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 408 including a hard disk; and a communication section 409 including a network interface card such as a LAN card or modem. Communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to I / O interface 405 as needed. Removable media 411, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 410 as needed, so that computer programs read from the removable media can be installed into storage section 408 as needed.

[0116] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.

[0117] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, a computer-readable storage medium may include ROM 402 and / or RAM 403 described above, and / or one or more memories other than ROM 402 and RAM 403.

[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0119] It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or couplings fall within the scope of the present invention.

[0120] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

Claims

1. An iterative quantization-aware training method for large language model on-device deployment, characterized in that: The method comprises: Randomly selecting text data samples used in the current training phase from the text training dataset, performing initial quantization of all original parameters in the target large language model deployed on the server for the current training phase, and obtaining an initial quantization parameter matrix; Determining a weight quantization ratio for the current training phase using a proportional scheduler, and generating a Boolean mask matrix for the current training phase using a weight mask function; Replacing some parameters in the initial quantization parameter matrix with original parameters of the target large language model using the weight quantization ratio and the Boolean mask matrix to obtain a quantization parameter matrix; Using the target large language model having the quantization parameter matrix to encode the text data sample of the current training stage based on the self-attention mechanism and to decode it based on the cross-attention mechanism, a text data processing result is obtained, and a loss value of the current training stage is obtained using the text data processing result; Updating parameters of the target large language model in the current training phase using the quantization parameter matrix, the loss value in the current training phase, and a preset learning rate; Repeat the operations of the current training stage for each training stage until preset training conditions are met, obtain a large language model with quantized parameters, and deploy the target large language model with quantized parameters to the client, wherein the large language model with quantized parameters is used to process text data of the client.

2. The method according to claim 1, characterized in that Also includes: Limiting the value range of original parameters of the pre-trained large language model through a parameter truncation operation to obtain the target large language model; The original parameters in the pre-trained large language model are initially scored through a parameter scoring operation to obtain an initial score matrix, and the weight quantization ratio used in the first round of training is initialized.

3. The method according to claim 2, characterized in that The original parameters of the target large language model and the original parameters of the pre-trained large language model have the same floating point precision; The initial score matrix has the same shape as the parameter matrix of the pre-trained large language model and is used to screen the quantized model parameters during the parameter quantization process.

4. The method according to claim 1, wherein Determining the weight quantization ratio of the current training phase using the proportional scheduler includes: Based on the total number of stages used in training, the weight quantization ratio of the current training stage is determined in a linearly increasing manner using the ratio scheduler.

5. The method according to claim 1, characterized in that The Boolean mask matrix has the same shape as the initial quantization parameter matrix; Part of the elements in the Boolean mask matrix are set to preset values ​​according to the weight quantization ratio of the current training stage.

6. The method according to claim 5, characterized in that Generating the Boolean mask matrix of the current training phase using the weight mask function includes: Creating or updating the score matrix of the current training phase using preset calibration data, the quantization parameter matrix of the previous training phase, and the original parameters of the target large language model; The score matrix and the weight quantization ratio of the current training stage are processed using a preset selection function to obtain a Boolean mask matrix of the current training stage.

7. The method according to claim 6, characterized in that Creating or updating the score matrix of the current training phase using preset calibration data, the quantization parameter matrix of the previous training phase, and the original parameters of the target large language model includes: Calculating the local quantization loss of the current training stage using the preset calibration data, the quantization parameter matrix of the previous training stage, and the original parameters of the target large language model; Calculating the second-order partial derivative of the original weight of the local quantization loss in the current training stage to obtain the Hessian matrix of the current training stage; The Hessian matrix of the current training stage, the quantization parameter matrix of the previous training stage, and the original parameters of the target large language model are processed to obtain a score matrix of the current training stage.

8. The method according to claim 1, characterized in that Updating the parameters of the target large language model in the current training phase using the quantization parameter matrix, the loss value in the current training phase, and a preset learning rate includes: Calculating a loss value of the current training phase using the quantization parameter matrix, original parameters of the target large language model, and the text data samples used in the current training phase; Derivative the loss value of the current training stage to obtain the gradient value of the current training stage; Utilizing the preset learning rate and the gradient value, a gradient descent operation is performed on the target large language model to update parameters of the current training phase.

9. The method according to any one of claims 1 to 8, characterized in that The target large language model includes a multilingual machine translation model, and the text training dataset includes a corpus with an aligned relationship between a source language and a target language; The client includes a mobile communication device and an embodied intelligent robot.

10. The method according to any one of claims 1 to 8, characterized in that The target large language model includes an intelligent question-answering model, and the text training dataset includes a structured knowledge base with question-text answer labels and a multimodal dataset with image-question-text answer triple labels; The client includes a smart home terminal, a smart medical terminal, a customer service terminal and a smart education terminal.

Citation Information

Patent Citations

  • Neural-network-model compression method, system and device and readable storage medium

    CN108229681A

  • Quantitative perception training method and related device

    CN115293324A

  • Neural network model training method and device, electronic equipment and storage medium

    CN118194954A

  • Model training method and device, computer equipment, storage medium and program product

    CN119047523A

  • Method and device for compressing generative pre-trained language models via quantization

    US20240104346A1