Distributed quantitative perception training method and system for large model
By employing a distributed quantization-aware training method, combined with an adaptive weighting mechanism of gradients and Hessian matrices, and selecting sensitive samples for parallel processing, the problem of high computational cost in large model training is solved, achieving efficient quantization training and parallel acceleration.
Patent Information
- Application Number
- CN202510844861.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-04-08
- Filing Date
- 2025-06-23
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies have high computational and time costs when training large models, and the quantization quality and training efficiency of quantization below 8-bit are poor, making it difficult to balance the preservation of quantization information and training efficiency.
A distributed quantization-aware training method is adopted, which combines the gradient criterion and Hessian matrix. Through an adaptive weighting mechanism, the sample data most sensitive to quantization-awareness is selected, and the core sample selection and quantization-awareness training are processed in parallel. Distributed parallel training technology is used.
It improves the efficiency and effectiveness of quantization training for large models, reduces computation time, and maintains the accuracy and parallel acceleration capabilities of the models.
Smart Images

Figure CN120930727A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of large model training technology, specifically relating to a distributed quantization-aware training method and system for large models. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] With the continuous increase in the scale of deep neural networks and the size of training data, especially with the rapid development in computer vision and large language models, it has become a consensus that increasing model size improves the model's generalization ability. Therefore, the industry must confront the problem that the computational and time costs of training these large models will become a major obstacle to model development.
[0004] Quantization-aware quantization can reduce the bit width of model parameters, lowering the computational and storage costs of deep neural networks while maintaining inference accuracy, significantly alleviating memory constraints encountered when deploying large models. However, in work involving quantization below 8 bits, the quantization quality and training efficiency are less than satisfactory compared to other quantization methods. Currently, the fusion of training and inference for large models faces the following challenges: how to retain more effective quantization information, and how to balance training efficiency and quantization performance. Summary of the Invention
[0005] To address the aforementioned issues, this invention proposes a distributed quantization-aware training method and system for large models. Combining gradient criteria and the Hessian matrix, it comprehensively considers both first-order gradient information perception and Hessian gradient information perception, and achieves quantization-aware training of large models through an adaptive weighting mechanism. Since calculating gradients increases model training time, this invention employs a parallel training approach to reduce the time spent calculating gradient information in order to improve quantization training efficiency and quantization effect.
[0006] According to some embodiments, the first solution of the present invention provides a distributed quantization-aware training method for large models, employing the following technical solution: A distributed quantization-aware training method for large models includes: Obtain the configuration parameter information of the large model to be trained; Based on the obtained configuration parameter information, the loss function of the large model is constructed, the gradient norm is calculated, and the first-order gradient information perception score of the large model is obtained. Add perturbations to the large model, construct the perturbation-included loss function of the large model, calculate the Hessian matrix, and obtain the Hessian gradient information perception score of the large model. An adaptive weighted combination of the first-order gradient information perception score and the Hessian gradient information perception score of the obtained large model is performed to obtain the quantitative perception score of the large model. Based on the obtained large model quantization perception score, determine the large model quantization perception training samples, and complete the distributed quantization perception training of the large model.
[0007] As a further technical limitation, in the process of determining the training samples for large-scale model quantization perception, the weights in the obtained large-scale model quantization perception scores are sorted by size, the training samples are determined according to the sorting, and the determined samples are processed in parallel to complete the distributed quantization perception training of the large-scale model.
[0008] As a further technical limitation, in the process of adaptive weighted combination, after standardizing the obtained first-order gradient information perception score and Hessian gradient information perception score, the weights of the first-order gradient information perception score and the Hessian gradient information perception score are calculated, and the sample scores are adaptively weighted and combined to obtain the quantitative perception score of the large model. Based on the quantitative perception score of the large model, the sample data suitable for quantitative perception training of the large model are determined.
[0009] As a further technical limitation, in the process of distributed quantization-aware training of a large model, the determined large model quantization-aware training samples are divided into several subsets, the subsets are processed and the gradients of the subsets are calculated, and the gradients of all subsets are aggregated in combination with the communication protocol to complete the update of the global model parameters; the distributed quantization-aware training of the large model is completed based on the updated model parameters.
[0010] As a further technical limitation, the obtained large model loss function is backpropagated, and the partial derivatives of the backpropagated loss function with respect to the model parameters are calculated to obtain the model gradient. The sum of the squares of all obtained model gradients is taken as the gradient norm, and the square root of the gradient norm is removed to obtain the first-order gradient information perception score of the large model.
[0011] As a further technical limitation, the second-order partial derivatives of the perturbation loss function of the constructed large model with respect to the model parameters are calculated. The obtained second-order partial derivatives are used as the diagonal elements of the Hessian matrix, and the sum of squares of all diagonal elements is calculated to obtain the Hessian gradient information perception score of the large model.
[0012] According to some embodiments, the second aspect of the present invention provides a distributed quantization-aware training system for large models, employing the following technical solution: A distributed quantization-aware training system for large models, comprising: The acquisition module is configured to acquire configuration parameter information of the large model to be trained; The first scoring module is configured to construct the loss function of the large model based on the acquired configuration parameter information, calculate the gradient norm, and obtain the first-order gradient information-aware score of the large model. The second scoring module is configured to add perturbations to the large model, construct the perturbation-containing loss function of the large model, calculate the Hessian matrix, and obtain the Hessian gradient information perception score of the large model. The weighting module is configured to adaptively weight and combine the first-order gradient information perception score and the Hessian gradient information perception score of the obtained large model to obtain the quantitative perception score of the large model. The training module is configured to determine the large model quantization perception training samples based on the obtained large model quantization perception score, and complete the distributed quantization perception training of the large model.
[0013] According to some embodiments, a third aspect of the present invention provides a computer-readable storage medium, employing the following technical solution: A computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the steps in the distributed quantization-sensing training method for a large model as described in the first aspect of the present invention.
[0014] According to some embodiments, the fourth aspect of the present invention provides an electronic device, which adopts the following technical solution: An electronic device includes a memory, a processor, and a program stored in the memory and running on the processor, wherein the processor executes the program to implement the steps in the distributed quantization-sensing training method for a large model as described in the first aspect of the present invention.
[0015] According to some embodiments, the fifth aspect of the present invention provides a computer program product, which adopts the following technical solution: A computer program product includes software code, wherein the program in the software code performs the steps in the distributed quantization-sensing training method for large models as described in the first aspect of the present invention.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention combines the gradient criterion and the Hessian matrix, considering both first-order gradient information perception and Hessian gradient information perception. Through an adaptive weighting mechanism, it rationally sets the weights of the scoring function, selects the sample data most sensitive to quantization perception training, and implements distributed training in the core sample selection and quantization perception process. This achieves both improved efficiency and quantization effect of large-scale model quantization training and parallel acceleration of large-scale models. Attached Figure Description
[0017] The accompanying drawings, which form part of this embodiment, are used to provide a further understanding of this embodiment. The illustrative embodiments and their descriptions are used to explain this embodiment and do not constitute an improper limitation of this embodiment.
[0018] Figure 1 This is a flowchart of the distributed quantization-aware training method for a large model in Embodiment 1 of the present invention. Figure 2(a) is a schematic diagram of the training time of the MobileNetV2 model in Embodiment 1 of the present invention; Figure 2(b) is a schematic diagram of the training accuracy of the MobileNetV2 model in Embodiment 1 of the present invention; Figure 3(a) is a schematic diagram of the parallel time of adding distributed data in Embodiment 1 of the present invention; Figure 3(b) is a schematic diagram of the model accuracy in Embodiment 1 of the present invention with the addition of distributed data. Detailed Implementation
[0019] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0020] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0021] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0022] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0023] Example 1 Embodiment 1 of this invention introduces a distributed quantization-sensing training method for large models.
[0024] like Figure 1As shown, this embodiment combines the gradient criterion and Hessian diagonal to select core samples that better represent the impact of data on the lost landscape. By considering both the first-order and second-order information of the gradient during the training process, it provides a more comprehensive selection criterion than existing methods. The adaptive weighting mechanism can dynamically adjust the importance of gradient and Hessian information according to their variance, ensuring a balanced and effective selection process based on the specific characteristics of the training data and the model. In the core sample selection process, since gradient calculation is too time-consuming, distributed parallel training is used to accelerate the training speed. Distributed parallel training is also used in the quantization perception part, which comprehensively speeds up the model training.
[0025] Gradient Hessian Quantization Aware Training (QAT-Grad Hessian) technique This embodiment uses Hessian matrix scoring and first-order gradient information scoring for the target quantization model; QAT-Grad Hessian consists of three core components, including: (1) First-order gradient information perception, i.e., identifying samples that are sensitive to the QAT model; (2) Hessian gradient information perception, that is, when the first-order information cannot fully represent the sensitivity of the sample to the quantization model, the second step requires Hessian matrix information to determine the sensitivity of the sample to the full-precision model, so as to realize the model to perceive the importance of the input sample in a more detailed way. (3) Weighted adaptive mechanism, which is used to weight and balance the first-order gradient information and Hessian gradient information to adaptively perceive the input sample.
[0026] This embodiment uses the first-order gradient to reflect the loss function, that is, the rate of change of the model parameters. When calculating the gradient of an input sample, we are actually evaluating how sensitive the model parameters are to the loss change caused by that sample; if the gradient norm of the sample is large, it means that the sample has a more significant impact on the model parameters, indicating that the model is more sensitive to the sample.
[0027] This embodiment describes the loss function. Perform backpropagation and calculate the gradient. ;in, The parameters representing the model, i Index representing the model parameters.
[0028] Calculate the sum of squared gradients as the gradient norm: (1) Taking the square root of formula (1), we get: (2) During the quantization-aware training of QAT, the model needs to adapt to quantization errors. By calculating and analyzing the gradient norm, the top K ranked samples are selected to identify the samples that have the greatest impact on the results during quantization-aware training. The obtained samples are used as a subset to participate in the model training process so that the model can still maintain high accuracy after quantization.
[0029] This embodiment perturbs the parameter weights and observes the increase in training loss.
[0030] Assuming a given parameter perturbation, a second-order Taylor approximation of the training loss is used. Apply small perturbations to it , making The change in training loss can be expressed as: (3) in, , representing the loss function relative parameters The expectation of the gradient, and This represents the i-th value of the Hessian matrix representing the loss. It is a loss function relative parameters The second derivative (i.e., the diagonal elements of the Hessian matrix).
[0031] Since the target model is a convergent model, it is assumed that... That is, formula (3) can be simplified to: (4) therefore, The impact of perturbations caused by training samples on the model training loss was quantified. The larger the value, the higher the sensitivity of the sample to the full-precision model, and the richer the information contained in the sample, so the greater the possibility of selecting it as a core sample.
[0032] In practical calculations, directly obtaining the Hessian matrix score is extremely complex and computationally intensive, which undoubtedly increases the additional computational cost of the quantization process. Therefore, this embodiment uses the Fisher information matrix (F) as an effective approximation; that is, it uses the sum of squared gradients to approximate the sum of squares of the diagonal elements of the Hessian matrix.
[0033] Since H is the Hessian matrix of the negative log-likelihood loss, H is equivalent to the Fisher information matrix, i.e.: (5) In this embodiment, the score based on the Hessian matrix is denoted as: (6) Weighted combinations of gradient scores and Hessian matrix scores can make core sample selection methods more reasonable. Different weights can reflect the different importance of gradient information and Hessian information in sample selection.
[0034] This embodiment introduces weighting coefficients to adjust the relative importance of these two scores, standardizing the gradient score and the Hessian matrix score respectively: (7) (8) in, It is the mean of the first-order gradient scores. It is the standard deviation of the first-order gradient information score. It is the mean of the Hessian scores. The standard deviation of the first-order gradient score is given; the variances of the first-order gradient score and the Hessian matrix score are respectively: (9) (10) The weights for obtaining the gradient and Hessian score are: (11) (12) Combining the first-order gradient weights and the Hessian matrix weights, we obtain the final total score: (13) This embodiment balances the influence of various types of samples and comprehensively selects the sample with the highest score as the core sample.
[0035] Parallel acceleration techniques for quantization-aware training processes In the field of deep learning, quantization-aware training is a key technique that aims to reduce the memory footprint and computational cost of models by quantizing model parameters while maintaining model performance. However, as model size and dataset size increase, single-machine training methods are no longer sufficient, and most research on quantization-aware training has not considered the computational efficiency required to train a quantized model. Therefore, a training method that considers both the training efficiency and the quantization effect is proposed.
[0036] To address the low training efficiency inherent in quantization-aware training, this embodiment proposes a core sample selection method based on the first-order gradient and Hessian matrix. The core samples selected based on scores are used as training data in the quantization-aware training phase. Parallel acceleration techniques such as data parallelism are employed during the core sample selection process and distributed data parallelism are used during the quantization-aware training process.
[0037] The gradient Hessian quantization-based perceptual training technique consists of three core components. The selection process requires gradient calculation, which significantly increases the training time. To ensure effective improvement in computational efficiency, this embodiment employs a data parallelism strategy. During core sample selection, input samples are distributed to different computing devices, and the forward and backward propagation processes for core sample selection are completed in parallel, reducing the additional training time consumed by gradient calculation.
[0038] In the quantization-aware training process, this embodiment divides the core dataset into multiple subsets and processes these subsets in parallel on two V100 computing devices. Each computing device stores a complete copy of the model and independently computes the gradients on its assigned data subset. Then, through an optimized communication protocol, such as AllReduce, the gradients from all devices are aggregated to update the global model parameters. This embodiment not only improves training speed but also reduces memory usage on a single device through parallel processing, enabling the training of larger models.
[0039] Distributed training technology is used in both stages, which makes the technical solution in this embodiment balance the disadvantage of quantization-aware training being time-consuming compared to other quantization methods, as well as the time required for gradient calculation during the inefficient core sample selection process.
[0040] Case Analysis In this embodiment, the data shown in Table 1 is used to compare the QAT Top-1 accuracy of MobileNetV2 on CIFAR-100. The bit width of the weights / activations is 2 / 32. The training time and accuracy of the MobileNetV2 model are shown in Figure 2(a) and Figure 2(b), respectively.
[0041] As can be seen, the gradient Hessian quantization-based perceptual training technique already performs very well. When the subset accounts for 40% of the entire dataset, the training accuracy of the method in this embodiment is improved by 1%, and the accuracy is the highest in most cases, which verifies the superiority of the quantization effect of the method in this embodiment.
[0042] Table 1. Comparison of Top-1 accuracy of different core sample selection methods on different subsets.
[0043] Table 2 shows the Top-1 accuracy comparison of different core sample selection methods under different subsets. The performance of adding data parallelism and distributed data parallelism techniques during core sample selection and quantization-aware training is illustrated in Figures 3(a) and 3(b), respectively. The experiment used two V100 GPUs. The time and model accuracy after adding distributed data parallelism are shown in Figures 3(a) and 3(b), respectively. It can be seen that the addition of data parallelism in this embodiment did not affect the model's quantization performance, and both methods outperformed the state-of-the-art method ACS. Table 3 shows the training time using the methods described in Table 2. Under different data subset selections, QAT_GradHessian_DDP improved performance by 34.73%, 39%, 42%, 45%, and 44% compared to QAT_GradHessian, respectively, thus verifying the efficiency of the quantization process in this embodiment.
[0044] Table 2. Comparison of Top-1 accuracy of different core sample selection methods on different subsets.
[0045] Table 3. Comparison of Top-1 times for different core sample selection methods on different subsets.
[0046] This embodiment combines the gradient criterion and the Hessian matrix, considering both first-order gradient information perception and Hessian gradient information perception. Through an adaptive weighting mechanism, the weights of the scoring function are reasonably set, and the sample data most sensitive to quantization perception training is selected. Combined with distributed training implemented in the core sample selection and quantization perception process, it achieves both improved efficiency and quantization effect of large model quantization training and parallel acceleration of large models.
[0047] Example 2 Embodiment 2 of this invention introduces a distributed quantization perception training system for large models.
[0048] A distributed quantization-aware training system for large models, comprising: The acquisition module is configured to acquire configuration parameter information of the large model to be trained; The first scoring module is configured to construct the loss function of the large model based on the acquired configuration parameter information, calculate the gradient norm, and obtain the first-order gradient information-aware score of the large model. The second scoring module is configured to add perturbations to the large model, construct the perturbation-containing loss function of the large model, calculate the Hessian matrix, and obtain the Hessian gradient information perception score of the large model. The weighting module is configured to adaptively weight and combine the first-order gradient information perception score and the Hessian gradient information perception score of the obtained large model to obtain the quantitative perception score of the large model. The training module is configured to determine the large model quantization perception training samples based on the obtained large model quantization perception score, and complete the distributed quantization perception training of the large model.
[0049] The detailed steps are the same as the distributed quantization perception training method for large models provided in Example 1, and will not be repeated here.
[0050] Example 3 Embodiment 3 of the present invention provides a computer-readable storage medium.
[0051] A computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the distributed quantization-sensing training method for large models as described in Embodiment 1 of the present invention.
[0052] The detailed steps are the same as the distributed quantization perception training method for large models provided in Example 1, and will not be repeated here.
[0053] Example 4 Embodiment 4 of the present invention provides an electronic device.
[0054] An electronic device includes a memory, a processor, and a program stored in the memory and running on the processor, wherein the processor executes the program to implement the steps in the distributed quantization perception training method for a large model as described in Embodiment 1 of the present invention.
[0055] The detailed steps are the same as the distributed quantization perception training method for large models provided in Example 1, and will not be repeated here.
[0056] Example 5 Embodiment 5 of the present invention provides a computer program product.
[0057] A computer program product includes software code, wherein the program in the software code performs the steps of the distributed quantization perception training method for large models as described in Embodiment 1 of the present invention.
[0058] The detailed steps are the same as the distributed quantization perception training method for large models provided in Example 1, and will not be repeated here.
[0059] The above description is merely a preferred embodiment of this practice and is not intended to limit the scope of this practice. Various modifications and variations can be made to this practice by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this practice should be included within the protection scope of this practice.
Claims
1. A distributed quantization-sensing training method for large models, characterized in that, include: Obtain the configuration parameter information of the large model to be trained; Based on the obtained configuration parameter information, the loss function of the large model is constructed, the gradient norm is calculated, and the first-order gradient information perception score of the large model is obtained. Add perturbations to the large model, construct the perturbation-included loss function of the large model, calculate the Hessian matrix, and obtain the Hessian gradient information perception score of the large model. An adaptive weighted combination of the first-order gradient information perception score and the Hessian gradient information perception score of the obtained large model is performed to obtain the quantitative perception score of the large model. Based on the obtained large model quantization perception score, determine the large model quantization perception training samples, and complete the distributed quantization perception training of the large model.
2. The distributed quantization-aware training method for a large model as described in claim 1, characterized in that, In determining the training samples for large-scale model quantization perception, the weights in the obtained large-scale model quantization perception scores are sorted by size, and the training samples are determined according to the sorting. The determined samples are then processed in parallel to complete the distributed quantization perception training of the large-scale model.
3. The distributed quantization-aware training method for a large model as described in claim 1, characterized in that, In the adaptive weighted combination process, after standardizing the obtained first-order gradient information perception score and Hessian gradient information perception score, the weights of the first-order gradient information perception score and the Hessian gradient information perception score are calculated. The sample scores are then adaptively weighted and combined to obtain the quantization perception score of the large model. Based on the quantization perception score of the large model, suitable sample data for quantization perception training of the large model are determined.
4. The distributed quantization-aware training method for a large model as described in claim 1, characterized in that, In the distributed quantization-aware training of a large model, the determined large model quantization-aware training samples are divided into several subsets, the subsets are processed and their gradients are calculated, and the gradients of all subsets are aggregated in combination with the communication protocol to complete the update of the global model parameters; the distributed quantization-aware training of the large model is completed based on the updated model parameters.
5. The distributed quantization-aware training method for a large model as described in claim 1, characterized in that, The obtained large model loss function is backpropagated, and the partial derivatives of the backpropagated loss function with respect to the model parameters are calculated to obtain the model gradient. The sum of the squares of all obtained model gradients is taken as the gradient norm, and the square root of the gradient norm is removed to obtain the first-order gradient information perception score of the large model.
6. The distributed quantization-aware training method for a large model as described in claim 1, characterized in that, Calculate the second-order partial derivatives of the perturbation loss function with respect to the model parameters of the constructed large model. Use the obtained second-order partial derivatives as the diagonal elements of the Hessian matrix. Calculate the sum of squares of all diagonal elements to obtain the Hessian gradient information perception score of the large model.
7. A distributed quantization-sensing training system for large models, characterized in that, include: The acquisition module is configured to acquire configuration parameter information of the large model to be trained; The first scoring module is configured to construct the loss function of the large model based on the acquired configuration parameter information, calculate the gradient norm, and obtain the first-order gradient information-aware score of the large model. The second scoring module is configured to add perturbations to the large model, construct the perturbation-containing loss function of the large model, calculate the Hessian matrix, and obtain the Hessian gradient information perception score of the large model. The weighting module is configured to adaptively weight and combine the first-order gradient information perception score and the Hessian gradient information perception score of the obtained large model to obtain the quantitative perception score of the large model. The training module is configured to determine the large model quantization perception training samples based on the obtained large model quantization perception score, and complete the distributed quantization perception training of the large model.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the distributed quantization-aware training method for large models as described in any one of claims 1-6.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements the steps of the distributed quantization-aware training method for large models as described in any one of claims 1-6.
10. A computer program product, comprising software code, characterized in that, The program in the software code performs the steps of the distributed quantization-aware training method for large models as described in any one of claims 1-6.