Word sequence prediction method for large-model mixing precision quantification driven by multiple weight saliency
By dynamically adjusting the number of model weight quantization bits through the simulated annealing algorithm and matrix sparsity distribution, the problem of inaccurate weight identification in word sequence prediction tasks of large language models is solved, and the deployment efficiency and accuracy of the model in resource-constrained environments are improved.
Patent Information
- Application Number
- CN202510735007.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Existing mixed-precision quantization methods have difficulty accurately identifying key weights in large language models, resulting in logical errors or semantic deviations in word sequence prediction tasks, and are unable to adapt to the imbalance between performance and resource utilization in different task requirements and resource-constrained environments.
The simulated annealing algorithm is used to dynamically allocate bit width. Combined with the matrix sparsity distribution and the optimal split point strategy, the weight significance is calculated through the Hessian matrix and Cholesky decomposition. The high significance, intermediate significance and low significance are stored in layers and quantized, and the number of quantization bits of the model weights is dynamically adjusted.
It achieves optimal resource allocation under complex weight distribution and different data characteristics, improves the inference speed and prediction accuracy of large language models, and especially improves the deployment efficiency of models in edge computing environments.
Smart Images

Figure CN120597871A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and AI question answering, and specifically to a large-model mixed-precision quantized word sequence prediction method driven by multiple weighted saliency. Background Art
[0002] Transformer-based large language models (LLMs) such as OPT and LLaMA have made revolutionary progress in natural language processing, demonstrating exceptional performance in word sequence prediction tasks such as text generation, machine translation, and dialogue systems. These models learn universal language representations through large-scale pre-training, enabling accurate prediction of word sequences in natural language. They are widely used in scenarios such as intelligent customer service, intelligent writing, and real-time translation. However, LLMs are large in scale, often with billions or even hundreds of billions of parameters. Their high computational requirements and memory usage limit their deployment on edge devices, mobile devices, and resource-constrained environments, becoming a key bottleneck restricting their practical application.
[0003] To address this issue, researchers have proposed a variety of model compression techniques, including low-bit quantization, pruning, knowledge distillation, and low-rank decomposition. Low-bit quantization has attracted considerable attention because it significantly reduces computational and storage requirements without modifying the network structure. Post-training quantization (PTQ) is particularly noteworthy because it allows pre-trained models to be directly quantized without requiring retraining, significantly saving time and resources and becoming a key method for the efficient deployment of large models.
[0004] Mixed-precision quantization is crucial in quantization techniques. It uses different bit counts for different weights based on their sensitivity. Specifically, critical weights (such as the Q, K, and V projection weights of self-attention) significantly impact model performance, and low-bit quantization can lead to a significant performance drop. Non-critical weights (such as scaling factors in layer normalization) have lower precision requirements, and high-bit quantization wastes resources. Despite the obvious advantages of mixed-precision quantization, traditional methods face challenges in practical applications. Due to the extremely complex weight distribution of large-scale language models (LLMs), traditional mixed-precision quantization methods lack the ability to accurately distinguish between key and non-key weights, making it difficult to scientifically evaluate the impact of each weight on word sequence prediction tasks. This directly leads to quality issues such as a significant decrease in accuracy and insufficient semantic coherence when generating long text content, and poor adaptability to resource-constrained scenarios such as edge devices. Word sequence prediction tasks in different scenarios exhibit significant differences in data distribution. For example, news summary generation has high requirements for conciseness, while dialogue systems focus more on fluency. In addition, the diversified design of model architectures, from hierarchical attention to global attention, results in different degrees of influence of each weight on model performance. However, traditional quantization methods cannot dynamically adjust the number of quantization bits for each weight based on these differences, resulting in rigid quantization strategies, inefficient resource utilization, and an inability to simultaneously meet high-performance prediction requirements and computing resource constraints.
[0005] These shortcomings make it difficult for traditional mixed-precision quantization to strike a balance between performance and resource efficiency, resulting in poor performance in the actual task of word sequence prediction. Therefore, designing an adaptive mixed-precision quantization strategy that can accurately identify weight sensitivity and optimize resource allocation while dynamically adapting to different task requirements is of great research value in improving the deployment efficiency of LLMs in word sequence prediction scenarios. Summary of the Invention
[0006] The purpose of this invention is to address the model efficiency bottleneck in word sequence prediction tasks by proposing a large-model mixed-precision quantization method driven by multiple weight saliency metrics. In text generation tasks, existing mixed-precision quantization methods struggle to accurately identify key and non-key parameters in self-attention weights, resulting in logical errors or semantic deviations in the generated content during low-bit quantization. Furthermore, in dialogue system deployment scenarios, differences in data distribution of different user inputs further exacerbate resource allocation imbalances. To this end, this invention dynamically allocates bit width by introducing a simulated annealing algorithm. Within a mixed-precision framework, the probability density function of the weight matrix is calculated to quantize the weight sparsity distribution. The weight matrix is then classified by saliency using an optimal split point strategy based on matrix sparsity. This approach achieves precise protection of key weights and efficient resource utilization in word sequence prediction tasks, alleviating the performance degradation of traditional methods caused by their inability to adapt to complex weight distributions. It also overcomes the performance fluctuations of the model under different data features and architectural designs caused by the lack of a dynamic adjustment mechanism, ultimately improving the inference speed and prediction accuracy of large language models in practical deployments.
[0007] In order to achieve the above object, the technical solutions specifically adopted by the present invention are as follows:
[0008] A multi-weighted saliency driven large model mixed precision quantization word sequence prediction method includes the following steps:
[0009] Step 1: Obtain the question-answering dataset and divide it into a training set and a test set in proportion. Convert the text data in the training set and the test set into token ID sequences.
[0010] Step 2: Build a large language model loaded with benchmark parameters and quantize the large language model to obtain the target function;
[0011] Step 2-1: Store the parameters of the large language model in layers, block the parameter weights of different layers by column, and obtain the weight matrix of each layer. The layers used for block processing include the self-attention query projection layer, the self-attention key projection layer, the self-attention value projection layer, the first linear transformation layer of the feedforward layer, and the second linear transformation layer of the feedforward layer;
[0012] Step 2-2: Calculate the Hessian matrix of the parameter weights of different blocks and obtain the inverse matrix corresponding to each block through Cholesky decomposition and inversion. Calculate the significance of the weight of each corresponding block based on the inverse matrix to obtain a significance list and divide it into high significance weights, medium significance weights, and low significance weights.
[0013] Step 2-3: Calculate the probability density function of the weight in each block to determine the weight sparsity distribution, and use the percentile method to search for the optimal split point of the matrix sparsity to divide the weight matrix of the corresponding block into the mask of the corresponding block;
[0014] Steps 2-4: quantize the weights in each block separately. If the weights are low-significance, they are quantized at low bits; if the weights are high-significance, they are quantized at high bits. The remaining weights are quantized at intermediate bits. Finally, the quantized parameters are used as the benchmark parameters of the large language model to obtain the objective function.
[0015] Step 3: Input the tokenID sequence in the test set into the objective function, predict the next word sequence, and obtain the optimal word sequence prediction result based on the probability distribution.
[0016] Preferably, the saliency calculation method is:
[0017] Calculate the Hessian matrix, the calculation formula is as follows:
[0018]
[0019] Where H represents the Hessian matrix, and the input vector x∈R t×m From the training set, where P is the number of samples, x [k] Represents the kth input sample vector.
[0020] Cholesky decomposition and inversion, Cholesky decomposition of the Hessian matrix can obtain the lower triangular matrix L, which satisfies H block =LL T By inverting the decomposed matrix, that is
[0021] The weight matrix of a block after block processing in step 2-1 is defined as W block , the corresponding inverse matrix block is defined as The calculation method is as follows:
[0022]
[0023] Where S represents the saliency of the block, n is the number of weight elements in the block, and W block (i) represents the i-th weight element in the block, is the i-th diagonal element of the corresponding block of the inverse matrix.
[0024] As an advantage, the method for dividing the significance weight is as follows: if S i ≥T high , then the weight w i Belongs to high significance weight; if S i ≤T low , then the weight w i Belongs to low significant weight; if T low i <Thigh , then the weight w i Belongs to the middle significant weight, where S i represents the significance value of the ith position in the significance list, T high represents the high significance weight threshold, T low Represents the low significance weight threshold.
[0025] Preferably, the high significance weight threshold is calculated as follows:
[0026]
[0027] The calculation method of the low significance weight threshold is:
[0028]
[0029] Here, n refers to the total number of weight parameters in a block, r1 refers to a threshold for high-saliency weights, and r2 refers to a threshold for low-saliency weights.
[0030] Preferably, the step 2-2 further includes iteratively updating and optimizing the significance weight division result through a simulated annealing algorithm.
[0031] Preferably, in step 2-3, the probability density function of each weight block is calculated for the high-significance weights, intermediate-significance weights and low-significance weights divided according to step 2-2, so as to determine the weight sparsity distribution, which includes non-significant dense weights, non-significant sparse weights and significant weights.
[0032] Preferably, in step 2-3, the method for determining the optimal segmentation point is:
[0033] Use the quantization error formula to determine the optimal split point. Assuming that the non-significant weights are in the interval [-m, m], first, the value range covered by the uniform quantizer is [X min ,X max ], the number of quantized intervals M is usually equal to 2 b , where b is the target quantization bit width and the quantization step size Δ is:
[0034]
[0035] The boundaries of the quantization interval are expressed as:
[0036] b q =X min +Δ·l
[0037] Among them, l∈0,1,…,M, the mean of each interval is: x q =X min+Δ·l-0.55Δ; thus, the quantization error can be obtained, and the calculation formula is as follows:
[0038]
[0039] When the split point p is used to divide the non-significant weights into sparse areas and dense areas, the new error measure calculation formula is as follows:
[0040]
[0041] Among them, A s and A c Represents the weight set of sparse area and dense area respectively, w s and w c is the weight element of the corresponding area, and are the means of the weights of sparse and dense areas respectively;
[0042] The goal is to find a p-value such that Minimum, this value is the optimal split point p * ,Right now:
[0043]
[0044] Use the percentile search method to traverse different p values within a certain value range and calculate the corresponding Finally determine the optimal split point p * .
[0045] Preferably, the weight matrix is divided according to the optimal segmentation point to obtain three masks mask1, mask2, and mask3, which correspond to the non-significant dense weights, non-significant sparse weights, and significant weights in each block, respectively. Then mask1, mask2, and mask3 are stored.
[0046] Preferably, in step 2-4, the low-bit quantization method is: using the matrix masks mask1, mask2, and mask3 related to the low-saliency weights obtained in step 2-3 to perform processing, and then performing quantization calculation on the weights of different sparsities: let the quantized binary weight matrix be B low , the scaling factor is α low , then the binary quantization calculation formula is as follows:
[0047] B low =α low sign(W low )
[0048]
[0049] Using the residual approximation strategy, first calculate the initial binary quantization result B 0,low and α 0,low , the calculation formula is as follows:
[0050]
[0051] Then we get the residual Binarize the residual again to get B r,low and α r,low , the calculation formula is as follows:
[0052]
[0053] The final quantitative results are:
[0054]
[0055] According to the above quantization process, W 1,low ,W 2,low ,W 3,low Quantitative calculations were performed to obtain the results Then add them together to get the final quantitative result of low significance weight:
[0056]
[0057] As an example, the intermediate bit quantization method is as follows: let the quantized weight be Q mid , first calculate the zero point Z mid and offset value Δ m id, assuming the quantization range is [X min,mid ,X max,mid ], the number of quantization bits is b mid , the number of quantized intervals Then the quantization step size Δ m The calculation formula for id is as follows:
[0058] Δ mid =X min,mid -X max,mid
[0059] Zero point Z mid The calculation method is:
[0060]
[0061] The round function is used to perform rounding operations, using the zero point Z mid , offset value Δ mid Perform preliminary quantization on the weights to obtain the preliminary quantization result Q 0,mid , the specific formula is as follows:
[0062]
[0063] Among them, the clip(x,a,b) function is used to limit the value of x to the range of [a,b] and calculate the residual matrix R before and after quantization. mid , the calculation formula is as follows:
[0064] R mid =W mid -(Δ mid ·(Q 0,mid -Z mid ))
[0065] Using the residual approximation optimization strategy, the scaling factor of the first calculation is set to The quantified results are recorded as The parameters for the second calculation are First calculate and The calculation formula is as follows:
[0066]
[0067] Then calculate the residual Then quantize the residual twice and calculate The calculation formula is as follows:
[0068]
[0069] The final quantitative results are:
[0070]
[0071] According to the above quantization process, W 1,mid ,W 2,mid ,W 3,mid Quantitative calculations were performed to obtain the results Then add them together to get the final quantitative result of the intermediate significance weight:
[0072] Preferably, the high-bit quantization method is the same as the middle-bit quantization method.
[0073] Preferably, in step 1, the method for obtaining the tokenID sequence is as follows: the original text in the dataset is processed into a token sequence through a word segmenter, and each token sequence is converted into a digital ID, the IDs are formed into an input tensor, and special tags [CLS] and [SEP] are added to identify sequence boundaries. Finally, a tokenID sequence of uniform length is obtained by padding and truncating the sequence.
[0074] Preferably, step 3 includes the following sub-steps:
[0075] Step 3-1: Input the token ID sequence into the target model and convert it into a high-dimensional vector representation through the embedding layer. Each ID will be converted into a 768-dimensional vector representation;
[0076] Step 3-2: Add position encoding information to each position represented by the vector, and finally form an input tensor with position information;
[0077] Step 3-3: Perform attention calculation in the Transformer layer on the input tensor with position information;
[0078] Steps 3-4: After all Transformer layers have processed the last layer, the output is a tensor of shape [1, 10, vocab_size], where each element output[0][t][i] represents the probability of the i-th word in the vocabulary appearing as the next word at the t-th position in the input sequence. This tensor fully records the model's prediction of the next word for each position in the entire input sequence.
[0079] Preferably, in step 3-3, the parameter precisions of the query projection layer, key projection layer, value projection layer, first linear transformation layer of the feedforward layer, and second linear transformation layer of the feedforward layer in the quantized Transformer layer have been secondary grouped according to the significance of the parameters, with 3-bit precision used for high-significance parameters, 2-bit precision used for medium-significance parameters, and 1-bit precision used for low-significance parameters.
[0080] Preferably, in step 3-3, the method for calculating attention in the Transformer layer is:
[0081] First, the input features are projected into the Query space, Key space, and Value space through linear transformation. The attention weights are calculated using the Query, Key, and Value. The Value is weighted and summed according to the attention weights. The attention output is then passed through the first fully connected layer for feature expansion. A nonlinear activation function is applied to the output of the first linear transformation layer of the feedforward layer. The activated result is restored to its original dimension through the second fully connected layer, resulting in a tensor that integrates contextual information and has enhanced expressiveness through nonlinear transformation.
[0082] The present invention has the following characteristics and beneficial effects:
[0083] The present invention can effectively solve the key problem of model deployment in word sequence prediction tasks, alleviate the problems of existing mixed-precision quantization methods that are difficult to adapt to complex weight distributions and lack a dynamic adjustment mechanism, resulting in the inability to achieve optimal resource allocation and the imbalance between performance and resource utilization under different data features and model architectures. At the same time, it overcomes the limitations of traditional methods in determining weight significance and allocating bits, especially in resource-constrained environments such as edge computing, and improves the inference efficiency of large language models while maintaining the accuracy and response speed of word sequence prediction.
[0084] 1) Accuracy is allocated through a simulated annealing strategy, achieving dynamic and precise accuracy grouping. This dynamic adjustment mechanism can escape local optimal solutions during the search process, more accurately capture differences in weight importance, adapt to complex weight distributions, and achieve optimal resource allocation under different data characteristics and model architectures.
[0085] 2) Design a saliency grouping strategy that adapts to the sparsity distribution of weights. High, medium, and low saliency weights are regrouped according to their Gaussian distribution. This allows the quantization process of different saliency weights after grouping to better adapt to different sparsity weights and reduce quantization errors.
[0086] 3) A more accurate representation of weights is achieved through a residual approximation strategy that performs a second-order approximation on the weights. This strategy can reduce quantization error and minimize quantization loss of weights when storing them at very low bit widths. It also balances accuracy and storage, avoiding the increase in average weight bit count caused by directly storing weights. This reduces the number of bits required for storage while maintaining model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 4 is a flow chart of an algorithm in an embodiment of the present invention.
[0088] Figure 2 This is a diagram briefly introducing the quantization model in an embodiment of the present invention.
[0089] Figure 3 Detailed architecture diagram of the quantization model in an embodiment of the present invention.
[0090] Figure 4 Result diagram of different group sizes of ablation experiment 1 of the present invention.
[0091] Figure 5 This is a result diagram of different bit width allocation strategies in the second ablation experiment of the present invention.
[0092] Figure 6 This is a graph showing the results of different quantization calculation strategies for the ablation experiment three of the present invention. DETAILED DESCRIPTION
[0093] The present invention is described in detail below in conjunction with specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0094] A large model mixed precision quantization word sequence prediction method driven by multiple weighted significance, such as Figure 1 As shown, the following steps are included:
[0095] Step 1: Obtain the question-answering dataset and convert the text data in the question-answering dataset into a tokenID sequence.
[0096] Specifically, in this embodiment, the original text "Chapter 1. The catsat on the mat." in the Wikitext-2 dataset is used and processed into a token sequence by a tokenizer. The tokenization result is ["Chapter","1",".","The","cat","sat","on","the","mat","."], and each token is converted into a numerical ID, such as [123,45,6,789,1011,1213,1415,1617,1819,20]. These IDs form an input tensor. Special tags such as [CLS] and [SEP] are added to mark sequence boundaries, and the sequence is padded or truncated to ensure uniform length.
[0097] Step 2: Build a large language model loaded with benchmark parameters and quantize the large language model to obtain the objective function. This includes the following sub-steps:
[0098] Step 2-1: Store the parameters of the large language model in layers, block the parameter weights of different layers by column, and obtain the weight matrix of each layer.
[0099] In this embodiment, the parameters of the benchmark large model are selected for loading. Taking the OPT model as an example, the self-attention query projection layer, self-attention key projection layer, self-attention value projection layer, first linear transformation layer of the feedforward layer, and second linear transformation layer of the feedforward layer, which are crucial to its performance, are stored in layers. This is because the self-attention mechanism-related layers play a core role in the processing of natural language by the large language model. The self-attention query projection layer, self-attention key projection layer, and self-attention value projection layer are responsible for generating the key information required for attention calculation, helping the model focus on important content and capture long-distance dependencies. The attention output layer integrates the calculation results, which directly affects the model's understanding and utilization of contextual information. The first linear transformation layer of the feedforward layer and the second linear transformation layer of the feedforward layer bear the important responsibilities of feature transformation and mapping, and play a decisive role in the quality of the text generated by the model. At the same time, the parameters of these layers account for a large proportion of the total model parameters. Quantizing them can significantly reduce the model's storage requirements and computational complexity, and their calculation mode is relatively regular, suitable for hardware acceleration. After quantization, hardware resources can be better utilized and computing efficiency can be improved. Other layers, however, are not critical to the model's core functionality, have complex parameter distributions and importance patterns, and are difficult to quantize. Quantization contributes little to model performance improvements and may increase hardware implementation complexity and reduce computational efficiency. Therefore, choosing to store these specific layers in layers provides a clear structure and defined goals for subsequent quantization operations, helping to achieve efficient model compression and fast inference.
[0100] It is understandable that in order to achieve efficient and accurate quantization effects, it is necessary to refine the weights of different layers of the model. Specifically, the weights of different layers are first grouped or blocked according to the column dimension, that is, the weights are divided into multiple groups or blocks in the column direction. The purpose of this operation is to better capture the characteristic differences of the weights in different dimensions. After the grouping / blocking is completed, the weights in the same block are quantized using unified parameters. This method can ensure that the weight quantization process within the same group remains consistent, avoiding the hardware calculation burden brought by the discrete allocation of bit width in traditional methods. Some traditional mixed precision methods require discrete allocation of bit width on the entire weight matrix, which not only increases the complexity of hardware calculation, but also affects the inference efficiency. The group quantization method can reduce this burden and improve the inference efficiency.
[0101] Step 2-2: Calculate the Hessian matrix of the parameter weights of different blocks and obtain the inverse matrix corresponding to different blocks through Cholesky decomposition and inversion. Calculate the significance of the weight of each corresponding block based on the inverse matrix to obtain a significance list and divide it into high significance weights, medium significance weights, and low significance weights.
[0102] Specifically, the steps for calculating the Hessian matrix, decomposing and inverting the matrix through Cholesky, and solving the significance of the weight matrix block are as follows:
[0103] Calculate the Hessian matrix, which reflects the quadratic approximation of the weight parameters on the model output loss and provides basic data for subsequent calculations of weight significance. In the post-training quantization (PTQ) scenario of large language models (LLMs), the Hessian matrix is generated using the Levenberg-Marquardt approximation, as is common practice. The calculation formula is as follows:
[0104]
[0105] Where H represents the Hessian matrix, and the input vector x∈R t×m From the training set, where P is the number of samples, x [k] Represents the kth input sample vector.
[0106] Cholesky decomposition and inversion, since the Hessian matrix needs to be operated in the following embodiment, the problem of zero diagonal elements and numerical instability may be encountered during the operation, which will seriously affect the accuracy and stability of the algorithm. To solve these problems, the Cholesky decomposition and inversion method is used in this embodiment to obtain the inverse matrix of the Hessian matrix. The Cholesky decomposition of the Hessian matrix can obtain the lower triangular matrix L, which satisfies H block =LL T By inverting the decomposed matrix, that is A more stable inverse matrix form can be obtained. This method is more reliable in numerical calculations and can effectively avoid algorithm errors caused by numerical instability.
[0107] The weight matrix of a block after the division in step 2-1 is W block , the corresponding inverse matrix block is For this block, calculate the sum of the ratio of the square of its weight to the square of the diagonal elements of the corresponding block in the inverse matrix, and use this as the saliency of the block. The calculation formula is as follows:
[0108]
[0109] Where S represents the saliency of the block, and n is the number of weight elements in the block. block (i) represents the i-th weight element in the block, The saliency calculated in this way can effectively measure the importance of each block in model quantization and provide a key basis for subsequent quantization decisions.
[0110] Furthermore, the significance is initialized and divided. In order to further optimize the quantization effect, the weights in the significance list obtained in step 3 need to be preliminarily divided. Since different weights have different degrees of influence on the performance of the large language model, they are divided into high significance weights, intermediate significance weights and low significance weights according to the distribution characteristics of the significance values. Specifically, the weights with significance values in the first r1% (such as r1=20) are divided into high significance weights, the weights in the last r2% (such as r2=20) are divided into low significance weights, and the weights in the middle part (accounting for 1-r1%-r2%) are classified as intermediate significance weights. Here, by setting the threshold T high and T low To achieve the division, the calculation formula is as follows:
[0111]
[0112] Among them S (k) Represents the significance value of the kth position after sorting the significance list S. If S i ≥T high , then the weight w i Belongs to high significance weight; if S i ≤T low , then the weight w i Belongs to low significant weight; if T low i <T high , then the weight w i It has an intermediate significant weight.
[0113] Further, such as Figure 2 and Figure 3 As shown, the simulated annealing algorithm is used to iteratively update and optimize the saliency weight division results, and the optimal weighted saliency division method is found and stored. The specific steps are as follows:
[0114] Initialize the simulated annealing algorithm parameters, which include the initial temperature T0, the cooling coefficient α, the number of iterations n, and the termination temperature T final . The initial temperature T0 is set to 1.0, which determines the search range of the algorithm in the initial stage and the probability of accepting a worse solution. A higher initial temperature enables the algorithm to explore in a wider solution space and avoid falling into the local optimal solution too early, but it also increases the amount of calculation; a lower initial temperature can speed up the convergence, but it may cause the algorithm to miss the global optimal solution. The cooling coefficient α is set to 0.95, which controls the rate of temperature drop, ensuring that the algorithm gradually narrows the search range and approaches the optimal solution. The number of iterations n is set to 50, which controls the number of algorithm iterations, as long as the temperature T>T final , the algorithm will iterate according to the number of iterations. The termination temperature T final Set to 1e-3 to determine whether the algorithm stops searching. When the temperature drops to this value, the algorithm ends the iteration.
[0115] The simulated annealing algorithm is iteratively updated. During the iterative process of the simulated annealing algorithm, each iteration attempts to adjust the current weight significance division method. Based on the current division state, one or more weights are randomly selected and their significance categories are changed (i.e., from one category of high, medium, and low significance weight categories to another category), thereby generating a new division scheme. The objective function value under the new division scheme is calculated. In the quantization scenario of this method, the objective function uses the relative entropy of the weight output calculated based on the Kullback-Leibler (KL) divergence to measure the information difference between the quantized weight and the original weight. The calculation formula is as follows:
[0116]
[0117] Among them D kl (·||·) represents the KL divergence, x is the input vector of the training set, w f is the original weight, Indicates that according to the partitioning scheme [g1,…g n ] for weight w f The quantified result, g i Represents the quantization bit width of the i-th group.
[0118] The Metropolis criterion is used to determine whether to accept the new partitioning scheme: if the objective function value of the new scheme is better than the current scheme (i.e., the objective function value decreases), the new scheme is directly accepted; if the objective function value of the new scheme is worse than the current scheme (i.e., the objective function value increases), the new scheme is accepted with a certain probability. This probability is related to the current temperature T and is calculated as follows:
[0119]
[0120] Where ΔE is the difference between the objective function values of the new solution and the current solution. As the iteration progresses, the temperature is reduced by the cooling coefficient α, that is, T = αT, so that the algorithm gradually reduces the probability of accepting poor solutions and focuses more on finding the local optimal solution.
[0121] The optimal partitioning method is determined and stored, and the simulated annealing algorithm is iteratively updated according to the number of iterations n until the temperature drops to the termination temperature T finalThe partitioning scheme obtained at this time is the optimal weight-significance partitioning method under the current conditions. This optimal partitioning method is stored, and the stored content includes the significance category of each weight (high significance weight, medium significance weight, or low significance weight). This stored information will play a key role in the subsequent quantization step. In this way, more accurate quantization of large language models can be achieved, effectively improving the performance of the model in low-bit conditions and laying a solid foundation for efficient model deployment.
[0122] Step 2-3: Calculate the probability density function of the weight in each block for the high-significance weight, medium-significance weight, and low-significance weight divided in step 2-2 to determine the weight sparsity distribution. Use the percentile method to search for the optimal split point of the matrix sparsity to divide the weight matrix of the corresponding block into mask1, mask2, and mask3. The specific steps are as follows:
[0123] Calculate the probability density function of the weights in each block. When processing each weight matrix, assume that the weight matrix block currently being processed is W and its elements are w ij , i=1,2,…,n,j=1,2,…,m. Mathematically, Gaussian distribution is defined as: if the random variable X obeys a mathematical expectation of μ and variance of σ 2 Normal distribution, denoted as X~N(μ,σ 2 ), its probability density function is:
[0124]
[0125] Among them, μ determines the center position of the distribution, that is, the mean; σ determines the degree of dispersion of the distribution. In the LLMs weight analysis scenario, Gaussian distribution can be used to approximately describe the distribution of weights. This is because the actual observed weight distribution is similar to the Gaussian distribution in characteristics. Most weight values are concentrated near a certain central value. The farther away from the central value, the lower the probability of the weight value appearing.
[0126] Based on this, in this step, the probability density function g(x) of the non-significant weight is calculated by kernel density estimation (KDE). Taking the Gaussian kernel as an example, its formula is:
[0127]
[0128] Among them, n is the number of weights involved in the calculation in the current block, x i is the value of the i-th weight, h is the bandwidth parameter, which controls the smoothness of the kernel function, is a Gaussian kernel function. Through this formula, the present embodiment can obtain the probability density function of the weight in each block, and then determine the sparsity distribution of the weight, providing a basis for subsequent division of the weight matrix.
[0129] Determine the optimal segmentation point. Based on the probability density function g(x) obtained in the previous step, determine the optimal segmentation point p * To divide the weight matrix. This method uses the quantization error formula to determine the optimal split point. When performing weight division, the goal of this embodiment is to achieve effective compression while preserving the weight information as much as possible, and the quantization error is the key indicator to measure the degree of achievement of this goal. Assuming that the non-significant weights are in the interval [-m, m], first, the value range covered by the uniform quantizer is [X min ,X max ], the number of quantized intervals M is usually equal to 2 b (b is the target quantization bit width), then the quantization step size Δ is:
[0130]
[0131] The boundaries of the quantization interval can be expressed as:
[0132] b q =X min +Δ·l
[0133] Among them, l∈0,1,…,M, the mean of each interval is: x q =X min +Δ·l-0.55Δ
[0134] From this, the quantization error can be obtained, and the calculation formula is as follows:
[0135]
[0136] Since g(x) is a symmetric function, the above formula is simplified to:
[0137]
[0138] When the split point p is used to divide the non-significant weights into sparse areas and dense areas, the new error measure calculation formula is as follows:
[0139]
[0140] Among them, A s and A c Represents the weight set of sparse area and dense area respectively, w s and w c is the weight element of the corresponding area, and are the means of the weights of the sparse and dense regions, respectively.
[0141] The goal in this example is to find a p-value such that Minimum, this value is the optimal split point p * ,Right now:
[0142]
[0143] Use the percentile search method to traverse different p values within a certain value range and calculate the corresponding , finally determine p * .
[0144] Divide the weight matrix according to the determined optimal split point p * , divide the weight matrix. For the current weight matrix block W, the division rules are as follows:
[0145] Mask mask1 corresponding to non-significant dense weights: if |w ij ∣≤p * And the weight is not a significant weight, then mask1 ij =1, otherwise mask1 ij =0.
[0146] Mask mask2 corresponding to non-significant sparse weights: if |w ij ∣>p * And the weight is not a significant weight, then mask2 ij =1, otherwise mask2 ij =0.
[0147] Mask mask3 corresponding to the significant weight: If the weight is a significant weight, then mask3 ij =1, otherwise mask3 ij =0.
[0148] In this way, the weight matrix is divided into three masks mask1, mask2, and mask3, which correspond to the non-significant dense weights, non-significant sparse weights, and significant weights in each block respectively, and then mask1, mask2, and mask3 are stored.
[0149] Steps 2-4: Quantize the weights in each block separately. If the weights are low-significance, perform low-bit quantization; if the weights are high-significance, perform high-bit quantization. The remaining weights perform intermediate-bit quantization. Finally, the quantized parameters are used as the benchmark parameters of the large language model to obtain the objective function.
[0150] Specifically, hierarchical saliency quantization operates on different layers within the larger model. Then, based on the significance of the weights, hierarchical quantization applies different quantization operations to the high-, low-, and intermediate-significance weights from step 2-2. High-significance weights are quantized with high bits; low-significance weights are quantized with low bits (e.g., binarization); and the remaining weights are quantized with intermediate bits.
[0151] For low-bit quantization, binary quantization is used, and the low-saliency weight-related matrix masks mask1, mask2, and mask3 obtained in steps 1-3 are used for processing. Taking mask1 as an example (mask2 and mask3 are processed in the same way), the low-saliency weight sub-matrix W is obtained by element-wise multiplication: 1,low =W low ⊙mask1. This step can be used to transform the original matrix W according to the mask matrix low The weights of different sparsity are accurately selected. Then the weights of different sparsity are quantized: let the quantized binary weight matrix be B low , the scaling factor is α low , then the binary quantization calculation formula is as follows:
[0152] B low =α low sign(W low )
[0153]
[0154] In order to reduce the quantization error, the residual approximation strategy is adopted. First calculate the initial binary quantization result B 0,low and α 0,low , the calculation formula is as follows:
[0155]
[0156] Then we get the residual Binarize the residual again to get B r,low and α r,low , the calculation formula is as follows:
[0157]
[0158] The final quantitative results are:
[0159]
[0160] According to the above quantization process, W 1,low ,W 2,low ,W 3,low Quantitative calculations were performed to obtain the results Then add them together to get the final quantitative result of low significance weight:
[0161] For the intermediate bit quantization, the matrix masks mask1, mask2, and mask3 related to the intermediate saliency weights obtained in steps 2-3 are processed to obtain W 1,mid ,W 2,mid ,W 3,mid The specific operation has been introduced in 6.2 and will not be repeated here. Then the weights of these different sparsities are quantized: let the quantized weight be Q mid , first calculate the zero point Z mid and offset value Δ m id. Assume that the quantization range is [X min,mid ,X max,mid ], the number of quantization bits is b mid , the number of quantized intervals Then the quantization step size Δ m The calculation formula of id is as follows:
[0162] Δ mid =X min,mid -X max,mid
[0163] Zero point Z mid The calculation method is:
[0164]
[0165] The round function is used to perform rounding operations, using the zero point Z mid , offset value Δ m id performs preliminary quantization on the weights and obtains the preliminary quantization result Q 0,mid , the calculation formula is as follows:
[0166]
[0167] The clip(x,a,b) function is used to limit the value of x to the range [a,b]. Calculate the residual matrix R before and after quantization mid :
[0168] R mid =W mid -(Δ mid ·(Q 0,mid -Z mid ))
[0169] In order to further reduce the quantization error, the residual approximation optimization strategy is adopted. Assume that the scaling factor of the first calculation is The quantified results are recorded as The parameters for the second calculation are First calculate and The calculation formula is as follows:
[0170]
[0171] Then calculate the residual Then perform secondary quantization on the residual and calculate
[0172]
[0173] The final quantitative results are:
[0174]
[0175] According to the above quantization process, W 1,mid ,W 2,mid ,W 3,mid Quantitative calculations were performed to obtain the results Then add them together to get the final quantitative result of the intermediate significance weight:
[0176] It should be noted that, in this embodiment, the high-bit quantization method is the same as the intermediate-bit quantization method, and the quantization operation is the same, only the number of quantization bits is different. Therefore, in this embodiment, the quantization operation is introduced by taking the intermediate-bit quantization as an example.
[0177] Step 3: Input the token ID sequence into the objective function, predict the next word sequence, and obtain the optimal word sequence prediction result based on the probability distribution.
[0178] Specifically, in this embodiment, these token ID sequences are first input into the target model and converted into high-dimensional vector representations through the embedding layer. Each ID will be converted into a 768-dimensional vector representation. For example, token ID 123 will be mapped to [0.12, -0.34, 0.56, ..., 0.78] (a total of 768 values), and the entire sequence will become a tensor with a shape of [1, 10, 768].
[0179] Position encoding information is added to each position. For example, the first token "Chapter" will be encoded with position 0 [0.01, -0.02, ...], and the second token "1" will be encoded with position 1 [0.03, -0.01, ...]. The dimension of the position encoding is consistent with the word vector (such as 768 dimensions), so that the model can capture word meaning information and perceive the relative position relationship of tokens in the sequence. This is crucial for correctly understanding phrases that rely on word order such as "the cat sat on the mat", and ultimately form an input tensor with position information.
[0180] The Transformer layer performs attention calculations on the positional input tensor. After quantization in step 6, the parameter precision of the query projection layer, key projection layer, value projection layer, first linear transformation layer of the feedforward layer, and second linear transformation layer of the feedforward layer has been regrouped based on parameter significance. High-significance parameters use 3-bit precision, medium-significance parameters use 2-bit precision, and low-significance parameters use 1-bit precision. First, the input features are projected into the query space, key space, and value space using linear transformations. Attention weights are calculated using the query, key, and value, and the value is weighted summed according to the attention weights. The attention output is then passed through the first fully connected layer for feature expansion. A nonlinear activation function is applied to the output of the first linear transformation layer of the feedforward layer. The activated result is restored to its original dimension through a second fully connected layer, resulting in a tensor that incorporates contextual information and has enhanced expressiveness through nonlinear transformations.
[0181] Predict the next most likely word. After all Transformer layers are processed, the last layer outputs a tensor of shape [1, 10, vocab_size], where each element output[0][t][i] represents the probability of the i-th word in the vocabulary appearing as the next word at the t-th position in the input sequence. This tensor fully records the model's prediction of the next word at each position in the entire input sequence. For example, at position t=4 corresponding to "cat", output[0][4] is a vector containing vocab_size elements, where the element corresponding to the word ID "sat" has the highest value, indicating that the model predicts that the probability of "cat" followed by "sat" is the highest. Similarly, the entire tensor fully describes the model's prediction of the probability distribution of the next word at each position in the input sequence, and ultimately selects the word ID with the highest probability at each position as the prediction result.
[0182] The perplexity of the prediction task is evaluated. Perplexity measures the uncertainty of the model's prediction of the next word. The lower the perplexity, the more accurate the prediction task. The system compares the distribution of the next word predicted by the model with the next word in the actual text, calculates the cross-entropy loss, and finally summarizes the perplexity score of the entire test set as the evaluation result. The calculation formula is as follows:
[0183]
[0184] Where N is the number of samples in the dataset, P(w i ) is the probability of the model predicting the i-th word in the sample.
[0185] The effects of the present invention will be further described below in conjunction with simulation experiments.
[0186] The simulation experiment of the present invention is aimed at the word sequence prediction task, and the dataset used is WikiText2, which is a text dataset extracted from Wikipedia articles for language modeling, containing about 2 million words. The experiment quantized the word sequence prediction OPT model by 3 bits, focusing on the impact of different quantization methods and parameter configurations on the perplexity of word sequence prediction. In the experimental framework, per-channel group quantization is adopted, and 128 is set as the group size. We randomly selected 128 samples from WikiText2 as calibration data, each sample containing 2048 tags. By allowing the model to perform word sequence prediction based on these sample data, the effects of different quantization methods in actual prediction are compared to verify the effectiveness of this method in the word sequence prediction task. In addition, we also conducted three ablation experiments. The first ablation experiment was to study the impact of different group sizes on the word sequence prediction results. The performance was evaluated at the 3-bit level when the group size was 128 columns, 256 columns, and 512 columns, focusing on verifying the effect of different weight group sizes on the word sequence prediction perplexity; the second ablation experiment was to explore the effect of different bit width allocation strategies on the word sequence prediction effect. The performance of the GPTQ allocation strategy, head-tail allocation, random allocation, SliM-LLM allocation strategy, and the simulated annealing strategy allocation used by this method were evaluated at the 3-bit level to verify the advantages of the simulated annealing strategy allocation used by this method in the word sequence prediction task; the third ablation experiment was to clarify the impact of the quantization calculation method in the method on word sequence prediction. The performance of different quantization calculation methods for low-bit quantization and different combinations of medium / high-bit quantization calculation methods were evaluated at the 3-bit level to verify the improvement effect of the residual approximate quantization method used in this method on the word sequence prediction task.
[0187] Table 1: Comparison of perplexity of different methods
[0188]
[0189] Result Analysis
[0190] The simulation experiments of the present invention employed the present invention and a prior art SliM-LLM quantization framework. The SliM-LLM quantization framework, proposed by Wei Huang et al. in "SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models," is a mixed-precision framework based on saliency-determined bit allocation (SBA) and saliency-weighted quantizer calibration (SQC).
[0191] In order to verify the effect of the simulation experiment of the present invention in the word sequence prediction task, the OPT model was used for comprehensive verification. The experimental results are shown in Table 1. In terms of word sequence prediction accuracy, the present method shows significant advantages. In the model OPT-1.3b, the perplexity of the original precision model (FP16) is 14.63, while the perplexity of the present method is similar to that of the original precision model by as much as 98.49%, which fully verifies the ability of the present method to retain the generation quality. Compared with the SliM-LLM method, the perplexity of the present method in word sequence prediction is reduced by 7.01%. Compared with other quantization methods such as RTN, OmniQuant, AWQ, and GPTQ, the perplexity is reduced by 87.52%, 7.64%, 9.01%, and 9.83%, respectively. Especially in the long sequence generation task, the present method effectively reduces the context break problem caused by low-bit quantization through precise weight significance classification, thereby improving the prediction performance.
[0192] In the larger model OPT-2.7b, the perplexity of the original precision model (FP16) is 12.47. The perplexity of this method is still as close as 97.91% to that of the original precision model, demonstrating excellent quantization compatibility. Compared with the SliM-LLM method, the perplexity is reduced by 3.99%. Compared with other quantization methods, the perplexity is reduced by 95.75%, 3.41%, 6.26%, and 7.01%, respectively. Especially in some high-precision semantic conversion tasks, this method significantly reduces the precision loss of key parameters through dynamic bit width allocation and second-order residual quantization strategies. Especially on low-resource devices, this method saves storage resources while improving generation efficiency by reducing the impact of quantization noise on word order-sensitive tasks.
[0193] In the first ablation experiment, Figure 4As can be seen in the figure, in model OPT-1.3b, we evaluated the performance of 128 columns, 256 columns, and 512 columns at the 3-bit level and observed that larger group sizes can improve GPU efficiency during inference. The results show that increasing group granularity does not significantly increase the perplexity of the model, which shows that this method is robust and conducive to more efficient deployment methods.
[0194] In the second ablation experiment, Figure 5 It can be seen that in the model OPT-1.3b, the GPTQ allocation strategy has the worst effect, while the head-tail allocation is to set the head group to low bits and select the same number of tail groups to high bits. The random allocation is to randomly select the same number of low / high bit groups for bit width allocation. One of these two strategies is too fixed, and the other is too random. SliM-LLM uses the head-tail strategy based on the calculation of the significance matrix, and the effect is significantly improved. This method uses the simulated annealing strategy for bit width allocation to achieve dynamic and accurate precision grouping. It captures the weight importance differences more accurately than the previous allocation strategy, adapts to complex weight distribution, and achieves optimal resource allocation under different data characteristics and model architectures.
[0195] In the third ablation experiment, Figure 6 As can be seen in the OPT-1.3b model, the traditional binary quantization method / xnor binary quantization method combined with the traditional high-bit quantization method has a mediocre effect. However, introducing residual approximation into binary quantization can achieve a more accurate representation of weights, and this strategy can reduce quantization error. This method further introduces residual approximation into the medium / high-bit quantization method, further reducing quantization error and further balancing accuracy and storage, reducing the number of storage bits while maintaining model accuracy.
[0196] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A large-model mixed-precision quantized word sequence prediction method driven by multiple weighted saliency, characterized by: The steps include: Step 1: Obtain the question-answering dataset and divide it into a training set and a test set in proportion. Convert the text data in the training set and the test set into token ID sequences. Step 2: Build a large language model loaded with benchmark parameters and quantize the large language model to obtain the target function; Step 2-1: Store the parameters of the large language model in layers, block the parameter weights of different layers by column, and obtain the weight matrix of each layer; Step 2-2: Calculate the Hessian matrix of the parameter weights of different blocks and obtain the inverse matrix corresponding to each block through Cholesky decomposition and inversion. Calculate the significance of the weight of each corresponding block based on the inverse matrix to obtain a significance list and divide it into high significance weights, medium significance weights, and low significance weights. Step 2-3: Calculate the probability density function of the weight in each block to determine the weight sparsity distribution, and use the percentile method to search for the optimal split point of the matrix sparsity to divide the weight matrix of the corresponding block into the mask of the corresponding block; Steps 2-4: quantize the weights in each block separately. If the weights are low-significance, they are quantized at low bits; if the weights are high-significance, they are quantized at high bits. The remaining weights are quantized at intermediate bits. Finally, the quantized parameters are used as the benchmark parameters of the large language model to obtain the objective function. Step 3: Input the token ID sequence in the test set into the objective function, predict the next word sequence, and obtain the optimal word sequence prediction result based on the probability distribution.
2. The method according to claim 1, characterized in that The significance calculation method is: Calculate the Hessian matrix, the calculation formula is as follows: Where H represents the Hessian matrix, and the input vector x∈R t×m From the training set, where P is the number of samples, x [k] Represents the kth input sample vector; Cholesky decomposition and inversion, Cholesky decomposition of the Hessian matrix can obtain the lower triangular matrix L, which satisfies H block =LL T , by inverting the decomposed matrix, that is The weight matrix of a block after block processing in step 2-1 is defined as W block , the corresponding inverse matrix block is defined as The calculation method is as follows: Where S represents the saliency of the block, n is the number of weight elements in the block, and W block (i) represents the i-th weight element in the block, is the i-th diagonal element of the corresponding block of the inverse matrix.
3. The method according to claim 1, characterized in that The method for dividing the significance weight is: if S i ≥T high , then the weight w i Belongs to high significance weight; If S i ≤T low , then the weight w i Belongs to low significance weight; If T low i <T high , then the weight w i Belongs to the middle significant weight, where S i represents the significance value of the ith position in the significance list, T high represents the high significance weight threshold, T low Represents the low significance weight threshold. 4. The method according to claim 3, characterized in that The high significance weight threshold is calculated as follows: The calculation method of the low significance weight threshold is: Here, n refers to the total number of weight parameters in a block, r1 refers to a threshold for high-saliency weights, and r2 refers to a threshold for low-saliency weights.
5. The method according to claim 4, characterized in that The step 2-2 further includes iteratively updating and optimizing the significance weight division result through a simulated annealing algorithm.
6. The method according to claim 5, characterized in that In step 2-3, the probability density function of each block of weights is calculated for the high-significance weights, intermediate-significance weights, and low-significance weights divided according to step 2-2, thereby determining the weight sparsity distribution, which includes non-significant dense weights, non-significant sparse weights, and significant weights.
7. The method according to claim 6, characterized in that In step 2-3, the method for determining the optimal segmentation point is: Use the quantization error formula to determine the optimal split point. Assuming that the non-significant weights are in the interval [-m, m], first, the value range covered by the uniform quantizer is [X min ,X max ], the number of intervals after quantization M is usually equal to 2 b , where b is the target quantization bit width and the quantization step size Δ is: The boundaries of the quantization interval are expressed as: b q =X min +D·l Among them, l∈0,1,…,M, the mean of each interval is: x q =X min +Δ·l-0.55Δ; thus, the quantization error can be obtained, and the calculation formula is as follows: When the split point p is used to divide the non-significant weights into sparse areas and dense areas, the new error measure calculation formula is as follows: Among them, A s and A c Represents the weight set of sparse area and dense area respectively, w s and w c is the weight element of the corresponding area, and are the means of the weights of sparse and dense areas respectively; The goal is to find a p-value such that Minimum, this value is the optimal split point p * ,Right now: Use the percentile search method to traverse different p values within a certain value range and calculate the corresponding Finally determine the optimal split point p * .
8. The method according to claim 7, characterized in that According to the optimal segmentation point, the weight matrix is divided into three masks mask1, mask2, and mask3, which correspond to the non-significant dense weights, non-significant sparse weights, and significant weights in each block respectively. Then mask1, mask2, and mask3 are stored.
9. The method according to claim 8, characterized in that In step 2-4, the low-bit quantization method is: using the matrix masks mask1, mask2, and mask3 related to the low-significance weights obtained in step 2-3 to process, and then performing quantization calculation on the weights of different sparsities: let the quantized binary weight matrix be B low , the scaling factor is α low , then the binary quantization calculation formula is as follows: B low =α low sign(W low ) Using the residual approximation strategy, first calculate the initial binary quantization result B 0,low and α 0,low , the calculation formula is as follows: Then we get the residual Binarize the residual again to get B r,low and α r,low , the calculation formula is as follows: The final quantitative results are: According to the above quantization process, W 1,low ,W 2,low ,W 3,low Quantitative calculations were performed to obtain the results Then add them together to get the final quantitative result of low significance weight:
10. The method according to claim 1, characterized in that The method of quantizing the intermediate bits is: assuming the weight after quantization is Q mid , first calculate the zero point Z mid and offset value Δ m id, assuming the quantization range is [X min,mid ,X max,mid ], the number of quantization bits is b mid , the number of quantized intervals Then the quantization step size Δ m The calculation formula for id is as follows: Δ mid =X min,mid -X max,mid Zero point Z mid The calculation method is: The round function is used to perform rounding operations, using the zero point Z mid , offset value Δ mid Perform preliminary quantization on the weights to obtain the preliminary quantization result Q 0,mid , the specific formula is as follows: Among them, the clip(x,a,b) function is used to limit the value of x to the range of [a,b] and calculate the residual matrix R before and after quantization. mid , the calculation formula is as follows: R mid =W mid -(Δ mid ·(Q 0,mid -Z mid )) Using the residual approximation optimization strategy, the scaling factor of the first calculation is set to The quantified results are recorded as The parameters for the second calculation are First calculate and The calculation formula is as follows: Then calculate the residual Then quantize the residual twice and calculate The calculation formula is as follows: The final quantitative results are: According to the above quantization process, W 1,mid ,W 2,mid ,W 3,mid Quantitative calculations were performed to obtain the results Then add them together to get the final quantitative result of the intermediate significance weight:
11. The method according to claim 10, characterized in that The high-bit quantization method is the same as the middle-bit quantization method.
12. The method according to claim 1, characterized in that In step 1, the token ID sequence is obtained by processing the original text in the dataset into a token sequence through a word segmenter, then converting each token sequence into a digital ID, forming the IDs into an input tensor, and adding special tags [CLS] and [SEP] to identify sequence boundaries. Finally, the sequence is padded and truncated to obtain a uniform-length token ID sequence.
13. The method according to claim 12, characterized in that The step 3 includes the following sub-steps: Step 3-1: Input the token ID sequence into the target model and convert it into a high-dimensional vector representation through the embedding layer. Each ID will be converted into a 768-dimensional vector representation; Step 3-2: Add position encoding information to each position represented by the vector, and finally form an input tensor with position information; Step 3-3: Perform attention calculation in the Transformer layer on the input tensor with position information; Steps 3-4: After all Transformer layers have processed the last layer, the output is a tensor of shape [1, 10, vocab_size], where each element output[0][t][i] represents the probability of the i-th word in the vocabulary appearing as the next word at the t-th position in the input sequence. This tensor fully records the model's prediction of the next word for each position in the entire input sequence.
14. The method according to claim 13, characterized in that In step 3-3, the parameter accuracies of the query projection layer, key projection layer, value projection layer, first linear transformation layer of the feedforward layer, and second linear transformation layer of the feedforward layer in the quantized Transformer layer have been secondary grouped according to the significance of the parameters. Parameters with high significance use 3-bit precision, parameters with medium significance use 2-bit precision, and parameters with low significance use 1-bit precision.
15. The method according to claim 13, characterized in that In step 3-3, the attention calculation method in the Transformer layer is: First, the input features are projected into the Query space, Key space, and Value space through linear transformation. The attention weights are calculated using the Query, Key, and Value. The Value is weighted and summed according to the attention weights. The attention output is then passed through the first fully connected layer for feature expansion. A nonlinear activation function is applied to the output of the first linear transformation layer of the feedforward layer. The activated result is restored to its original dimension through the second fully connected layer, resulting in a tensor that integrates contextual information and has enhanced expressiveness through nonlinear transformation.
Citation Information
Patent Citations
Hybrid precision neural network quantification method and device,equipment
CN114049530A
Large language model software and hardware collaborative quantitative accelerated calculation method and system
CN117574976A
Lightweight vehicle-mounted target detection model and mixing precision quantification method thereof
CN118397598A
Large language model efficient reasoning method and system based on parallel decoding
CN118627629A
Remote sensing information processing method based on deep learning training
CN119314036A