Language task processing method and electronic equipment

By analyzing task calibration data to determine key weight parameters, and using fitted quantization parameters and optimal scaling factors to quantize the language model, the problem of high storage and computing resource requirements and low accuracy of large-scale language models on mobile terminals and edge computing nodes is solved, thereby improving the accuracy of task processing.

CN120873148AActive Publication Date: 2025-10-31LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Patent Information

Application Number
CN202511376899.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2025-10-31
Estimated Expiration
2045-09-25

AI Technical Summary

Technical Problem

Existing technologies for deploying large-scale language models on mobile terminals or edge computing nodes suffer from high storage and computing resource requirements, and the quantization bit width leads to low accuracy in language task processing results.

Method used

By analyzing task calibration data, the target weight parameters that have a significant impact on the accuracy of task processing results are identified. The gradient is obtained by fitting the quantization parameters, and the optimal scaling factor is found. The target weight parameters are then scaled and quantized based on the quantization accuracy to reduce storage and computing resource requirements.

Benefits of technology

While reducing resource requirements, it effectively improves the accuracy of language task processing, solves the problem of precision loss caused by quantization bit width, and realizes systematic and automated weight scaling and quantization error compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873148A_ABST
    Figure CN120873148A_ABST
Patent Text Reader

Abstract

The invention discloses a language task processing method and electronic equipment, and relates to the technical field of artificial intelligence. The method comprises the following steps: selecting a target weight parameter according to task calibration data and an activation function type of an original task processing model; according to each target weight parameter, the maximum weight parameter and the task calibration data, determining a fitting processing result of the quantization parameter; and determining a first quantization error function and a second quantization error function carrying a scaling factor according to a fitting processing result, and obtaining a numerical value of the scaling factor by minimizing a ratio of the first quantization error function and the second quantization error function. And performing scaling processing on each target weight parameter by using the scaling factor, and quantifying each scaling weight parameter based on the quantization precision parameter to obtain a task processing model for executing the to-be-processed language task. According to the method and the device, the problem that high accuracy of a language task processing result cannot be ensured on the basis of reducing resources used in a language task execution process in related technologies can be solved, and the language task processing precision is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a language task processing method and electronic device. Background Technology

[0002] With the rapid development of artificial intelligence technology, language models are becoming increasingly large, making them unsuitable for deployment in resource-constrained scenarios such as mobile terminals or edge computing nodes.

[0003] Related technologies directly convert the parameters of trained language models into a low-precision format, eliminating the need for retraining or fine-tuning the model. While this significantly reduces the demand for storage and computing resources, it still struggles to meet users' high-precision requirements in scenarios with extremely high parameter counts. Summary of the Invention

[0004] This invention provides a language task processing method and electronic device, which can not only reduce the demand for storage and computing resources, but also solve the problem of low accuracy of language task processing results caused by quantization bit width, and effectively improve the accuracy of language task processing.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a language task processing method, comprising: Based on task calibration data matching the language task to be processed and the activation function type of the original task processing model, target weight parameters that affect the task processing results are selected from the weight parameters of the original task processing model. The fitting result of the quantization parameters is determined based on each target weight parameter, the maximum weight parameter of the original task processing model, and the task calibration data. A first quantization error function and a second quantization error function carrying a scaling factor are determined based on the fitting result. The scaling factor is obtained by minimizing the ratio of the first quantization error function to the second quantization error function. Each target weight parameter is scaled using the scaling factor, and each scaled weight parameter is quantized based on the quantization accuracy parameter to obtain the task processing model used to perform the language task to be processed.

[0006] The present invention also provides an electronic device, including a memory and a processor, wherein the processor is configured to implement the steps of any of the above-described language task processing methods when executing a computer program stored in the memory.

[0007] The advantages of the technical solution provided by this invention lie in its ability to analyze the output of task calibration data matching the language task to be processed, determine the target weight parameters that significantly impact the accuracy of the task processing results, and scale these target weight parameters. This reduces the computational and storage resources required during language task execution, effectively minimizing errors caused by weight parameter quantization and improving the accuracy of language task processing results. Furthermore, by fitting the quantization parameters to obtain the gradient and optimizing the scaling factor, a systematic and automated weight scaling process is achieved, compressing the quantization error range and improving the task execution performance of the quantized language task processing model. This not only reduces the demand for storage and computational resources but also solves the problem of low accuracy in language task processing results due to quantization bit width, effectively improving language task processing precision. In addition, this invention provides corresponding electronic equipment for the language task processing method, further enhancing its practicality. The electronic equipment offers corresponding advantages. Attached Figure Description

[0008] To more clearly illustrate the technical solutions of the present invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 A schematic diagram of the hardware framework applicable to the language task processing method provided by the present invention; Figure 2 A flowchart illustrating a language task processing method provided by the present invention; Figure 3 A schematic diagram of the neural network provided by this invention; Figure 4 A schematic diagram of the quantization results provided by this invention; Figure 5 A schematic diagram of fitting the quantization factor provided by the present invention in an exemplary scenario; Figure 6 A schematic diagram illustrating the fitting of the quantization function provided by the present invention in an exemplary scenario; Figure 7 A flowchart illustrating another language task processing method provided by the present invention; Figure 8 This is a structural framework diagram of an exemplary embodiment of the language task processing device provided by the present invention; Figure 9 This is a structural diagram of an exemplary embodiment of the electronic device provided by the present invention. Detailed Implementation

[0010] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. In this specification and the aforementioned drawings, the terms "first," "second," "third," "fourth," etc., are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. The term "exemplary" means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.

[0011] In recent years, the number of parameters in deep learning models has grown exponentially, with some advanced language models having hundreds of billions or even trillions of parameters. These models rely on computing resources and massive storage resources provided by high-performance GPUs (Graphics Processing Units) and TPU (Tensor Processing Units) clusters during both training and inference phases. This results in high hardware deployment costs and limitations in data transmission bandwidth and memory access latency, making it difficult to run on mobile terminals or edge computing nodes. To overcome this resource bottleneck, model quantization compression technology has been proposed. By quantizing 32-bit floating-point numbers into 8-bit integer representations, it significantly reduces storage space requirements and computational complexity while maintaining model performance to some extent. Although the storage and computational costs of the model decrease with decreasing bit width, quantization bit widths (such as 8 bits, 4 bits, etc.) introduce information loss, leading to a decrease in model output accuracy. For large models with hundreds of billions of parameters, converting high-precision floating-point parameters to low-precision values ​​can cause a significant drop in model accuracy. To address this issue, PQT (Post-Quantization Training) and PTQ (Post-Training Quantization) are currently used for model compression. PQT optimizes model performance through further training after quantization, which can restore performance to some extent. However, for large models, PQT not only increases cost but also easily leads to overfitting. PTQ, on the other hand, quantizes the model after training, directly performing numerical conversion on the trained model. This converts model parameters from high-precision (e.g., 32-bit floating-point numbers) to low-precision (e.g., 8-bit integers) format without retraining or fine-tuning, saving significant time and resources, making it more suitable for compressing large-scale language models. However, related technologies that use PQT to compress large models, such as GPTQ (Gradient-based Post-training Quantization) and AWQ (Activation-aware Weight Quantization), although they alleviate accuracy decay to some extent, introduce errors during the quantization process, leading to a decrease in model performance. In scenarios with extremely high parameter counts, they still cannot meet users' high-precision requirements.

[0012] To further improve the language task processing accuracy of the task processing model after PTQ with extremely low computational resource consumption, this invention analyzes the output results of task calibration data that matches the language task to be processed, determines the target weight parameters that have a significant impact on the accuracy of the task processing results, obtains the gradient by fitting the quantization parameters, and uses the optimal scaling factor obtained by optimization to scale these target weight parameters. The weight parameters of the original task processing model with hundreds of billions / hundreds of parameters that have been trained are quantized with low bit width, reducing storage volume and computational overhead, and mitigating the accuracy loss caused by the reduction of quantization bit width.

[0013] This section describes the specific application environment architecture or hardware architecture upon which the execution of language task processing methods depends. The following section will further elaborate on this. Figure 1 Examples of possible application scenarios related to the technical solutions of this invention are provided below: The hardware framework may include a high-performance server 11, an edge computing device 12, and a user terminal 13. The high-performance server 11 and the edge computing device 12 are connected via a network. The high-performance server 11 can pre-train a large-scale task processing model for executing the language task to be processed, and send the task processing model to the edge computing device 12. The edge computing device 12 receives the task processing model as the original task processing model and deploys a processor for executing the language task processing method described in any of the above embodiments. The user terminal 13 is deployed to provide a human-computer interaction interface and is also connected to the edge computing device 12 via a network. The edge computing device 12 receives the language task to be processed sent by the user terminal 13 via the network, obtains task calibration data matching the language task to be processed from a public database through an external interface, obtains a task processing model by executing the language task processing method described in any of the above embodiments, processes the language task to be processed using the task processing model, and sends the final task processing result to the user terminal via the network.

[0014] It should be noted that the above application scenarios are only shown to facilitate understanding of the ideas and principles of the present invention, and the embodiments of the present invention are not limited in any way. On the contrary, the embodiments of the present invention can be applied to any applicable scenario. After introducing the technical solution of the present invention, various non-limiting embodiments of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0015] Please see first. Figure 2 , Figure 2 This is a flowchart illustrating a language task processing method provided in this embodiment. This embodiment may include the following: S201: Based on the task calibration data that matches the language task to be processed and the activation function type of the original task processing model, select the target weight parameters that affect the task processing results from the weight parameters of the original task processing model.

[0016] The language task to be processed can be any natural language task, including but not limited to text classification tasks (assigning one or more predefined categories or labels to a text, such as topic classification and spam identification), sentiment analysis tasks (identifying and extracting subjective information in text, such as opinions, sentiments, and emotions), entity recognition tasks (identifying named entities with specific meanings in text and classifying them into predefined categories), relation extraction tasks (identifying semantic relationships between entities in text), semantic role labeling (analyzing the relationship between predicates and related components in a sentence), and natural language generation tasks (machine translation, text summarization, dialogue generation, human-computer question answering, code generation and interpretation), etc. The original task processing model refers to a large-scale language model that has been trained and can directly perform the task to be processed, which can be any large language model or multimodal language model. Task calibration data can be training sample data used to train the original task processing model to handle the language task to be processed, or a small portion of the training sample data for fine-tuning. It can also be historical data used by other language models to process the same language task to be processed. It can be obtained through public databases. For example, for tasks such as text classification and sentiment analysis, task calibration data consists of text data, audio and video data, and image data with sentiment type labels or text type labels. For natural language generation tasks, task calibration data consists of a large amount of plain text. The model constructs its own input (previous text) and output (next word) through self-supervised learning. For question-answering tasks, task calibration data consists of a set of text data, audio and video data, and image data with a question-answer-reasoning process.

[0017] In this embodiment, the original task processing model includes multiple network layers. The output of the previous layer serves as the input to the next layer, and so on until the last layer. The vector of the input layer is generally called the input vector, the intermediate layers are called hidden layers (each corresponding to a hidden layer vector), and the vector of the output layer is called the output vector. Each layer has multiple neurons (i.e., nodes), such as... Figure 3As shown, each line corresponds to a weight parameter. Generally, multiple nodes in the same layer can correspond to a vector. After inputting the task calibration data, the calculation is completed through multiple weighted lines and passed to the next layer. The output of the previous layer becomes the input of the next layer, and so on until the last layer. The more nodes and the deeper the neural network, the stronger the expressive power. The original task processing model in this embodiment has at least hundreds of billions of parameters, thus requiring extremely high computing power. This step evaluates the importance of each weight parameter in the original task processing model. The importance evaluation considers the activation function type and is achieved by analyzing the contribution of the weight to the model output, i.e., the task processing result. For example, the gradient information of the weight, the absolute value of the weight, or the sensitivity of the weight can be used to measure the importance of the weight. The target weight parameter is the weight parameter that has a significant impact on the accuracy of the task processing result. The more important the weight, the more carefully its quantification process needs to be handled to avoid a significant negative impact on the model performance.

[0018] S202: Based on the target weight parameters, the maximum weight parameter of the original task processing model, and the task calibration data, determine the fitting result of the quantization parameters. Based on the fitting result, determine the first quantization error function and the second quantization error function carrying a scaling factor. The scaling factor is obtained by minimizing the ratio of the first quantization error function to the second quantization error function.

[0019] The quantization parameters include a quantization factor and a quantization function. The quantization factor and quantization function can be determined using any model quantization method. The quantization error function can be represented using the corresponding quantization error function determination method. Any optimization algorithm can be used to find the optimal value of the function ratio relationship; this invention does not impose any limitations on this. The value of the scaling factor obtained through optimization calculation is the optimal scaling factor.

[0020] For example, a general quantization function It can be determined by the following relationship: ; Where w is the weighting parameter, N is the number of quantization bits (e.g., INT8 is quantization to 8 bits), Δ is the quantization factor, Round is the rounding function (usually rounding to the nearest integer), and |w| represents the absolute value of w. The quantization error function introduced by the general quantization method... It can be represented as: x is the input element value, and RoundErr represents the rounding of the error function.

[0021] Quantization methods based on scaling factors, such as AWQ (Activation-aware Weight Quantization), use the following quantization function: ; Where x is the input element value, It is the quantization factor corresponding to this method. , This introduces scaled parameters. The quantization error function after introducing the scaling factor can be expressed as follows: .

[0022] Only when the quantization factor is the same before and after scaling, that is... Furthermore, RoundErr is independent of internal variables (always within the range [0, 0.5]). Only quantization based on scaling factors, compared to general quantization, can reduce the error to 1 / s of the original value. However, current technologies, during the scaling process, extract the column with the largest activation value (w) and multiply it by s. This could potentially alter the largest w, thus affecting... The quantization factor differs before and after scaling. Furthermore, when the parameters are purely random, the average error of the Round function is 0.25, which can be considered a constant. However, the parameters are often not random but rather exhibit a regular distribution; RoundErr is usually not 0.25, so RoundErr is often correlated with internal variables. To further reduce the error introduced by quantization through scaling, it is necessary to find the optimal value of the scaling factor. Since both the quantization function and the quantization factor are step functions, they are difficult to optimize using gradient methods. This step involves fitting the quantization parameters; that is, the purpose of fitting the step function is to approximate the quantization factor and the quantization function, thereby achieving the gradient optimization process of the loss function and obtaining the optimized scaling factor. The fitting results include whether the quantization factor has been approximated or fitted, and whether the quantization function has been fitted.

[0023] S203: Scale each target weight parameter using a scaling factor, and quantize each scaled weight parameter based on the quantization precision parameter to obtain a task processing model for performing the language task to be processed.

[0024] In the scaling process using the scaling factor determined in the previous step, to accommodate different hardware environments, a scaling scale parameter can be set based on the actual deployment environment resource configuration and the user's task accuracy requirements. The scaling scale parameter can be set manually or automatically generated by fine-tuning any large language model capable of dialogue. During scaling, this scaling scale parameter can be obtained. A number of weight parameters matching the scaling scale parameter are selected from each target weight parameter. Each weight parameter to be scaled is then multiplied by the scaling factor to obtain the scaling weight parameters. For example, if the target weight parameter selected in S201 is column K, the scaling scale parameter can be M, where M is less than K. For situations with ample computing resources, multiple M values ​​can be set, and the M value that best results in the quantized task processing model can be selected. For situations with limited computing resources, M can generally be manually specified.

[0025] Among them, the quantization precision parameter is the preset quantization precision, which corresponds to the data precision of the parameters of the large language model or multimodal language model in the computer. It is generally in the format of FP32, FP16, BF16, FP8, INT8, FP4, etc. FP32 (single precision floating point) is a floating point number format that uses 32 bits of binary representation. The specific structure is as follows: (1) Sign bit: 1 bit, used to indicate the positive or negative value. (2) Exponent bit: 8 bits, used to indicate the range of the value. (3) Fractional part: 23 bits, used to indicate the precision of the value. The numerical range of FP32 is approximately 1.18e-38 to 3.40e-38, and the precision is approximately 6 to 9 significant digits. It provides a good balance between numerical range and precision, and is also well supported by hardware. Generally speaking, the quantization precision can be set to W8A16, W4A16, W8A8, W4A4, etc. W corresponds to the precision of the weight parameters in a large or multimodal language model, and A corresponds to the precision of the activation vectors in the same model. Quantization precision parameters may include, for example, the number of quantization bits, the quantization range (i.e., the minimum and maximum values, which can be determined by analyzing task calibration data and user requirements), and the quantization step size. For important weights, a higher number of quantization bits (e.g., 8 bits or higher) may be chosen to retain more precision information; while for less important weights, a lower number of quantization bits (e.g., 4 bits or lower) can be used to further reduce storage and computational overhead. The quantization step size maps high-precision values ​​to low-precision values ​​by dividing the quantization range by the range represented by the low precision. For example, in W8A8 quantization, if the original range of the weights is from -1.0 to 1.0, the quantization step size is 2.0 divided by 255 (for unsigned integers) or divided by 256 (for signed integers). This step size is used to convert the weight values ​​from floating-point numbers to 8-bit integers.

[0026] After scaling the target weight parameters selected in S201, this step can quantize the original task processing model using any quantization method described in the relevant documentation. This involves converting the original task processing model and activation values ​​from high precision (e.g., 32-bit floating-point numbers) to low precision (e.g., 8-bit integers) to reduce the storage space and computational complexity of the task processing model when performing the language task, while maintaining the accuracy of the task processing results output by the model. The original image texture is clearly visible, but storing such data occupies a large amount of memory. Therefore, it can be quantized to only have black and white colors. Figure 4 As shown, this significantly reduces the image storage space without affecting target recognition. After quantizing the original task processing model, the resulting model is defined as the task processing model in this step. The task processing model is used to execute the language task to be processed and output the corresponding task processing result. If a similar language task to be processed is received subsequently, the task processing model obtained in S203 can be used directly to execute the corresponding task.

[0027] In the technical solution provided in this embodiment, by analyzing the output results of task calibration data matching the language task to be processed, the target weight parameters that have a significant impact on the accuracy of the task processing results are determined. These target weight parameters are then scaled, which effectively reduces the errors caused by weight parameter quantization while decreasing the computational and storage resources required during language task execution, thereby improving the accuracy of the language task processing results. Furthermore, by fitting the quantization parameters to obtain the gradient and optimizing the scaling factor, a systematic and automated weight scaling is achieved, compressing the quantization error range and improving the task execution performance of the quantized language task processing model. This not only reduces the demand for storage and computational resources but also solves the problem of low accuracy in language task processing results due to quantization bit width, effectively improving the precision of language task processing.

[0028] In the above embodiments, no limitation is made on how to obtain the task calibration data of the language task to be processed. Based on the above embodiments, the present invention also provides an exemplary implementation method, which may include the following: If the task type of the language task to be processed has a public dataset, a predetermined number of language sample data are selected from the public dataset to construct a task calibration dataset. If the task type of the language task to be processed does not have a public dataset, a corresponding number of language sample data are obtained from multiple relevant databases according to a predetermined ratio based on the task characteristics of the language task to be processed, to serve as the task calibration dataset.

[0029] In this embodiment, a dataset generally corresponds to a collection of sample data, which consists of samples and labels. For large language models, samples can be combinations of content such as code, formulas, books, articles, and web pages, while labels can be category labels or content such as code, formulas, books, articles, and web pages. For multimodal language models, samples are generally a combination of multimodal data and text data, such as images and corresponding image descriptions. For models with publicly available datasets, it is generally assumed that the publicly available dataset is the pre-training dataset. A small amount of data can be randomly extracted from this dataset according to the classification as a calibration set X (e.g., if the dataset contains three sub-datasets A, B, and C, 0.01% of the data can be extracted from each sub-dataset). For models with unpublished datasets, the type of task processing model corresponds to the task characteristics, i.e., the task type, such as a language model or a multimodal model. For example, if the language task to be processed is a visual language task, and the corresponding task processing model is a visual language model, then the corresponding task calibration data is a combination of image and text data. The required data can be divided into three categories: College-level Problems, Math, and General Visual Question Answering. Then, multiple open-source datasets are selected in each category. For example, in Math, three datasets are selected: Math Vista, MATH-Vision, and Math Verse. A certain number of data (e.g., 0.01%) are randomly extracted from each dataset to form the task calibration dataset.

[0030] As can be seen from the above, this embodiment sets different methods for generating task calibration data according to different task types. By improving the correlation between task calibration data and the language task to be processed, it is beneficial to accurately select target weight parameters that have a significant impact on task processing results, thereby ensuring that quantization loss is reduced and improving the accuracy of task processing results.

[0031] The above embodiments do not limit how to select the target weight parameters that affect the task processing results. Based on the above embodiments, the present invention also provides an exemplary implementation method, which may include the following: If the activation function type is a parameterless activation function, the task calibration data is input into the original task processing model. During one forward propagation of the original task processing model, the matrix multiplication results of the weight parameters of each network layer and the corresponding input task data are obtained. A preset number of target matrix multiplication results are selected from the matrix multiplication results in descending order, and the weight parameters corresponding to each target matrix multiplication result are used as target weight parameters. If the activation function type is a parameterized activation function, the task calibration data is input into the original task processing model. During one forward propagation of the original task processing model, the weight parameters and gating parameters of each network layer are multiplied element-wise with the product of the corresponding input task data. A preset number of target matrix multiplication results are selected from the matrix multiplication results in descending order, and the weight parameters corresponding to each target matrix multiplication result are used as target weight parameters.

[0032] In this embodiment, the activation function is the type of function used by the task processing model to determine the task processing result. The type of activation function can be directly determined based on the formula or graph of the activation function. Activation function types include parameterless activation functions and parameterized activation functions. Parameterless activation functions include Sigmoid, Tanh, and ReLU (Rectified Linear Unit). Parameterized activation functions include GLU (Gated Linear Unit) and SwiGLU (Swish-Gated Linear Unit).

[0033] For nonparametric activation functions x corresponds to the input data to the model; in this example, it is task calibration data. Act represents the activation function. For large language models, it is text data; for multimodal language models, it is image + text data. w is the model weight, and h is the output value of the activation function of the large or multimodal language model, generally used for subsequent processing, or directly corresponding to the output text (or image + text). Under a specific input x, the magnitude of wx can reflect the response strength of w to the current input, thus indirectly reflecting the importance of w under that input. It can be directly used... The magnitude of w represents the importance of the weight parameter w. The value of wx is the vector result obtained after multiplying w and x in one forward propagation. That is, it is the tensor obtained after multiplying the current layer weight tensor and the input tensor according to the specified dimensions in the forward propagation. It can be obtained using the framework's built-in one forward propagation without additional training or optimization. As long as the model structure, weights w, and input x are determined, the value of wx is uniquely determined. The top m parameters (the value can be flexibly determined based on the actual situation, ranging from 0.1% to 3% of the total parameters) are taken as the target weight column. For cases where no other parameters are introduced, this saves on the calculation and analysis process, thereby saving computational and storage resources. For cases using parameterized activation functions... ,in This indicates element-wise multiplication. This refers to the gating parameter. In this case, since activation is also affected by the gating parameter, only the gating parameter is considered. This would result in finding non-critical weight parameters. Therefore, in this embodiment, for such scenarios, the m parameter columns corresponding to the largest values ​​h after analysis are used as the target weight columns.

[0034] As can be seen from the above, this embodiment selects weight parameters that have a significant impact on the task processing results based on different activation function types. This can further reduce the required computing and storage resources while ensuring the accuracy of weight parameter selection.

[0035] The above embodiments do not limit how the fitting result of the quantization function is determined. The present invention also provides an exemplary implementation method, which may include the following: For the quantization factor, if the maximum weight parameter of the original task processing model is one of the target weight parameters, then the quantization factor is determined based on the maximum weight parameter, the scaling factor, and the number of quantization bits. If the maximum weight parameter of the original task processing model is not any of the target weight parameters, but is greater than the product of the maximum target weight parameter and the scaling factor among the target weight parameters, then the maximum weight parameter is used as the quantization factor. If the maximum weight parameter of the original task processing model is not among the target weight parameters, and is less than the product of the maximum target weight parameter and the scaling factor among the target weight parameters, then the quantization factor is determined based on the maximum target weight parameter, the scaling factor, and the number of quantization bits, and a smooth and non-monotonic linear function with learnable parameters is used as the fitting result corresponding to the quantization factor.

[0036] For the quantization function, the quantization gradient is determined based on the difference between the quantization interval and the task calibration data; the objective polynomial is determined based on the quantization gradient and the quantization interval; and the quantization function is fitted using the objective polynomial to obtain the fitting result corresponding to the quantization function.

[0037] In this embodiment, for the quantization factor, if the maximum weight parameter of the original task processing model... If the parameters are already among the target weight parameters selected by S201, then no approximation or fitting processing is performed on them. That is, the quantization factor... If the maximum weight parameter of the original task processing model It is already among the target weight parameters selected by S201, so If the maximum weight parameter of the original task processing model If the parameter is not in the target weight parameters, then fitting is performed on a case-by-case basis. hour, It is a constant value No fitting is required when When, it can also be expressed as: ,in This refers to the largest element in the key weight column. As s increases linearly, the quantization factor is very similar in form to the ReLU function and is not differentiable. It can be approximated using a smooth, non-monotonic linear function with learnable parameters, making the quantization factor a continuously differentiable function. In other words, in the process of fitting the quantization factor, a function approximation method is used for the second case. For example, the scaling factor can be used as the independent variable, and the exponential function can be determined based on the negative product of a positive integer as the base, the learnable parameter, and the scaling factor as the exponent. The fitting result of the quantization factor is generated by using the sum of the exponential function and the target constant value as the denominator and the scaling factor as the numerator. For example, the base is the natural constant e, and the relation can be used... To approximate the quantization factor, such as Figure 5 As shown, the relation approximates the quantization factor by using a learnable parameter β, where a can be any target constant value, such as 1.

[0038] After completion After approximation, the quantization function can be fitted. Generally, a polynomial (i.e., the target polynomial) can be used for fitting to obtain a continuously differentiable fitting result. The quantization interval can be a pre-set constant. For integer quantization, this value can be 0.5 ± the fluctuation value. For example, the target polynomial can be: ;in, Represents a polynomial. Indicates the quantization interval, sign represents the sign function, and x represents the task calibration data. When this happens, sign takes the value +1. When k is used, sign takes a value of -1, and k is the power of the power chosen flexibly according to the actual scenario, with a common range of [3, 100]. A higher power of k results in a better fit, but also a steeper gradient, such as... Figure 6 As shown, this may lead to instability during the optimization process.

[0039] As can be seen from the above, this embodiment selectively approximates the quantization factor and fits the quantization function with a polynomial. The resulting quantization error function is a relation to the scaling factor, which can be used to calculate the optimal scaling factor using the gradient update method, making the implementation process simpler.

[0040] Furthermore, in this embodiment, the ratio of the first quantization error function to the second quantization error function can be used as the loss function of the original task processing model; the task calibration data is input into the original task processing model, the original task processing model is iteratively trained based on the loss function, and the value of the scaling factor when the loss function is minimized is determined.

[0041] Once the quantization factor Δ and the quantization function Q are fitted, the loss function... If a function becomes a continuously differentiable function that depends only on the scaling factor s, then s can be updated by calculating the gradient to minimize L, at which point the optimal s can be determined.

[0042] The above embodiments do not limit how the original task processing model is quantified. Based on the above embodiments, the present invention also provides an exemplary implementation, which may include the following: The maximum positive value among the positive values ​​corresponding to each original weight parameter and each scaling weight parameter can be used as the maximum positive weight parameter, and the new quantization factor can be determined based on the maximum positive weight parameter and the number of quantization bits. For each original weight parameter and each scaling weight parameter, the integer value of the ratio of the new quantization factor to the current weight parameter can be calculated, and the quantization weight parameter corresponding to the current weight parameter can be determined based on the integer value and the new quantization factor.

[0043] In this embodiment, the original weight parameters are those weight parameters in the original task processing model that are not the target weight parameters, and the scaled weight parameters are the results of scaling the target weight parameters. When determining the quantization factor, the absolute values ​​of each original weight parameter and each scaled weight parameter can be processed first, and then sorted from largest to smallest. The maximum value is selected, and for ease of distinction and description, the selected maximum value can be defined as the maximum positive weight parameter. The quantization factor can be determined according to the quantization range and quantization step size. All weight parameters are quantized. Furthermore, the activation function can also be quantized, such as by quantizing both weights and activation values ​​into 8-bit integers, that is, mapping the range of weights and activation values ​​to the range of 8-bit integers, taking into account the case of signed integers, i.e., from -128 to 127 or from 0 to 255.

[0044] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention also provides an exemplary implementation, such as... Figure 7 As shown, it may include the following: S1: Obtain the original task processing model and quantization precision parameters. The original task processing model can be represented as O, for example.

[0045] S2: Based on the original task processing model, determine that the language task samples that match the language task to be processed have a public dataset, and obtain a small number of language task samples from the public dataset to construct the task calibration dataset X.

[0046] S3: The language task samples that match the language task to be processed do not have a public dataset. A small number of language task samples are obtained from relevant open source datasets to construct the task calibration dataset X.

[0047] S4: Obtain the activation function of the original task processing model. Determine the activation function type based on its characteristics (whether it contains parameters). If it is a parameterless activation function, proceed to S5; if it is a parameterized activation function, proceed to S6.

[0048] S5: Use directly The magnitude of w represents the importance of the weight parameter w. The top m parameters with the largest wx are taken as the target weight parameters to construct the key weight column.

[0049] S6: Use the m parameter columns with the largest values ​​in the activated values ​​h as key weight columns, and then execute S7.

[0050] S7: Determine the scaling parameter as the number of columns to scale, m.

[0051] S8: Determine the scaling factor s by optimizing the loss function L through gradient optimization.

[0052] For quantification factors If the largest element It is already in the key weight column selected by S7, then No approximation is needed; the original can be used directly. Calculate the gradient. For the quantization factor... If the largest element If it is not in the key weight column, then This function can be approximated by fitting a specified continuously differentiable function. approximate . use right Perform fitting. Update s using gradient calculation, so that... When the value is minimized, the optimal value s can be determined.

[0053] S9: Multiply each target weight parameter in the m key weight columns by the scaling factor s to obtain a new weight parameter matrix w`.

[0054] S10: Perform the standard quantization process on w` to complete the quantization of the original task processing model.

[0055] As can be seen from the above, this embodiment determines the key weight columns for scaling based on the output analysis results of the task calibration data, obtains the gradient by fitting the quantization function with a univariate multivariate function, and optimizes the scaling factor through backpropagation, thereby achieving systematic and automated weight scaling and compressing the quantization error range. The fine-grained weight quantization compression method provided in this embodiment improves the accuracy of the task processing results of the task processing model while reducing resource usage.

[0056] It should be noted that there is no strict order of execution between the steps in this invention. As long as they conform to the logical order, these steps can be executed simultaneously or in a certain preset order. Figure 2 This is just an illustrative example and does not mean that this is the only possible execution order.

[0057] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0058] This invention also provides a corresponding apparatus for the language task processing method, further enhancing the practicality of the method. The apparatus can be described from both a functional module perspective and a hardware perspective. The language task processing apparatus provided by this invention is described below. This apparatus is used to implement the language task processing method provided by this invention. In this embodiment, the language task processing apparatus may include or be divided into one or more program modules. These program modules are stored in a storage medium and executed by one or more processors to complete the language task processing method disclosed in Embodiment 1. The program module referred to in this embodiment is a series of computer program instruction segments capable of performing a specific function, which is more suitable than the program itself for describing the execution process of the language task processing apparatus in the storage medium. The following description will specifically introduce the functions of each program module in this embodiment. The language task processing apparatus described below can be referred to in correspondence with the language task processing method described above.

[0059] From the perspective of functional modules, see Figure 8 , Figure 8 This is a structural diagram of the language task processing device provided in this embodiment under one specific implementation. The device may include: The key weight identification module 801 is used to select target weight parameters that affect the task processing results from the weight parameters of the original task processing model based on the task calibration data that matches the language task to be processed and the activation function type of the original task processing model.

[0060] The parameter fitting module 802 is used to determine the fitting result of the quantization parameters based on the target weight parameters, the maximum weight parameter of the original task processing model, and the task calibration data.

[0061] The scaling processing module 803 is used to determine the first quantization error function and the second quantization error function carrying the scaling factor based on the fitting processing result, obtain the value of the scaling factor by minimizing the ratio of the first quantization error function to the second quantization error function, and perform scaling processing on each target weight parameter using the scaling factor.

[0062] The quantization module 804 is used to quantize each scaling weight parameter based on the quantization precision parameter to obtain a task processing model for performing the language task to be processed.

[0063] For example, in some embodiments of this example, the parameter fitting module 802 can also be used to: if the maximum weight parameter of the original task processing model is one of the target weight parameters, then determine the quantization factor based on the maximum weight parameter, the scaling factor, and the number of quantization bits; if the maximum weight parameter of the original task processing model is not any of the target weight parameters, and is greater than the product of the maximum target weight parameter and the scaling factor among the target weight parameters, then use the maximum weight parameter as the quantization factor; if the maximum weight parameter of the original task processing model is not among the target weight parameters, and is less than the product of the maximum target weight parameter and the scaling factor among the target weight parameters, then determine the quantization factor based on the maximum target weight parameter, the scaling factor, and the number of quantization bits, and use a linear function with learnable parameters, smoothness, and non-monotonicity as the fitting result corresponding to the quantization factor.

[0064] As an exemplary implementation of the above embodiments, the parameter fitting module 802 can be further used to: determine an exponential function with the scaling factor as the independent variable and the negative number of the product of the learnable parameter and the scaling factor as the exponent, using a positive integer as the base; and generate a fitting result of the quantization factor by using the sum of the exponential function and the target constant value as the denominator and the scaling factor as the numerator.

[0065] For example, in some other embodiments of this embodiment, the parameter fitting module 802 can also be used to: determine the quantization gradient based on the difference between the quantization interval and the task calibration data; determine the target polynomial based on the quantization gradient and the quantization interval; and use the target polynomial to fit the quantization function to obtain the fitting result corresponding to the quantization function.

[0066] As an exemplary implementation of the above embodiments, the parameter fitting module 802 can be further used to: call a target polynomial to fit the quantization function; the target polynomial is: ;in, Represents a polynomial. The quantization interval is represented by , sign represents the sign function, x represents the task calibration data, and k is the power of quantization.

[0067] For example, in some other embodiments of this embodiment, the quantization module 804 can also be used to: take the maximum value among the positive values ​​corresponding to each original weight parameter and each scaling weight parameter as the maximum positive weight parameter, and determine a new quantization factor based on the maximum positive weight parameter and the number of quantization bits; calculate the integer value of the ratio of the new quantization factor to the current weight parameter for each original weight parameter and each scaling weight parameter, and determine the quantization weight parameter corresponding to the current weight parameter based on the integer value and the new quantization factor.

[0068] For example, in some other embodiments of this embodiment, the scaling processing module 803 can also be used to: obtain scaling scale parameters; select a number of weight parameters to be scaled that match the scaling scale parameters from each target weight parameter; and multiply each weight parameter to be scaled with a scaling factor to obtain each scaling weight parameter.

[0069] For example, in some other embodiments of this embodiment, the key weight identification module 801 can also be used to: if the activation function type is a parameterless activation function, input the task calibration data into the original task processing model, obtain the matrix multiplication results of the weight parameters of each network layer and the corresponding input task data during one forward propagation of the original task processing model, and select a preset number of target matrix multiplication results from each matrix multiplication result in descending order, and use the weight parameters corresponding to each target matrix multiplication result as target weight parameters; if the activation function type is a parameterized activation function, input the task calibration data into the original task processing model, and multiply the weight parameters and gating parameters of each network layer with the corresponding input task data element by element during one forward propagation of the original task processing model, and select a preset number of target matrix multiplication results from each matrix multiplication result in descending order, and use the weight parameters corresponding to each target matrix multiplication result as target weight parameters.

[0070] For example, in some other embodiments of this embodiment, the scaling processing module 803 may also be used to: use the ratio of the first quantization error function to the second quantization error function as the loss function of the original task processing model; input task calibration data into the original task processing model; perform iterative training on the original task processing model based on the loss function; and determine the value of the scaling factor when the loss function is minimized.

[0071] For a description of the features in the embodiment corresponding to the language task processing device, please refer to the relevant description in the embodiment corresponding to the language task processing method, which will not be repeated here.

[0072] The language task processing device mentioned above is described from the perspective of functional modules. Furthermore, the present invention also provides an electronic device, which is described from the perspective of hardware. Figure 9 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present invention. The electronic device includes a memory 91 and a processor 92. The memory 91 stores a computer program, and the processor 92 is configured to run the computer program to perform the steps in any of the language task processing method embodiments described above.

[0073] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described language task processing method embodiments at runtime.

[0074] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0075] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described language task processing method embodiments.

[0076] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described language task processing method embodiments.

[0077] The foregoing has provided a detailed description of a language task processing method and electronic device provided by the present invention. The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. Whether the units and algorithm steps of the various examples described in the disclosed embodiments are executed in electronic hardware or computer software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, and such implementations should not be considered beyond the scope of the present invention. Several improvements and modifications can be made to the present invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A language task processing method, characterized in that, include: Based on the task calibration data that matches the language task to be processed and the activation function type of the original task processing model, select the target weight parameters that affect the task processing results from the weight parameters of the original task processing model. Based on the target weight parameters, the maximum weight parameter of the original task processing model, and the task calibration data, the fitting result of the quantization parameters is determined; Based on the fitting results, a first quantization error function and a second quantization error function carrying a scaling factor are determined. The value of the scaling factor is obtained by minimizing the ratio of the first quantization error function to the second quantization error function. The scaling factor is used to scale each target weight parameter, and the scaling weight parameter is quantized based on the quantization accuracy parameter to obtain a task processing model for performing the language task to be processed.

2. The language task processing method according to claim 1, characterized in that, The quantization parameters include quantization factors. Based on the target weight parameters, the maximum weight parameter of the original task processing model, and the task calibration data, the fitting result of the quantization parameters is determined, including: If the maximum weight parameter of the original task processing model is one of the target weight parameters, then the quantization factor is determined based on the maximum weight parameter, the scaling factor, and the number of quantization bits. If the maximum weight parameter of the original task processing model is not any target weight parameter, and is greater than the product of the maximum target weight parameter and the scaling factor among all target weight parameters, then the maximum weight parameter is used as the quantization factor. If the maximum weight parameter of the original task processing model is not among the target weight parameters and is less than the product of the maximum target weight parameter and the scaling factor among the target weight parameters, then the quantization factor is determined according to the maximum target weight parameter, the scaling factor and the number of quantization bits, and a smooth and non-monotonic linear function with learnable parameters is used as the fitting result corresponding to the quantization factor.

3. The language task processing method according to claim 2, characterized in that, The fitting results using a smooth, non-monotonic linear function with learnable parameters as the quantization factor include: Using the scaling factor as the independent variable, an exponential function is determined based on the negative of the product of the positive integer as the base, the learnable parameter, and the scaling factor as the exponent. The fitting result of the quantization factor is generated by using the sum of the exponential function and the target constant value as the denominator and the scaling factor as the numerator.

4. The language task processing method according to claim 1, characterized in that, The quantization parameters include a quantization function. Based on each target weight parameter, the maximum weight parameter of the original task processing model, and the task calibration data, the fitting result of the quantization parameters is determined, including: The quantization gradient is determined based on the difference between the quantization interval and the task calibration data. The target polynomial is determined based on the quantization gradient and the quantization interval; The quantization function is fitted using the target polynomial to obtain the fitting result corresponding to the quantization function.

5. The language task processing method according to claim 4, characterized in that, Using the target polynomial, the quantization function is fitted, including: The quantization function is fitted using a target polynomial; the target polynomial is: ; in, Represents a polynomial. The quantization interval is represented by , sign represents the sign function, x represents the task calibration data, and k is the power of exponentiation.

6. The language task processing method according to claim 1, characterized in that, The scaling weight parameters are quantized based on the quantization precision parameter, including: The maximum value among the positive values ​​corresponding to each original weight parameter and each scaling weight parameter is taken as the maximum positive weight parameter, and a new quantization factor is determined based on the maximum positive weight parameter and the number of quantization bits. For each original weight parameter and each scaling weight parameter, calculate the integer value of the ratio of the new quantization factor to the current weight parameter, and determine the quantization weight parameter corresponding to the current weight parameter based on the integer value and the new quantization factor.

7. The language task processing method according to claim 1, characterized in that, The scaling process for each target weight parameter using the scaling factor includes: Get the scaling parameters; Select a number of target weight parameters that match the scaling scale parameter from each target weight parameter, and multiply each target weight parameter with the scaling factor to obtain each scaling weight parameter.

8. The language task processing method according to any one of claims 1 to 7, characterized in that, Select target weight parameters that influence the task processing results from the weight parameters of the original task processing model, including: If the activation function type is a parameterless activation function, the task calibration data is input into the original task processing model to obtain the matrix multiplication results of the weight parameters of each network layer and the corresponding input task data during one forward propagation of the original task processing model. A preset number of target matrix multiplication results are selected from the matrix multiplication results from large to small, and the weight parameters corresponding to each target matrix multiplication result are used as target weight parameters. If the activation function type is a parameterized activation function, the task calibration data is input into the original task processing model. During one forward propagation of the original task processing model, the weight parameters and gating parameters of each network layer are multiplied element-wise with the corresponding input task data. A preset number of target matrix multiplication results are selected from the matrix multiplication results in descending order, and the weight parameters corresponding to each target matrix multiplication result are used as target weight parameters.

9. The language task processing method according to any one of claims 1 to 7, characterized in that, The scaling factor is obtained by minimizing the ratio of the first quantization error function to the second quantization error function, including: The ratio of the first quantization error function to the second quantization error function is used as the loss function of the original task processing model. The task calibration data is input into the original task processing model, and the original task processing model is iteratively trained based on the loss function to determine the value of the scaling factor when the loss function is minimized.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the language task processing method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Deep learning network model optimization method based on parameter quantization

    CN116524173A

  • Natural language processing task execution method, device, equipment, system and medium

    CN118520849A

  • Method and device for post-training quantization based on large language model

    CN118863088A

  • Hierarchical non-training large model hybrid quantification method and system

    CN120124688A

  • Systems and methods for improving performance of a large language model by controlling training content

    US20250225400A1

Cited By

  • Method for determining reasoning result through large language model and electronic equipment

    CN121562833A