Model quantification method, device, electronic device, storage medium and program product
By quantizing the high-precision neural network model on the embedded device and converting it into low-bit integers, the problem of insufficient computing resources and memory is solved, and the operating efficiency and prediction accuracy of the model on the embedded device are improved.
Patent Information
- Application Number
- CN202510799724.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-16
AI Technical Summary
Large-scale pre-trained models deployed on embedded devices may have errors in their output results due to limited computing and memory resources.
The weight parameters and input activation values of the high-precision neural network model are converted from high-digit floating-point numbers to low-digit integers, keeping the quantization accuracy of the weights and activation values consistent, thus forming an equi-quantized neural network model.
It significantly reduces computing resource consumption, improves operating efficiency on embedded devices, maintains the model's predictive performance within a reasonable range, and reduces the error rate of output results.
Smart Images

Figure CN120317308B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a model quantization method, device, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] With the rapid development of artificial intelligence (AI), especially in natural language processing (NLP), large-scale pre-trained models have demonstrated superior performance across a variety of tasks. To expand the application scope of large models, they are now being deployed on an increasing number of embedded devices to enhance the user experience when interacting with them.
[0003] However, the large models deployed on embedded devices in related technologies are limited by computing resources and memory resources, resulting in errors in the output results. Summary of the Invention
[0004] The present application provides a model quantization method, device, electronic device, computer-readable storage medium, and computer program product to at least solve the problem of errors in the output results of large models deployed on embedded devices in the related art.
[0005] The present application provides a model quantization method, including: determining a first neural network model to be quantized, wherein the first neural network model includes a multi-layer neural network, and the weight parameters of each layer of the neural network in the multi-layer neural network and the output activation value of each layer of the neural network are both represented by floating-point numbers of a first preset number of bits; converting the weight parameters of the neural networks of multiple target layers of the first neural network model from floating-point numbers of the first preset number of bits to integers of a second preset number of bits to obtain a second neural network model, wherein the second preset number of bits is smaller than the first preset number of bits; when the second neural network model receives data to be inferred, converting the input activation values of the neural networks of the multiple target layers from floating-point numbers of the first preset number of bits to integers of the second preset number of bits; wherein the neural network of each target layer is used to output a corresponding output activation value according to the weight parameters and the input activation value.
[0006] The present application also provides a model quantization device, including: a first model determination module, used to determine a first neural network model to be quantized, wherein the first neural network model includes a multi-layer neural network, and the weight parameters of each layer of the neural network in the multi-layer neural network and the output activation value of each layer of the neural network are all represented by floating-point numbers of a first preset number of bits; a first quantization module, used to convert the weight parameters of the neural networks of multiple target layers of the first neural network model from floating-point numbers of a first preset number of bits to integers of a second preset number of bits, to obtain a second neural network model, wherein the second preset number of bits is smaller than the first preset number of bits; a second quantization module, used to convert the input activation values of the neural networks of multiple target layers from floating-point numbers of a first preset number of bits to integers of a second preset number of bits when the second neural network model receives data to be inferred; wherein the neural network of each target layer is used to output a corresponding output activation value based on the weight parameters and the input activation value.
[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of the quantization method of any of the above models when executing the computer program.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the quantization method of any of the above models are implemented.
[0009] The present application also provides a computer program product, comprising a computer program, which implements the steps of the quantization method of any of the above models when executed by a processor.
[0010] Through this application, a first neural network model to be quantized is first determined. The first neural network model is a pre-trained high-precision neural network model that can achieve high-precision reasoning and prediction result output. The weight parameters of each layer of the neural network of the first neural network model and the output activation values of each layer of the neural network are all represented by floating-point numbers of a first preset number of digits. However, due to the high precision of the first neural network model, the model calculation amount and model scale are large. The memory resources and computing resources of the embedded device are insufficient to support the first neural network model. Therefore, it is necessary to quantize the first neural network model, convert the weight parameters of the neural networks of multiple target layers of the first neural network model from floating-point numbers of a first preset number of digits to integers of a second preset number of digits, and obtain a second neural network model. In addition, the input activation values of the neural networks of multiple target layers are also converted from floating-point numbers of a first preset number of digits to integers of a second preset number of digits. Thus, by converting some of the weight parameters and activation values of the neural network model into integers of a lower number of digits, the computing resource consumption in the model inference process is significantly reduced, and the operating efficiency of the model on the embedded device or mobile device is improved. Furthermore, the weight parameters and input activation values of the neural network of the target layer are subjected to equal-precision quantization processing, so that the representation of the weight parameters and input activation values of the neural network of the target layer is kept completely consistent. Thus, although the accuracy of the output of the second neural network model after quantization is reduced, the prediction performance of the model is maintained within a reasonable range, avoiding the problem of non-convergence or garbled code caused by inconsistent quantization accuracy of the weight parameters and input activation values, and reducing the error rate of the output results of the second neural network model. Since the weight parameters and input activation values of the neural network of the target layer are subjected to equal-precision quantization processing, the errors caused by precision mismatch can be avoided, thereby achieving the technical effect of reducing the error rate of the output results of the second neural network model. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 This is a hardware structure block diagram of a server device according to a model quantization method according to an embodiment of the present application;
[0013] Figure 2 is a flow chart of a quantization method of a model according to an embodiment of the present application;
[0014] Figure 3 This is a second flow chart of a quantization method of a model according to an embodiment of the present application;
[0015] Figure 4 This is a third flow chart of a quantization method of a model according to an embodiment of the present application;
[0016] Figure 5 This is a fourth flow chart of a quantization method of a model according to an embodiment of the present application;
[0017] Figure 6 This is a fifth flow chart of a quantization method of a model according to an embodiment of the present application;
[0018] Figure 7 This is a sixth flowchart of a quantization method of a model according to an embodiment of the present application;
[0019] Figure 8 This is a seventh flowchart of a quantization method of a model according to an embodiment of the present application;
[0020] Figure 9 This is a structural block diagram of a quantization device of a model according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0022] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0023] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0024] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the quantization method of the model depends, the specific application environment architecture or specific hardware architecture is described here.
[0025] The quantization method embodiment of the model provided in the embodiment of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1This is a hardware structure diagram of a server device for a quantization method of a model in an embodiment of the present application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. The server device may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above server device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0026] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the quantization method of the model in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to the server device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0027] Transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a communication provider of the server device. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0028] The embodiment of the present application provides a model quantization method, which is applied to the above-mentioned server device, and the method is described in detail in conjunction with the execution process of the model quantization method. Figure 2 As shown, the method includes the following steps S200-220:
[0029] Step S200: Determine a first neural network model to be quantized.
[0030] The first neural network model includes a multi-layer neural network, and the weight parameters of each layer of the neural network in the multi-layer neural network and the output activation value of each layer of the neural network are both represented by floating-point numbers of a first preset number of bits.
[0031] Specifically, a trained first neural network model consisting of a multi-layer neural network is identified. The model's weight parameters and activation values are currently represented using higher-bit floating-point numbers (such as 16-bit half-precision (FP16) floating-point numbers or FP32). This provides higher prediction accuracy but places higher demands on computation and memory.
[0032] A neural network model consists of multiple interconnected neurons (nodes), with each layer responsible for a specific computational task. Weight parameters are the numerical values connecting neurons and influence the transmission of network signals. Activation values (or feature maps) are the outputs of neurons, which, after being processed by an activation function, become the inputs of the next layer of the neural network.
[0033] Step S210: Convert the weight parameters of the neural networks of the multiple target layers of the first neural network model from floating-point numbers of a first preset number of digits to integers of a second preset number of digits to obtain a second neural network model.
[0034] The second preset number of bits is smaller than the first preset number of bits.
[0035] Specifically, several layers in the model are selected as target layers, and the weight parameters of these layers are converted from high-precision floating-point numbers to integers with lower digits (such as 8-bit integers (Integer8, INT8)) to obtain the second neural network model.
[0036] Among them, quantization is the process of converting model parameters from high precision to low precision, with the aim of reducing the storage space and computing resource requirements of the model.
[0037] Step S220, when the second neural network model receives the data to be inferred, converts the input activation values of the neural networks of multiple target layers from floating-point numbers of a first preset number of bits to integers of a second preset number of bits.
[0038] Specifically, when the model performs inference, the input activation values of the target layer are also converted from floating-point numbers to integers with the same lower number of digits as the weights.
[0039] Among them, inference refers to the process of performing forward propagation to make predictions or generate outputs after the model receives new input data.
[0040] The neural network of each target layer outputs a corresponding output activation value based on the weight parameter represented as an integer of a second preset number of digits and the input activation value represented as an integer of a second preset number of digits. In the quantized model, the neural network of each target layer uses the weights and activation values represented by integers to perform calculations and output the corresponding activation value, and the entire process is performed in the integer domain of a lower number of digits.
[0041] Specifically, if the quantization precision of weights and activations is inconsistent, it means that they discard different precisions during the quantization process. For example, if weights are quantized to int8 (8-bit integers) and activations are quantized to int4 (4-bit integers), the activation quantization interval is larger, resulting in larger roundoff errors. In each layer of a neural network, weights and activations are multiplied and summed. Different quantization precisions can cause differences in the errors in the calculations at each layer. Neural networks typically have multiple layers, and computational errors in each layer are propagated to the next layer. If the quantization precision of weights and activations is inconsistent, the accumulation of errors across layers becomes complex and unpredictable. Large errors introduced by low-precision activations can be amplified after propagation through multiple layers, ultimately leading to significant deviations in model output. When weights and activations are quantized to the same precision, the errors introduced during quantization have similar magnitudes and distributions. For example, if both are quantized to int8, their quantization intervals are the same, and the range of roundoff errors is roughly the same. In each layer of a neural network, when weights and activations are multiplied, their errors can compensate or cancel each other out, rather than adding and amplifying each other. Consistent quantization precision ensures statistical consistency in the propagation of errors between layers. Because the error characteristics of weights and activations in each layer are similar, errors do not undergo sudden or erratic changes after propagating through multiple layers, making them more easily absorbed and mitigated by the model's nonlinear activation functions and subsequent layer processing. When the quantization precision of weights and activations is consistent, frequent data type conversions during calculations can be avoided, reducing the additional errors introduced by data type conversions. For example, if weights and activations are both int8, they can be directly multiplied without truncation or rounding, thus ensuring computational accuracy.
[0042] For example, the output activation value of the current layer of neural network is obtained by weighted summing the input activation value of the previous layer of neural network and the weight parameter of the current layer. This process can be expressed by the following formula:
[0043]
[0044] Where W is the weight parameter (matrix) of the current layer, which contains the weights for each neuron. a is the input activation value (or input data) of the previous layer. b is the bias term, an additional parameter that allows the model to produce output even without input. z is the result of the weighted summation (the pre-output activation value of the current layer), which is passed as input to the activation function for nonlinear transformation.
[0045] Then, the output activation value of the current layer is further obtained by applying the activation function, which is as follows:
[0046]
[0047] in, is the output activation value of the current layer, z is the result of weighted summation (the pre-output activation value of the current layer), is the activation function.
[0048] In this embodiment, a first neural network model to be quantized is first determined. The first neural network model is a pre-trained high-precision neural network model that can achieve high-precision inference prediction result output. The weight parameters of each neural network layer of the first neural network model and the output activation values of each neural network layer are all represented by floating-point numbers with a first preset number of digits. However, due to the high precision of the first neural network model, the model computational complexity and model scale are large. The memory resources and computing resources of the embedded device are insufficient to support the first neural network model. Therefore, it is necessary to quantize the first neural network model. The weight parameters of the neural networks of multiple target layers of the first neural network model are converted from floating-point numbers with a first preset number of digits to integers with a second preset number of digits to obtain a second neural network model. The input activation values of the neural networks of the multiple target layers are also converted from floating-point numbers with a first preset number of digits to integers with a second preset number of digits. Thus, by converting some of the weight parameters and activation values of the neural network model to integers with a lower number of digits, the computing resource consumption during the model inference process is significantly reduced, thereby improving the operating efficiency of the model on the embedded device or mobile device. Furthermore, the weight parameters and input activation values of the neural network of the target layer are subjected to equal-precision quantization processing, so that the representation of the weight parameters and input activation values of the neural network of the target layer is kept completely consistent. Thus, although the accuracy of the output of the second neural network model after quantization is reduced, the prediction performance of the model is maintained within a reasonable range, avoiding the problem of non-convergence or garbled code caused by inconsistent quantization accuracy of the weight parameters and input activation values, and reducing the error rate of the output results of the second neural network model. Since the weight parameters and input activation values of the neural network of the target layer are subjected to equal-precision quantization processing, the errors caused by precision mismatch can be avoided, thereby achieving the technical effect of reducing the error rate of the output results of the second neural network model.
[0049] In one embodiment, Figure 3 As shown, step S210 converts the weight parameters of the neural networks of multiple target layers of the first neural network model from floating point numbers of a first preset number of digits to integers of a second preset number of digits. It includes: steps S300-S330:
[0050] Step S300: determining a first numerical range of a plurality of weight parameters of a neural network of a target layer to be quantized.
[0051] The neural network of the target layer to be quantized is one of the neural networks of multiple target layers. Therefore, the quantization granularity is the neural network of each target layer. The quantization granularity is the basic unit for dividing the range of quantization parameters. The weight parameters of each target layer neural network are grouped and quantized, which ensures both quantization accuracy and computational efficiency. In other words, each target layer neural network needs to execute steps S300-S320 separately to convert the weight parameters.
[0052] Specifically, before model quantization begins, the weight parameters of the target layer are first analyzed to determine the range of their numerical distribution. This is an important basis in the quantization process and is used to calculate subsequent quantization parameters.
[0053] The target neural network layer is the neural network layer selected for quantization. The weight parameter represents the connection strength between neurons in the model, which determines how information is transmitted between layers. The value range is the range between the minimum and maximum values of the weight parameter, which determines the integer range after quantization.
[0054] Step S310 : determining a range of values that can be represented by an integer of a second preset number of bits as a quantization range.
[0055] Specifically, according to the selected number of quantization bits (such as INT8), the range of integer values that can be represented is determined. The quantization range is the target range in the quantization process, and all floating-point weights will be mapped into this range.
[0056] The second preset number of bits is the number of bits of the quantized weight parameter, for example, an 8-bit integer. The quantization range is the minimum and maximum integer values that can be represented by the quantized weight parameter.
[0057] For example, for INT8, the quantization range is fixed to [-127, 127]. For other quantization bit numbers, such as INT4 or INT16, the corresponding quantization range needs to be determined based on the number of bits.
[0058] Step S320 , establishing a weight mapping relationship of the neural network of the target layer to be quantized according to the first numerical range and the quantization range of the plurality of weight parameters of the neural network of the target layer to be quantized.
[0059] Specifically, based on the numerical range of the target layer weight parameter and the numerical range after quantization, the scaling factor and the zero point are calculated, and a mapping relationship is established to map the floating-point weight to the quantized integer weight.
[0060] The weight mapping relationship is a formula or function for converting floating-point weights into quantized integer weights.
[0061] Step S330 , according to the weight mapping relationship, converting multiple weight parameters of the neural network of the target layer to be quantized from floating-point numbers of a first preset number of bits into integers of a second preset number of bits.
[0062] In this embodiment, by determining the first numerical range, it is ensured that quantization is performed based on the actual distribution of the weight parameters of the target layer, thereby avoiding information loss during the quantization process. During the entire quantization process, through precise weight parameter mapping, the prediction accuracy of the model on the embedded device is maintained within a reasonable range, avoiding a significant decrease in accuracy due to quantization. By accurately converting the target layer of the first neural network model into a second neural network model, the latter exhibits lower resource consumption, higher computing efficiency and stable prediction performance when running on an embedded device. This enables large models to be deployed efficiently and accurately in a resource-constrained environment, solving the dual challenges of running efficiency and accuracy on embedded devices after model quantization.
[0063] In one embodiment, Figure 4 As shown, step S320, based on the first numerical range and quantization range of multiple weight parameters of the neural network of the target layer to be quantized, establishes the weight mapping relationship of the neural network of the target layer to be quantized. It includes: steps S400-S420:
[0064] Step S400 , determining an upper limit value and a lower limit value of a first numerical range of a plurality of weight parameters of a neural network of a target layer to be quantized.
[0065] Specifically, we identify the numerical distribution range of the weight parameters in the neural network layer that needs to be quantized, that is, the maximum and minimum values. This step is the starting point of the quantization process. It helps us understand the original numerical characteristics of the weight parameters and provides a basis for subsequent quantization parameter calculations.
[0066] The first numerical range is the numerical distribution range of the model weight parameter before quantization, expressed as floating point numbers. The upper limit and lower limit are the maximum and minimum values in the numerical distribution range of the weight parameter.
[0067] Exemplarily, all weight parameters of the current target layer are traversed, and the maximum value (upper limit value) and the minimum value (lower limit value) of the weight parameters are statistically obtained.
[0068] Step S410 : determining a scaling factor according to the upper limit value, the lower limit value, and the quantization range.
[0069] The scaling factor is used to map the upper limit value and the lower limit value represented by the floating-point number of the first preset number of bits to the quantization range.
[0070] Specifically, the scaling factor is a key parameter used to map weight parameters from floating-point numbers to the quantized integer range. It ensures that the values of the weight parameters can be properly scaled to fit the quantized range while maintaining the proportional relationship of the original values as much as possible.
[0071] Among them, the scaling factor is used to adjust the scaling ratio of the weight parameter from the floating point number to the quantized integer range to ensure that information loss is minimized during the quantization process.
[0072] Exemplarily, the scaling factor is calculated based on the first value range (i.e., the upper and lower value limits) of the weight parameter and the quantization range. The formula may be as follows:
[0073]
[0074] in, is the scaling factor, is the upper limit of the quantization range, is the lower limit of the quantization range, is the upper limit value of the weight parameter, is the lower limit value of the weight parameter.
[0075] For example, the size of the quantization range for INT8 is 255 (ie, (-127) to 127), and the value range of the weight parameter is the difference between the upper limit value and the lower limit value.
[0076] Step S420 : determining a zero point according to the scaling factor, the lower limit value, and the lower limit value of the quantization range.
[0077] Among them, the scaling factor and the zero point are used to indicate the weight mapping relationship.
[0078] Specifically, the zero point is another key parameter used to adjust the weight parameter mapping during the quantization process. It ensures that the quantized weight value is based on the center of the quantization range, thereby further reducing the quantization error.
[0079] The zero point is used to map the zero point of the weight value to an integer value within the quantization range during the quantization process, so as to optimize the distribution of the weight value in the entire quantization range.
[0080] Exemplarily, the zero point is calculated based on the scaling factor, the lower limit value of the weight parameter, and the lower limit value of the quantization range. The calculation formula is:
[0081]
[0082] in, is zero point, is the rounding symbol, is the lower limit of the quantization range, is the lower limit value of the weight parameter, is the scaling factor.
[0083] In this embodiment, by accurately calculating the scaling factor and zero point, the weight parameters are ensured to be reasonably mapped within the quantization range, reducing information loss during the quantization process and maintaining the model's prediction accuracy. The quantized weight parameters occupy less storage space, speeding up the calculation process, reducing deployment costs and energy consumption on embedded devices, and enhancing applicability in resource-constrained environments. Through the weight mapping relationship, the quantized weight values are more evenly distributed within the integer domain, reducing computational instability caused by quantization and ensuring consistency in the model's prediction results across different hardware platforms. By adjusting the specific values of the scaling factor and zero point, the model's quantization process can maintain computational efficiency while maintaining the model's prediction capability within a reasonable range, avoiding significant accuracy degradation caused by quantization. By determining a first numerical range for the weight parameters, calculating the scaling factor and zero point within the quantization range, and constructing a precise weight mapping relationship, not only can the weight parameters of the neural network be efficiently converted from floating-point numbers to integers, but it also ensures that the quantized model maintains high prediction accuracy and computational stability when running on embedded devices, achieving the dual goals of model simplification and performance preservation.
[0084] In one embodiment, Figure 5 As shown, step S220, when the second neural network model receives the data to be inferred, converts the input activation values of the neural networks of multiple target layers from floating-point numbers of a first preset number of digits to integers of a second preset number of digits. It includes: steps S500-S540:
[0085] Step S500: Input a preset training data set into the first neural network model to determine the test activation value output by each layer of the neural network of the first neural network model.
[0086] Specifically, to quantize the input activations of a neural network, we first need to use a pre-set training dataset to evaluate the distribution of activation outputs at each layer of the model under normal operating conditions. This step is the calibration phase of the quantization process, ensuring that the quantization parameters reflect the model's actual runtime behavior. Because activations (i.e., the outputs of each neural network layer) are dynamic and dependent on the input data, it is necessary to calculate the dynamic range of activations. This involves inputting the model with the current training set and recording the output distribution of activations at each layer during the model's forward propagation. Histogram quantiles can then be used to determine the quantization range (e.g., between 10% and 90%).
[0087] The preset training dataset is a fixed dataset used for model evaluation and parameter calibration. It differs from the dataset used during model training and typically contains representative samples. The activation value is the output of a node in a neural network layer after receiving an input signal and processing it through an activation function. It is a key intermediate result in the model inference process.
[0088] Step S510: Determine a second numerical range of multiple test activation values based on the test activation values output by each layer of the neural network.
[0089] Specifically, based on the obtained test activation values, the distribution range of these test activation values, including the minimum and maximum values, is counted and determined. This step is the prerequisite for the subsequent calculation of the quantization parameters.
[0090] The second numerical range is the numerical distribution interval of the test activation value, which is used to calculate the quantization parameter to ensure that the quantized integer activation value can cover the range of the floating-point activation value.
[0091] Step S520 : determining a range of values that can be represented by an integer of a second preset number of bits as a quantization range.
[0092] Step S530: establishing an activation value mapping relationship of the first neural network model according to the second numerical range and the quantization range.
[0093] Specifically, based on the second numerical range and quantization range of the test activation value, a scaling factor and a zero point of the activation value are calculated to establish a mapping relationship between the floating-point activation value and the quantized integer activation value. The calculation method of the scaling factor and the zero point has been described in the above embodiment and will not be repeated here.
[0094] Step S540: Convert the input activation values of the neural networks of the multiple target layers from floating-point numbers of a first preset number of digits to integers of a second preset number of digits according to the activation value mapping relationship.
[0095] Specifically, the activation value mapping relationship is applied to convert the floating-point activation value of the neural network input of the target layer into an integer activation value.
[0096] Specifically, maintaining consistent quantization precision (the same bit width) for weights and activations ensures statistical consistency of errors when propagating between layers, preventing sudden errors. This allows the entire computational graph to run in a unified integer computational domain, reducing additional errors caused by data type conversion.
[0097] For example, the activation value quantization formula can be as follows:
[0098]
[0099] Where x is the floating activation value before quantization, Q(x) is the integer activation value after quantization, scale is the scaling factor of the activation value, and zero_point is the zero point of the activation value. Is the rounding symbol.
[0100] In this embodiment, by quantizing the input activation values, the memory usage and computing time of the model during the inference process are significantly reduced, because the quantized integer representation requires fewer resources than the floating-point representation. The established activation value mapping relationship ensures that the quantized activation value can be as close to the original value as possible, thereby maintaining the prediction accuracy of the model. Through carefully designed scaling factors and zero points, the information loss in the quantization process is minimized. Calibration using a preset training data set helps the quantization parameters better adapt to different types of input data, thereby improving the generalization ability and prediction stability of the model on unseen data. Through the above-mentioned quantization process, not only the resource consumption of the model on the embedded device is reduced, but also the accuracy of the model prediction and the matching of the activation value and weight parameter are maintained through the precise activation value mapping relationship.
[0101] In one embodiment, the method further comprises quantizing the bias term.
[0102] Specifically, the output activation value of each layer of the neural network is associated with a bias term, which is an additional parameter that allows the model to produce output even without input.
[0103] Specifically, the quantization range of the bias term is determined.
[0104] Similar to quantizing weights, the first step in quantizing biases is to determine a reasonable quantization range that encompasses all possible values of the bias while ensuring that the quantized bias still accurately reflects its role in the calculation.
[0105] For example, first, the minimum and maximum possible values of the bias term are determined by analyzing the model or using a training dataset, thereby determining the quantization range. The quantization range must match the actual dynamic range of the bias term to avoid information loss.
[0106] Calculate the quantization parameter of the bias term.
[0107] Specifically, based on the quantization range of the bias term determined in step 1, a scaling factor and a zero point required for quantizing the bias term are calculated to establish a mapping relationship from the original floating-point number to the quantized integer.
[0108] Quantization parameter to which the bias term is applied.
[0109] Specifically, the bias term is converted from a floating-point representation to an integer representation using the quantization parameter of the bias term, ensuring that it is computationally consistent with the quantized weight.
[0110] Quantized computation during model inference.
[0111] Specifically, during the model inference process, quantized weights and bias terms are used for calculations to ensure that the activation values are calculated at the specified integer precision, thereby improving inference efficiency and reducing resource consumption while maintaining effective convergence of the model output.
[0112] Dequantization calculation output.
[0113] Specifically, after the model inference process is completed, the activation values represented by integers are dequantized to floating-point numbers for accurate output. This process is necessary to ensure the accuracy and readability of the model output.
[0114] For example, the integer activation values ultimately obtained from model inference are dequantized. The dequantization formula is as follows: a = (Q(a) - zero_point) / scale, where a is a high-precision floating-point activation value, Q(a) is an integer activation value, zero_point is the activation zero point, and scale is the activation scaling factor. This step is typically performed at the model's output layer to ensure that the final output text or prediction results are highly accurate and understandable.
[0115] In this embodiment, the bias term is quantized so that the entire model can run under a lower precision data type, reducing the data type conversion and storage requirements during the calculation process, and significantly improving the inference efficiency and speed of the model on resource-constrained embedded systems. The quantized bias term takes up less memory space, reducing the storage requirements of the model, which is particularly important for embedded devices because they often have limited storage resources. The quantization accuracy of the bias term is consistent with that of the weight, avoiding the output instability problem caused by the mismatch of different precision parameters during the calculation process, ensuring the effective convergence of the model output, and improving the accuracy and stability of the overall prediction. Quantizing the bias term and weight allows the entire calculation graph to run in a unified integer calculation domain, reducing the additional errors caused by data type conversion, and improving the consistency of the calculation process and the reliability of the prediction results.
[0116] In one embodiment, Figure 6 As shown, the method further includes: steps S600-S610:
[0117] Step S600: Perform performance evaluation on each of the multi-layer neural networks of the first neural network model to determine the degree of influence of each layer of the multi-layer neural network on the output result of the first neural network model.
[0118] Specifically, we conduct a detailed performance analysis of each layer in the neural network model to assess its contribution to the overall model output. This analysis helps identify which layers have the most significant impact on the model output, providing guidance for subsequent quantitative strategies.
[0119] Performance evaluation refers to the assessment of a neural network layer's functional utility, information transfer capabilities, and influence on the final output. The degree of influence represents the contribution of a neural network layer to the overall model output, typically measured through layer sensitivity analysis, gradient analysis, or feature importance.
[0120] For example, methods such as neural network pruning, sensitivity analysis, or gradient attribution can be used to evaluate the impact of each layer.
[0121] Step S610: Determine multiple target layers of neural networks based on the degree of influence of each layer of neural network on the output result of the first neural network model.
[0122] Among them, the influence of the neural networks of multiple target layers on the output results of the first neural network model is less than a preset degree.
[0123] Specifically, based on the assessed impact, layers whose impact on the model output exceeds a preset threshold are selected as target layers. These layers will undergo further quantization to optimize their performance on embedded devices.
[0124] The target neural network layer is the one selected for parameter reduction during quantization. The preset threshold is used to distinguish the importance of neural network layers. Layers exceeding this threshold are considered to have a significant impact on the model output.
[0125] Exemplarily, a threshold is set, the influence of each layer of the neural network is compared with the threshold, and the layer with an influence less than the threshold is selected as the target layer.
[0126] The weight parameters and input activation values of the neural networks in the non-target layers of the multi-layer neural network of the first neural network model are represented by floating-point numbers with a first preset number of digits. Except for the target layer, the other layers (non-target layers) in the model continue to maintain the original floating-point number representation with the first preset number of digits to maintain the prediction accuracy and stability of the overall model.
[0127] In this embodiment, by quantizing only those layers (target layers) that have a minimal impact on the output, rather than quantizing all layers, we are able to significantly reduce the model's storage space and computing resource requirements on embedded devices while maintaining model prediction accuracy. This is because quantizing the target layer neural network reduces data storage and data type conversion overhead during the computation process. Non-target layers (i.e., layers that have a greater impact on the model output) retain their original floating-point precision representation, helping to ensure that the model's overall predictive ability is not significantly affected. This is because while the target layer of the model is quantized, the non-target layers maintain high precision, jointly supporting the model's predictive accuracy. By selectively quantizing the target layers in the model, rather than quantizing the entire model, we not only achieve efficient model deployment on embedded devices, but also maintain the model's prediction accuracy while enhancing the model's flexibility and adaptability.
[0128] In one embodiment, Figure 7 As shown, the method further includes: steps S700-S720:
[0129] Step S700: Receive the bytes output by the decoding layer of the second neural network model through the output filtering layer.
[0130] Wherein, the second neural network model includes an output filtering layer.
[0131] Specifically, the output of the second neural network model's decoding layer, a series of byte streams, is input into a specially designed output filtering layer, which checks and filters these bytes to ensure they conform to certain format standards.
[0132] The output filtering layer is a new layer added to the neural network architecture, used to check and filter the model output to ensure its correctness and readability. The decoding layer is the layer in the neural network model responsible for converting internal representations into user-understandable outputs, typically located at the end of the model. A token is the smallest unit of text processing in the model and can be a word, subword, character, or even a specific tag.
[0133] Step S710: Outputting the text result corresponding to the byte through the output filtering layer when the byte meets the preset format.
[0134] Specifically, when the bytes received by the output filter layer meet the preset format requirements, these bytes are converted into meaningful text results that are easy for users to understand and use. This is a key step in ensuring the quality of model output.
[0135] The default format is the format specification that the model output byte stream should comply with, which is used to determine whether the output is normal. The text result is the user-readable output information, such as natural language text.
[0136] For example, a check and conversion logic is implemented in a custom output filter layer. First, the received byte stream is checked to see if it conforms to the preset format, for example, verifying whether each byte falls within a valid encoding range. If the byte stream meets the preset format, the corresponding decoding function is used to convert the byte stream into a text result.
[0137] Step S720 : Buffering the bytes into a buffer area through the output filtering layer when the bytes do not meet the preset format.
[0138] Specifically, if the output filtering layer detects that the bytes do not conform to the preset format requirements, these bytes will not be immediately converted into text results, but will be temporarily stored in a cache area for possible subsequent merging processing or error correction.
[0139] For example, in the output filtering layer, if bytes are detected that do not conform to the preset format, these bytes are saved in a cache list or array to store the bytes that do not conform to the format and track their source and location to facilitate subsequent processing or merging.
[0140] In this embodiment, the format check of the output filter layer ensures the accuracy of the model's output byte stream before it is converted into text results, thereby improving the quality of the final output text, avoiding garbled or incomprehensible output, and enhancing the user experience. For bytes that do not conform to the format, they are cached instead of being discarded directly, providing a potential error recovery path. If the subsequent output can be merged with these cached bytes to form valid characters, the output anomalies caused by quantization can be repaired to a certain extent, improving the robustness and reliability of the model. The existence of the output filter layer avoids unnecessary data processing and conversion, reduces unnecessary consumption of computing resources, and especially when the model outputs a large amount of unformatted bytes, this mechanism can effectively save computing time and storage space.
[0141] In one embodiment, Figure 8 As shown, in step S720, after the output filtering layer caches the bytes into the buffer area when the bytes do not meet the preset format, the method further includes steps S800-S810:
[0142] Step S800: After the output filter layer caches the previous byte in the cache area, if a new byte is received, the new byte is cached in the cache area.
[0143] Specifically, when the output filter receives a byte that cannot be directly decoded into valid text, it is cached in a specific buffer. If a new byte is subsequently received, it is not analyzed and is directly defined as a byte that does not conform to the preset format. It is also cached in the same area for future decoding. This mechanism ensures that unformatted bytes are not immediately discarded, but have the opportunity to be combined with other bytes to form a decodable string.
[0144] The preset number is used to trigger a threshold for merging and decoding bytes in the buffer, that is, when the number of bytes in the buffer reaches this threshold, merging and decoding operations will be performed.
[0145] Step S810: When the number of bytes cached in the buffer reaches a preset number, multiple bytes cached in the buffer are merged and then decoded to obtain a corresponding text result.
[0146] Specifically, once the number of bytes accumulated in the buffer reaches a preset threshold, these bytes are merged into a byte sequence, and then an attempt is made to decode the entire sequence to obtain a complete, correctly formatted text result. This merge decoding strategy effectively handles situations where multiple bytes represent a single word due to quantization.
[0147] For example, due to issues with Chinese word segmentation, the characters (tokens) output by the quantized model may not correctly represent a single Chinese character. Two or even multiple tokens may be required to represent the same character. However, the model does not preprocess the tokens during output, causing them to be directly converted into an unusual, undefined character during output, resulting in garbled text. The main reason for Chinese word segmentation causing a single Chinese character to be split into multiple tokens is that the word segmentation algorithm (especially subword segmentation) processes Chinese characters. This is due to factors such as the frequency of occurrence of the characters during word segmentation training, the word segmentation design (byte-level or subword-level), and whether it is optimized for Chinese. This can result in the output being a word token. Because the output tokens are continuous, the next token must also be a subword token. Directly outputting a subword token can lead to garbled text. In most cases, we receive tokens in the form of strings, but these strings may be garbled (because a Chinese character is broken down into multiple sub-tokens, each of which is a one-byte character but invalid when viewed individually). Therefore, we need to collect consecutive bytes belonging to the same Chinese character and then combine them for decoding. First, based on the fixed pattern of the preset Chinese character encoding, we can determine that a Chinese character consists of three bytes, each of which is within a specific range. The conversion function then uses the collected tokens for decoding, and if the decoding is successful, it will output a complete Chinese character.
[0148] In this embodiment, by caching and merging bytes that cannot be decoded immediately, the chance of decoding the correct text is increased, and the frequency of garbled characters is reduced, even when model quantization causes byte representation anomalies. Caching bytes instead of immediately discarding them avoids possible information loss in the model output. At the same time, merging decoding reduces the waste of computing resources and improves overall decoding efficiency. The caching and merging strategy provides an error recovery method that can attempt to restore the model output by waiting for subsequent bytes when an anomaly is encountered during the output process. It also facilitates the management of these anomalies and simplifies the error handling process of the model output. It effectively improves the accuracy and readability of the model output and reduces the problem of garbled characters.
[0149] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0150] The embodiment of the present application also provides a model quantization device, Figure 9 : is a structural block diagram of a quantization device of a model according to an embodiment of the present application, the device comprising:
[0151] The first model determination module 901 is used to determine the first neural network model to be quantized, wherein the first neural network model includes a multi-layer neural network, and the weight parameters of each layer of the neural network in the multi-layer neural network and the output activation value of each layer of the neural network are both represented by floating-point numbers with a first preset number of bits.
[0152] The first quantization module 902 is used to convert the weight parameters of the neural networks of multiple target layers of the first neural network model from floating-point numbers of a first preset number of bits to integers of a second preset number of bits to obtain a second neural network model, wherein the second preset number of bits is smaller than the first preset number of bits.
[0153] The second quantization module 903 is configured to convert the input activation values of the neural networks of the multiple target layers from floating-point numbers of a first preset number of digits to integers of a second preset number of digits when the second neural network model receives the data to be inferred. The neural network of each target layer outputs a corresponding output activation value based on the weight parameter represented as an integer of the second preset number of digits and the input activation value represented as an integer of the second preset number of digits.
[0154] In an exemplary embodiment, the first quantization module 902 is also used to determine a first numerical range of multiple weight parameters of the neural network of the target layer to be quantized, wherein the neural network of the target layer to be quantized is one of the neural networks of multiple target layers. Determine the range of values that can be represented by integers of a second preset number of bits as the quantization range. Based on the first numerical range and quantization range of the multiple weight parameters of the neural network of the target layer to be quantized, establish a weight mapping relationship of the neural network of the target layer to be quantized. Based on the weight mapping relationship, the multiple weight parameters of the neural network of the target layer to be quantized are converted from floating-point numbers of the first preset number of bits to integers of the second preset number of bits.
[0155] In an exemplary embodiment, the first quantization module 902 is further configured to determine an upper limit value and a lower limit value of a first numerical range of multiple weight parameters of a neural network of a target layer to be quantized. A scaling factor is determined based on the upper limit value, the lower limit value, and the quantization range, wherein the scaling factor is used to map the upper limit value and the lower limit value represented by a floating-point number of a first preset number of bits to the quantization range. A zero point is determined based on the scaling factor, the lower limit value, and the lower limit value of the quantization range. The scaling factor and the zero point are used to indicate a weight mapping relationship.
[0156] In an exemplary embodiment, the second quantization module 903 is further used to input a preset training data set into the first neural network model to determine the test activation value output by each layer of the neural network of the first neural network model. Based on the test activation value output by each layer of the neural network, a second numerical range of multiple test activation values is determined. The range of numerical values that can be represented by integers of a second preset number of bits is determined as a quantization range. Based on the second numerical range and the quantization range, an activation value mapping relationship of the first neural network model is established. Based on the activation value mapping relationship, the input activation values of the neural networks of the multiple target layers are converted from floating-point numbers of the first preset number of bits to integers of the second preset number of bits.
[0157] In an exemplary embodiment, the apparatus further comprises:
[0158] The impact determination module is used to perform performance evaluation on the multi-layer neural networks of the first neural network model respectively, and determine the degree of influence of each layer of the multi-layer neural network on the output result of the first neural network model.
[0159] A target layer determination module is configured to determine a plurality of target layers of neural networks based on the degree of influence of each layer of neural networks on the output results of the first neural network model, wherein the degree of influence of the neural networks of the plurality of target layers on the output results of the first neural network model is less than a preset degree. The weight parameters and input activation values of the neural networks of the non-target layers in the multi-layer neural network of the first neural network model are represented by floating-point numbers with a first preset number of digits.
[0160] In an exemplary embodiment, the apparatus further comprises:
[0161] The receiving module is used to receive the bytes output by the decoding layer of the second neural network model through the output filtering layer.
[0162] The filtering module is used to output the text result corresponding to the byte through the output filtering layer when the byte meets the preset format.
[0163] The first cache module is configured to cache bytes into a cache area through an output filtering layer when the bytes do not meet a preset format.
[0164] In an exemplary embodiment, the apparatus further comprises:
[0165] The second cache module is configured to cache the previous byte into the cache area through the output filter layer, and if a new byte is received, cache the new byte into the cache area.
[0166] The decoding module is used to merge multiple bytes cached in the buffer area and then decode them to obtain corresponding text results when the number of bytes cached in the buffer area reaches a preset number.
[0167] For the description of the features in the embodiment corresponding to the quantization device of the model, reference can be made to the relevant description of the embodiment corresponding to the quantization method of the model, which will not be repeated here.
[0168] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in the embodiment of the quantization method of any of the above models.
[0169] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned model quantization method embodiments when run.
[0170] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0171] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in the quantization method embodiment of any of the above models are implemented.
[0172] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the quantization method embodiment of any of the above models are implemented.
[0173] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0174] The above is a detailed introduction to the quantification method, device, electronic device, computer-readable storage medium, and computer program product of a model provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core idea of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
Claims
1. A quantization method for a model, characterized in that: The method comprises: Determining a first neural network model to be quantized, wherein the first neural network model includes a multi-layer neural network, and weight parameters of each layer of the neural network in the multi-layer neural network and output activation values of each layer of the neural network are both represented by floating-point numbers with a first preset number of bits; Converting weight parameters of neural networks of multiple target layers of the first neural network model from floating-point numbers of the first preset number of digits to integers of a second preset number of digits to obtain a second neural network model, wherein the second preset number of digits is smaller than the first preset number of digits; When the second neural network model receives the data to be inferred, converting the input activation values of the neural networks of the multiple target layers from floating-point numbers of the first preset number of digits to integers of the second preset number of digits; The neural network of each target layer is used to output a corresponding output activation value according to the weight parameter and the input activation value; The weight parameter of the second neural network model is an integer and the second preset number of digits of the weight parameter of the second neural network model is smaller than the first preset number of digits of the weight parameter of the first neural network model, so that the consumption of computing resources by the second neural network during the inference process is less than the consumption of computing resources by the first neural network model during the inference process, and the second neural network model is deployed on an embedded device or a mobile device; The second neural network model includes an output filtering layer, and the method also includes: receiving the bytes output by the decoding layer of the second neural network model through the output filtering layer; outputting the text result corresponding to the byte through the output filtering layer when the byte meets the preset format; the text result is natural language text; caching the byte into a cache area through the output filtering layer when the byte does not meet the preset format; caching the new byte into the cache area if a new byte is received after the previous byte is cached into the cache area through the output filtering layer; when the number of bytes cached in the cache area reaches a preset number, merging multiple bytes cached in the cache area and then decoding to obtain the corresponding text result.
2. The quantification method of the model according to claim 1, characterized in that The converting weight parameters of the neural networks of the plurality of target layers of the first neural network model from floating point numbers of the first preset number of digits to integers of the second preset number of digits includes: Determining a first numerical range of multiple weight parameters of a neural network of a target layer to be quantized currently, wherein the neural network of the target layer to be quantized currently is one of the neural networks of the multiple target layers; Determining a range of numerical values that can be represented by the integer of the second preset number of bits as a quantization range; Establishing a weight mapping relationship of the neural network of the target layer to be quantized according to the first numerical range of multiple weight parameters of the neural network of the target layer to be quantized and the quantization range; According to the weight mapping relationship, multiple weight parameters of the neural network of the target layer to be quantized are converted from floating-point numbers of the first preset number of bits to integers of the second preset number of bits.
3. The quantification method of the model according to claim 2, characterized in that The step of establishing a weight mapping relationship of the neural network of the target layer to be quantized according to the first numerical range of the plurality of weight parameters of the neural network of the target layer to be quantized and the quantization range includes: Determine an upper limit value and a lower limit value of a first numerical range of multiple weight parameters of the neural network of the target layer to be quantized; Determining a scaling factor according to the upper limit value, the lower limit value, and the quantization range, wherein the scaling factor is used to map the upper limit value and the lower limit value represented by floating-point numbers with the first preset number of bits to the quantization range; determining a zero point according to the scaling factor, the lower limit value, and the lower limit value of the quantization range; The scaling factor and the zero point are used to indicate the weight mapping relationship.
4. The quantification method of the model according to claim 1, characterized in that The converting the input activation values of the neural networks of the multiple target layers from floating-point numbers of the first preset number of digits to integers of the second preset number of digits includes: Inputting a preset training data set into the first neural network model, and determining a test activation value output by each layer of the neural network of the first neural network model; Determining a second numerical range of a plurality of the test activation values according to the test activation values output by each layer of the neural network; Determining a range of numerical values that can be represented by the integer of the second preset number of bits as a quantization range; Establishing an activation value mapping relationship of the first neural network model according to the second numerical range and the quantization range; According to the activation value mapping relationship, the input activation values of the neural networks of the multiple target layers are converted from floating-point numbers of the first preset number of bits to integers of the second preset number of bits.
5. The quantification method of the model according to any one of claims 1 to 4, characterized in that: The method further comprises: Performing performance evaluation on each of the multi-layer neural networks of the first neural network model to determine the degree of influence of each layer of the neural network in the multi-layer neural network on the output result of the first neural network model; Determining multiple target layers of neural networks based on the degree of influence of each layer of the neural network on the output result of the first neural network model, wherein the degree of influence of the neural networks of the multiple target layers on the output result of the first neural network model is less than a preset degree; Among them, the weight parameters and input activation values of the neural network of the non-target layer in the multi-layer neural network of the first neural network model are represented by floating-point numbers of the first preset number of bits.
6. A quantization device for a model, characterized in that: include: A first model determination module is configured to determine a first neural network model to be quantized, wherein the first neural network model includes a multi-layer neural network, and weight parameters of each layer of the neural network in the multi-layer neural network and output activation values of each layer of the neural network are both represented by floating-point numbers with a first preset number of bits; A first quantization module is configured to convert weight parameters of neural networks of multiple target layers of the first neural network model from floating-point numbers of the first preset number of bits to integers of a second preset number of bits, to obtain a second neural network model, wherein the second preset number of bits is smaller than the first preset number of bits; A second quantization module is configured to convert the input activation values of the neural networks of the multiple target layers from floating-point numbers of the first preset number of bits to integers of the second preset number of bits when the second neural network model receives the data to be inferred; wherein the neural network of each target layer is configured to output a corresponding output activation value according to the weight parameter and the input activation value; wherein the weight parameter of the second neural network model is an integer and the second preset number of bits of the weight parameter of the second neural network model is smaller than the first preset number of bits of the weight parameter of the first neural network model, so that the consumption of computing resources by the second neural network during the inference process is less than the consumption of computing resources by the first neural network model during the inference process, and the second neural network model is configured to be deployed on an embedded device or a mobile device; The second neural network model includes an output filtering layer, and the device is also used to receive the bytes output by the decoding layer of the second neural network model through the output filtering layer; output the text result corresponding to the byte through the output filtering layer when the byte meets the preset format; the text result is natural language text; cache the byte into the cache area through the output filtering layer when the byte does not meet the preset format; after caching the previous byte into the cache area, if a new byte is received through the output filtering layer, cache the new byte into the cache area; when the number of bytes cached in the cache area reaches a preset number, merge multiple bytes cached in the cache area and then decode them to obtain the corresponding text result.
7. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method according to any one of claims 1 to 5 when executed by a processor.
Citation Information
Patent Citations
Neural network model quantification method, system and device and computer readable medium
CN114021691A
Neural network quantification method and system, electronic equipment and storage medium
CN120087423A