A piecewise linear quantization method and related device
By quantizing the floating-point network model to be quantized into an integer candidate quantization model, and quantizing it in segments into multiple target quantization models, and deploying it on the NPU side, the problem of model accuracy loss caused by low-bit quantization in the existing technology and the strong dependence of QTA on data is solved, and efficient model deployment and calculation are achieved.
Patent Information
- Application Number
- CN202211710556.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-12-29
AI Technical Summary
The prior art causes serious model accuracy losses during low-bit quantization, and QTA is highly dependent on data, making it difficult to use in actual industrial applications.
A piecewise linear quantization method is provided, which is deployed on the NPU side by quantizing the floating-point network model to be quantized into an integer candidate quantization model and quantizing it into at least two target quantization models.
It ensures the model accuracy of the network model deployed on the NPU side, reduces the consumption of NPU, and improves the operation and computing speed of the network model.
Smart Images

Figure CN115951859B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a piecewise linear quantization method and related devices. Background Art
[0002] The commonly used quantization methods include PTQ (post-training quantization) and QTA (quantization-aware training). Among them, PTQ (post-training quantization) will cause serious loss of model accuracy after quantization when quantizing at low bits (for example, 4 bits, etc.). Although QTA (quantization-aware training) can ensure the accuracy of the model after quantization, it is highly dependent on data and requires more training data. However, in actual industrial applications, training data is often difficult to obtain, which limits the use of QTA (quantization-aware training).
[0003] In order to solve the above problems, some researchers have proposed PWLQ (piecewise linear quantization), which truncates the floating point by finding one or more suitable ones in the floating point domain, and then quantizes different data bits in different intervals. However, due to the overlap of different quantization intervals, the data after piecewise quantization needs to be dequantized and returned to float32 for calculation. But in actual operation, the calculation of float32 by NPU (neural-network processing units) will cause huge consumption of NPU, and the transfer of float32 data on NPU will be affected by bandwidth limitations, resulting in slow operation and calculation of float32.
[0004] Therefore the prior art still needs to be improved and enhanced. Summary of the invention
[0005] The technical problem to be solved by the present application is to provide a piecewise linear quantization method and related devices in view of the deficiencies in the prior art.
[0006] In order to solve the above technical problem, the first aspect of the embodiment of the present application provides a piecewise linear quantization method, the method comprising:
[0007] Quantizing the network model to be quantized to obtain a candidate quantization model, wherein the data type of the network model to be quantized is a floating point type, and the data type of the candidate quantization model is an integer type;
[0008] The candidate quantization model is quantized into at least two target quantization models, and the at least two target quantization models are deployed on the NPU side, wherein the data type of each target quantization model is an integer.
[0009] The piecewise linear quantization method, wherein the data type of the network model to be quantized is float32, and the data type of the candidate quantization model is int8.
[0010] The piecewise linear quantization method, wherein the number of data bits of each target quantization model of the at least two target quantization models is smaller than the number of data bits of the candidate quantization model.
[0011] The piecewise linear quantization method, wherein the step of quantizing the candidate quantization model into at least two target quantization models specifically includes:
[0012] For a parameter to be quantized in a candidate quantization model, dividing the parameter to be quantized into at least two quantization intervals;
[0013] The number of data bits corresponding to each quantization interval is obtained, and the candidate quantization models are quantized according to the number of data bits corresponding to each quantization interval to obtain at least two target quantization models, wherein the at least two target quantization models correspond one-to-one to the at least two quantization intervals.
[0014] The piecewise linear quantization method, wherein for the parameter to be quantized in the candidate quantization model, dividing the parameter to be quantized into at least two quantization intervals specifically includes:
[0015] For a parameter to be quantized in a candidate quantization model, finding at least one breakpoint corresponding to the parameter to be quantized;
[0016] The parameter to be quantized is divided into at least two quantization intervals based on the at least one breakpoint.
[0017] The piecewise linear quantization method, wherein after the at least two target quantization models are deployed on the NPU end, the method includes:
[0018] Dequantize each target quantization model through the NPU to obtain a candidate quantization model;
[0019] Model inference is performed based on the candidate quantization model by the NPU to obtain an inference result.
[0020] The piecewise linear quantization method, wherein the computing unit for performing inverse quantization in the NPU end is stored in a memory relocation instruction, so that when data is imported into the buffer based on the memory relocation instruction, each target quantization model is inversely quantized to obtain a candidate quantization model.
[0021] A second aspect of an embodiment of the present application provides a piecewise linear quantization system, the system comprising:
[0022] A first quantization module is used to quantize the network model to be quantized to obtain a candidate quantization model, wherein the data type of the network model to be quantized is a floating point type, and the data type of the candidate quantization model is an integer type;
[0023] The second quantization module is used to quantize the candidate quantization model into at least two target quantization models, wherein the data type of each target quantization model is an integer.
[0024] A deployment module is used to deploy the at least two target quantization models on the NPU side.
[0025] A third aspect of an embodiment of the present application provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in any of the above-described piecewise linear quantization methods.
[0026] A fourth aspect of the embodiments of the present application provides a terminal device, comprising: a processor, a memory, and a communication bus; the memory stores a computer-readable program that can be executed by the processor;
[0027] The communication bus realizes the connection and communication between the processor and the memory;
[0028] When the processor executes the computer-readable program, the steps in any of the above-described piecewise linear quantization methods are implemented.
[0029] Beneficial effect: Compared with the prior art, the present application provides a piecewise linear quantization method and related devices, the method comprising quantizing a network model to be quantized to obtain a candidate quantization model; quantizing the candidate quantization model into at least two target quantization models, and deploying the at least two target quantization models on the NPU side. The present application first quantizes the floating-point network model to be quantized into an integer candidate quantization model, and then quantizes the candidate quantization model into multiple target quantization models through piecewise quantization. On the one hand, this can ensure the model accuracy of the network model deployed on the NPU side, and on the other hand, the NPU does not need to perform floating-point calculations, thereby reducing the consumption of the NPU side. On the other hand, the integer network model obtained by inverse quantization will not be limited by the NPU bandwidth, thereby improving the operation and calculation speed of the deployed network model. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without inventive work.
[0031] Figure 1 This is a flowchart of the piecewise linear quantization method provided in this application.
[0032] Figure 2 This is an example diagram of the piecewise linear quantization method provided in this application.
[0033] Figure 3 An example diagram of the inference process after deploying at least two target quantization models on the NPU side.
[0034] Figure 4 This is a schematic diagram of the structure of the piecewise linear quantization system provided in this application.
[0035] Figure 5 This is a schematic diagram of the structure of the terminal device provided in this application. DETAILED DESCRIPTION
[0036] The present application provides a piecewise linear quantization method and related devices. To make the purpose, technical solution and effect of the present application clearer and more specific, the present application is further described in detail with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0037] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.
[0038] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with those in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.
[0039] It should be understood that the sequence numbers and sizes of the steps in this embodiment do not mean the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.
[0040] After research, it was found that the commonly used quantization methods include PTQ (post-training quantization) and QTA (quantization-aware training). Among them, PTQ (post-training quantization) will cause serious loss of model accuracy after quantization when quantizing at low bits (for example, 4 bits, etc.). Although QTA (quantization-aware training) can ensure the accuracy of the model after quantization, it is highly dependent on data and requires more training data. However, in actual industrial applications, training data is often difficult to obtain, which limits the use of QTA (quantization-aware training).
[0041] In order to solve the above problems, some researchers have proposed PWLQ (piecewise linear quantization), which truncates the floating point by finding one or more suitable ones in the floating point domain, and then quantizes different data bits in different intervals. However, due to the overlap of different quantization intervals, it is necessary to dequantize the piecewise quantized data and calculate the floating point data. But in actual operation, the NPU (neural-network processing units) has a high computational cost for floating point calculations, making the quantization model using PWLQ quantization unusable on the NPU side.
[0042] In order to solve the above problems, in an embodiment of the present application, the network model to be quantized is quantized to obtain a candidate quantization model; the candidate quantization model is quantized into at least two target quantization models, and the at least two target quantization models are deployed on the NPU side. The present application first quantizes the floating-point network model to be quantized into an integer candidate quantization model, and then quantizes the candidate quantization model into multiple target quantization models through segmented quantization. On the one hand, this can ensure the model accuracy of the network model deployed on the NPU side, and on the other hand, the NPU does not need to perform floating-point calculations, thereby reducing the consumption of the NPU side. On the other hand, the integer network model obtained by inverse quantization will not be limited by the NPU bandwidth, thereby improving the operation and calculation speed of the deployed network model.
[0043] The application content is further explained below through the description of embodiments in conjunction with the accompanying drawings.
[0044] This embodiment provides a piecewise linear quantization method, such as Figure 1 As shown, the method includes:
[0045] S10. Quantize the network model to be quantized to obtain a candidate quantization model.
[0046] Specifically, the network model to be quantized is a floating-point network model, that is, the data type of the network model to be quantized is a floating-point type, and the candidate quantization model is an integer network model, that is, the data type of the candidate quantization model is an integer type. Among them, the quantization of the network model to be quantized can adopt a PTQ (post-training quantization) type quantization algorithm, or a QTA (quantization-aware training) type quantization algorithm to quantize the network model to be quantized. For example, the KL divergence quantization method, the percentile quantization method, the ACIQ quantization method, etc. are adopted.
[0047] Furthermore, when the network model to be quantized is quantized, the model weights and activation layer output items of the network model to be quantized are represented by floating-point model data, for example, Figure 2 As shown, the data type of the model weights in the network model to be quantized is float32, so quantizing the network model to be quantized is to quantize the model weights in the model to be quantized, so as to quantize the data type of the model weights to int8.
[0048] S20. quantize the candidate quantization model into at least two target quantization models, and deploy the at least two target quantization models on the NPU end.
[0049] Specifically, each target quantization model is obtained by quantizing part of the data of the parameters to be quantized in the candidate quantization model, wherein the data type of each target quantization model is an integer, and the number of bits of data of each target quantization model is smaller than that of the candidate quantization model. Figure 2 The at least two target quantization models include a target quantization model A and a target quantization model B, the data type of the target quantization model A is int4, the data type of the target quantization model B is int3, etc. In addition, in practical applications, the at least two target quantization models may have some target quantization models with the same data type and some target quantization models with different data types; the data types of each target quantization model may be different, or the data types of each target quantization model may be the same.
[0050] In one implementation, quantizing the candidate quantization model into at least two target quantization models specifically includes:
[0051] For a parameter to be quantized in a candidate quantization model, dividing the parameter to be quantized into at least two quantization intervals;
[0052] The number of data bits corresponding to each quantization interval is obtained, and the candidate quantization models are quantized according to the number of data bits corresponding to each quantization interval to obtain at least two target quantization models.
[0053] Specifically, the parameter to be quantized may be a model weight in a candidate quantization model, etc. Dividing the parameter to be quantized into at least two quantization intervals means truncating the parameter to be quantized in its data domain to form at least two data segments, and each data end is a quantization interval.
[0054] After obtaining at least two quantization intervals, the number of data bits is configured for each quantization interval, wherein the number of data bits configured for the quantization interval is the number of data bits after quantization corresponding to the data segment of the parameter to be quantized belonging to the quantization interval. In other words, the data type of the data segment of the parameter to be quantized belonging to the quantization interval is quantized to the number of data bits corresponding to the quantization interval. It can be seen that at least two quantization models correspond to at least two quantization intervals one by one, and each target quantization model is obtained by quantizing the quantization interval corresponding to it. For example, the parameter to be quantized is divided into two quantization intervals, respectively denoted as Tail and Middle, wherein Tail corresponds to int4 and Middle corresponds to int3, then Tail is quantized to int4 to obtain a target quantization model, and Middle is quantized to int3 to obtain a target cool model.
[0055] In one implementation, for the parameter to be quantized in the candidate quantization model, dividing the parameter to be quantized into at least two quantization intervals specifically includes:
[0056] For a parameter to be quantized in a candidate quantization model, finding at least one breakpoint corresponding to the parameter to be quantized;
[0057] The parameter to be quantized is divided into at least two quantization intervals based on the at least one breakpoint.
[0058] Specifically, finding at least one breakpoint corresponding to the parameter to be quantized refers to finding a breakpoint in the data domain to which the data to be quantized belongs. For example, if the data type of the candidate quantization model is int8, then at least one breakpoint is found on the int8 value domain. The breakpoint finding method can adopt the breakpoint finding method in PWLQ, which will not be described here.
[0059] Further, after the breakpoints are found, the parameter to be quantized is divided based on the breakpoints to obtain several candidate sub-regions, and then the symmetric candidate sub-regions in the several candidate sub-regions are merged to obtain at least two quantization intervals. In one implementation, since the model weights in the candidate quantization model conform to the bell-shaped curve, after obtaining and finding n breakpoints, the parameter to be quantized can be divided into 2n+1 segments based on the breakpoints to be represented as symmetrical n+1 segments, that is, the number of quantization intervals is equal to the number of breakpoints 1. For example, if the number of breakpoints is 1, then the number of quantization intervals is 2. This is because the model weights in the candidate quantization model conform to the bell-shaped curve, so after finding n breakpoints, the parameter to be quantized can be divided into 2n+1 segments based on the breakpoints to be represented as symmetrical n+1 segments. For example, there are 1 breakpoint, and the candidate quantization model is divided into 2*1+1=3 segments, which are recorded as [-∞, -bkp], [-bkp, bkp] and [bkp, -∞]. The first and third segments are symmetrical, so two symmetrical intervals are determined, which are [±∞, ±bkp] and [-bkp, bkp]. There are 2 breakpoints, and the candidate quantization model is divided into 2*2+1=5 segments. Since the first end is symmetrical with the fifth segment, and the second segment is symmetrical with the fourth segment, three symmetrical intervals are determined, which are [±∞, ±bkp1], [±bkp1, ±bkp2], and [-bkp2, bkp2].
[0060] In addition, since each quantization interval corresponds to two quantization parameters, 5 is scale and zero point, respectively, the candidate quantization
[0061] After the model is segmented and quantized into at least two target quantization models, 2(n+1) parameters are obtained, where n is the number of breakpoints.
[0062] In one implementation, the at least two target quantization models are deployed on the NPU end.
[0063] Thereafter, the method comprises:
[0064] 0 Dequantize each target quantization model through the NPU to obtain a candidate quantization model;
[0065] Model inference is performed based on the candidate quantization model by the NPU to obtain an inference result.
[0066] Specifically, the NPU side deploys several target quantization models, so that the NPU side can obtain the quantization parameters corresponding to the target quantization models, that is, the NPU side can obtain at least two sets of quantization parameters, at least
[0067] The two sets of quantization parameters correspond to at least two sets of target quantization models one by one. In addition, when performing model inference on the NPU side, the quantization parameters corresponding to each target quantization model can be obtained, and then the quantization parameters can be used to calculate the target quantization model.
[0068] The candidate quantization model is obtained by dequantizing the number of quantization models, so that the candidate quantization model can be used for quantization. In this way, at least two low-bit target quantization models are used on the NPU to store the candidate quantization model, and when inference is performed, the candidate quantization model can be obtained by dequantization and the candidate quantization model can be obtained by the candidate quantization model.
[0069] The quantization model is selected for inference, which ensures the inference performance of the NPU. At the same time, the candidate quantization model obtained by dequantization is an integer type, so that the NPU does not need to perform floating-point calculations, which can avoid NPU
[0070] The bandwidth limitation of the end can be overcome, thereby improving the computing speed of the NPU end.
[0071] For example, if the candidate quantization model is the int8 model, and at least two target quantization models include the int4 target quantization model A and the int3 target quantization model B, then the NPU side deploys the target quantization model
[0072] Type A and int3 target quantization model B, the NPU side is inferring, such as Figure 3 As shown, first load the quantization model parameters of the target quantization model A and the target quantization model B, and then quantize the target
[0073] Model A and target quantization model B are dequantized to obtain a candidate quantization model. Finally, model inference is performed through the candidate quantization model to obtain an inference result.
[0074] In one implementation, the computing unit for performing inverse quantization in the NPU is stored in a memory relocation instruction, so that when data is imported into the buffer based on the memory relocation instruction, each target quantization model is inverse quantized to obtain a candidate quantization model.
[0075] Specifically, the memory relocation instruction is used to control the import of external data into the buffer of the NPU side, wherein the memory relocation instruction stores a computing unit, so that when the external data is imported into the buffer, the target quantization model can be synchronously dequantized back to the candidate quantization model, so that it can be determined that the dequantization process can be calculated during the data transfer process, and the calculation and transfer of the target quantization model can be completed without adding any additional consumption to the NPU side.
[0076] In summary, the present embodiment provides a piecewise linear quantization method and related devices, the method comprising quantizing a network model to be quantized to obtain a candidate quantization model; quantizing the candidate quantization model into at least two target quantization models, and deploying the at least two target quantization models on the NPU side. The present application first quantizes the floating-point network model to be quantized into an integer candidate quantization model, and then quantizes the candidate quantization model into multiple target quantization models through piecewise quantization. In this way, on the one hand, the model accuracy of the network model deployed on the NPU side can be guaranteed, and the NPU does not need to perform floating-point calculations, thereby reducing the consumption of the NPU side. On the other hand, the integer network model obtained by inverse quantization will not be limited by the NPU bandwidth, thereby improving the operation and calculation speed of the deployed network model.
[0077] Based on the above piecewise linear quantization method, this embodiment provides a piecewise linear quantization system, such as Figure 4 As shown, the system comprises:
[0078] The first quantization module 100 is used to quantize the network model to be quantized to obtain a candidate quantization model, wherein the data type of the network model to be quantized is a floating point type, and the data type of the candidate quantization model is an integer type;
[0079] The second quantization module 200 is used to quantize the candidate quantization model into at least two target quantization models, wherein the data type of each target quantization model is an integer.
[0080] The deployment module 300 is used to deploy the at least two target quantization models on the NPU side.
[0081] Based on the above-mentioned piecewise linear quantization method, this embodiment provides a computer-readable storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps in the piecewise linear quantization method as described in the above-mentioned embodiment.
[0082] Based on the above piecewise linear quantization method, the present application also provides a terminal device, such as Figure 5As shown, it includes at least one processor (processor) 20; display screen 21; and memory (memory) 22, and may also include a communications interface (Communications Interface) 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22 and the communication interface 23 can communicate with each other through the bus 24. The display screen 21 is configured to display a preset user guide interface in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logic instructions in the memory 22 to execute the method in the above embodiment.
[0083] In addition, the logic instructions in the memory 22 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0084] The memory 22 is a computer-readable storage medium that can be configured to store software programs, computer executable programs, such as program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions or modules stored in the memory 22, that is, implementing the methods in the above embodiments.
[0085] The memory 22 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, a variety of media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, may also be a transient storage medium.
[0086] In addition, the specific process of loading and executing the multiple instruction processors in the above-mentioned storage medium and the terminal device has been described in detail in the above-mentioned method, and will not be described one by one here.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A piecewise linear quantization method, characterized in that: The method comprises: Quantize the network model to be quantized to obtain a candidate quantization model, wherein the data type of the network model to be quantized is float32, and the data type of the candidate quantization model is int8; Quantize the candidate quantization model into at least two target quantization models, and deploy the at least two target quantization models on the NPU end, wherein the data type of each target quantization model is an integer type, and the number of data bits of each target quantization model in the at least two target quantization models is smaller than the number of data bits of the candidate quantization model; Dequantize each target quantization model through the NPU to obtain a dequantized candidate quantization model; Model inference is performed on the NPU side based on the candidate quantization model of the dequantization to obtain an inference result, wherein the computing unit for performing dequantization in the NPU side is stored in a memory relocation instruction, so that when data is imported into the buffer based on the memory relocation instruction, each target quantization model is dequantized to obtain a candidate quantization model of the dequantization.
2. The piecewise linear quantization method according to claim 1, characterized in that: The step of quantizing the candidate quantization model into at least two target quantization models specifically includes: For a parameter to be quantized in a candidate quantization model, dividing the parameter to be quantized into at least two quantization intervals; The number of data bits corresponding to each quantization interval is obtained, and the candidate quantization models are quantized according to the number of data bits corresponding to each quantization interval to obtain at least two target quantization models, wherein the at least two target quantization models correspond one-to-one to the at least two quantization intervals.
3. The piecewise linear quantization method according to claim 2, characterized in that: For the parameter to be quantized in the candidate quantization model, dividing the parameter to be quantized into at least two quantization intervals specifically includes: For a parameter to be quantized in a candidate quantization model, finding at least one breakpoint corresponding to the parameter to be quantized; The parameter to be quantized is divided into at least two quantization intervals based on the at least one breakpoint.
4. A piecewise linear quantization system, characterized in that: The system comprises: A first quantization module is used to quantize the network model to be quantized to obtain a candidate quantization model, wherein the data type of the network model to be quantized is float32, and the data type of the candidate quantization model is int8; A second quantization module, configured to quantize the candidate quantization model into at least two target quantization models, wherein the data type of each target quantization model is an integer type, and the number of data bits of each target quantization model in the at least two target quantization models is smaller than the number of data bits of the candidate quantization model; A deployment module is used to deploy the at least two target quantization models on the NPU end, and to dequantize each target quantization model through the NPU end to obtain a dequantized candidate quantization model; and to perform model inference based on the dequantized candidate quantization model through the NPU end to obtain an inference result, wherein the computing unit for performing dequantization in the NPU end is stored in a memory relocation instruction, so that when data is imported into the buffer based on the memory relocation instruction, each target quantization model is dequantized to obtain a dequantized candidate quantization model.
5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the piecewise linear quantization method as described in any one of claims 1-3.
6. A terminal device, characterized in that: include: Processor, memory and communication bus; The memory stores a computer-readable program executable by the processor; The communication bus realizes the connection and communication between the processor and the memory; When the processor executes the computer-readable program, the steps in the piecewise linear quantization method according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Multi-base pair quantization method and device for depth neural network
CN109376854A
4-bit quantization method and system of neural network
CN111882058A