Method for accelerating processing of deep learning models

By monitoring quantization acceleration instructions in the deep learning model, acquiring real-world data for dynamic quantization processing, generating target quantization parameters, and accelerating processing, the problem of low processing efficiency caused by differences in calibration data is solved, achieving more efficient quantization acceleration.

CN115688905BActive Publication Date: 2026-03-27ALIBABA (CHINA) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-09
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing deep learning models suffer from low processing efficiency in quantization calculations due to discrepancies between calibration data and actual data, and there is a lack of effective solutions.

Method used

By monitoring quantization acceleration commands, a pre-trained neural network model is triggered to acquire real-world data for dynamic quantization processing, generating target quantization parameters, and then using an accelerator processor for accelerated processing to achieve dynamic quantization.

Benefits of technology

While ensuring quantization accuracy, it improves the processing speed of deep learning models, solves the problem of low processing efficiency, and reduces additional overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115688905B_ABST
    Figure CN115688905B_ABST
Patent Text Reader

Abstract

The application discloses a method for accelerating a deep learning model. The method comprises the following steps: monitoring a quantization acceleration instruction, triggering a deep learning model to be executed for acceleration, wherein the deep learning model is a pre-trained neural network model; acquiring target data collected in a real environment; performing dynamic quantization processing on the deep learning model based on the target data to obtain target quantization parameters of the deep learning model; and performing acceleration processing on the deep learning model based on the target quantization parameters by using an acceleration processor to obtain a target result of the target data. The application solves the technical problem of low processing efficiency of the deep learning model in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of deep learning model processing, in particular, to a method for accelerating deep learning model processing. BACKGROUND

[0002] At present, more and more system services depend on artificial intelligence systems based on deep learning. The inference hardware of the current deep learning model usually uses low-precision quantization calculation format for acceleration, and most of the quantization inference engines generate quantization tables using calibration data. Since there is a difference between the calibration data and the data generated in the actual process, it will cause difficulties in the specific use process, thereby causing the processing efficiency of the deep learning model to be low.

[0003] In view of the above problems, no effective solution has been proposed so far. SUMMARY

[0004] The embodiments of the present application provide a method for accelerating deep learning model processing, which at least solves the technical problem of low processing efficiency of the deep learning model in the related art.

[0005] According to an aspect of the embodiments of the present application, a method for accelerating deep learning model processing is provided, comprising: monitoring a quantization acceleration instruction to trigger a deep learning model to be executed for acceleration processing, wherein the deep learning model is a pre-trained neural network model; acquiring target data collected in a real environment; performing dynamic quantization processing on the deep learning model based on the target data to obtain target quantization parameters of the deep learning model; and performing acceleration processing on the deep learning model based on the target quantization parameters through an acceleration processor to obtain a target result of the target data.

[0006] According to another aspect of the embodiments of the present application, a data processing apparatus is also provided, comprising: a triggering module configured to monitor a quantization acceleration instruction to trigger a deep learning model to be executed for acceleration processing, wherein the deep learning model is a pre-trained neural network model; an acquisition module configured to acquire target data collected in a real environment; a dynamic quantization module configured to perform dynamic quantization processing on the deep learning model based on the target data to obtain target quantization parameters of the deep learning model; and an acceleration network layer configured to perform acceleration processing on the deep learning model based on the target quantization parameters through an acceleration processor to obtain a target result of the target data.

[0007] According to another aspect of the embodiments of the present application, a computer readable storage medium is also provided, which includes a stored program, wherein when the program is running, the computer readable storage medium controls the device where the computer readable storage medium is located to execute the method of any one of the above embodiments.

[0008] According to another aspect of the embodiments of the present application, an electronic device is also provided, including a memory storing an executable program, and a processor configured to execute the program, wherein the program performs the method of any of the above embodiments when executed.

[0009] In the embodiments of the present application, the quantization acceleration instruction can be monitored to trigger a deep learning model to be executed for acceleration processing, wherein the deep learning model is a pre-trained neural network model; target data collected in a real environment is obtained; the deep learning model is dynamically quantized based on the target data to obtain target quantization parameters of the deep learning model; and the deep learning model is processed by an acceleration processor based on the target quantization parameters to obtain a target result of the target data, thereby achieving the purpose of improving the processing efficiency of the deep learning model. It is easy to note that the dynamic quantization of the deep learning model based on the target data can generate a quantization table online, so that the quantization process does not generate additional overhead, and the speed can be improved while ensuring the quantization accuracy, thereby solving the technical problem of low processing efficiency of the deep learning model in the related art.

[0010] It is easy to note that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation on the present application. BRIEF DESCRIPTION OF DRAWINGS

[0011] The drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The schematic embodiments of the present application and the description thereof are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:

[0012] Figure 1 is a schematic diagram of end-to-end inference time according to an embodiment of the present application;

[0013] Figure 2 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing the method of accelerating the deep learning model according to an embodiment of the present application;

[0014] Figure 3 is a flowchart of the method of accelerating the deep learning model according to Embodiment 1 of the present application;

[0015] Figure 4 is a flowchart of a dynamic quantization process according to an embodiment of the present application;

[0016] Figure 5 is a schematic diagram of symmetric quantization and asymmetric quantization according to an embodiment of the present application;

[0017] Figure 6 is a matrix multiplication schematic diagram according to an embodiment of the present application;

[0018] Figure 7 is a fusion quantization operator according to an embodiment of the application;

[0019] Figure 8 is a performance test comparison diagram according to an embodiment of the application;

[0020] Figure 9 is a flow chart of another dynamic quantization process according to an embodiment of the application;

[0021] Figure 10 is a block diagram of a deep learning model acceleration process according to an embodiment of the application;

[0022] Figure 11 is a flow chart of a deep learning model acceleration process according to an embodiment of the application;

[0023] Figure 12 is a performance test comparison diagram according to an embodiment of the application;

[0024] Figure 13 is a data processing device according to an embodiment of the application;

[0025] Figure 14 is a structure block diagram of a computer terminal according to an embodiment of the application. DETAILED DESCRIPTION

[0026] In order to make the personnel in the technical field better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] First, some of the nouns or terms that appear in the description of the embodiments of the present application are explained as follows:

[0029] Full precision: usually refers to the calculation of float type (float32);

[0030] Quantitative calculation: usually refers to the calculation format of scalar, for example, integer type (int8) refers to 8-bit integer (integer) format;

[0031] Heterogeneous inference engine (HIE): inference service is applied to different hardware, which needs the support of multiple types of heterogeneous computing engines;

[0032] Inference acceleration engine (TensorRT, referred to as TRT): inference acceleration engine can provide low latency and high throughput inference deployment for deep learning applications;

[0033] Deep learning acceleration library (CUDADeep Neural Network, referred to as CUDNN): deep learning acceleration library can accelerate widely used deep learning framework;

[0034] Heterogeneous computing intelligence (HCI): heterogeneous computing can improve computing power and performance, reduce power consumption and cost;

[0035] Matrix multiplication (GEMM): usually refers to f32, which can also include int8 matrix multiplication calculation;

[0036] Quantized GEMM calculation (QuantGEMM): for example, int8 matrix multiplication calculation;

[0037] Each inference quantization (PerTensor): the entire Tensor uses the same quantization parameter;

[0038] Each channel quantization (PerChannel): the entire Tensor is divided into multiple quantization parameters in the channel dimension;

[0039] Open neural network exchange (onnx): a general representation format of deep learning model;

[0040] Inference framework (OnnxRuntime, referred to as ORT): users can conveniently run one of the onnx models with the inference framework;

[0041] Kernel function (Kernel): here mainly refers to (High Performance Computing, referred to as HPC) calculation;

[0042] Quantization: the process of scaling the input from FP32 representation range to int8 or uint8 representation range, and save some quantization parameters.

[0043] Dequantization: the operation of converting the int8 or uint8 output to FP32 representation range by the quantization parameters calculated before.

[0044] Operator (OP): here refers to the operator in the deep learning model, such as convolution layer (Conv) is an operator, activation function (Relu) is an operator, and maximum pooling (MaxPooling) is an operator.

[0045] Generally, quantization is divided into dynamic quantization and static quantization, as follows:

[0046] Static quantization: a set of verification sets is used to estimate the quantization parameters of each layer by statistical method.

[0047] Dynamic quantization: according to the current input, the quantization parameters of the input are calculated in real time, so as to achieve more accurate simulation results.

[0048] In terms of accuracy, the accuracy of dynamic quantization is higher than that of static quantization.

[0049] The accuracy of static quantization is greatly affected by the verification set and verification algorithm, so only "appropriate" verification set + "appropriate" verification algorithm can better guarantee the accuracy of the final model. Although dynamic quantization uses the most direct way to quantize, it can dynamically calculate the quantization parameters of each input, and HIE uses PerChannel quantization granularity to ensure that the quantization of inputs is independent of each other, so the final model quantization is more stable.

[0050] In terms of performance, the performance of static quantization is higher than that of dynamic quantization.

[0051] Although the performance of static quantization is higher than that of dynamic quantization, the specific performance has a great relationship with the specific implementation process. In HIE, the extra loss brought by dynamic quantization can be controlled within 5% by implementing efficient quantization and dequantization Kernel, hiding the quantization process and reasonable quantization calculation process.

[0052] Figure 1is a schematic diagram of end-to-end inference time according to an embodiment of the present application, is a real Transformer, and the additional loss of statistical dynamic quantization only accounts for 2.8% of the total end-to-end time, wherein Dynamic Quantize (Fused) is to hide the quantization process in other processes, such as the normalization layer (LayerNorm), in HIE. For integer INT8, other processes account for 92, the dynamic quantization fusion process accounts for 15, the dynamic quantization process accounts for 4.6, and the matrix multiplication process accounts for 49; for floating-point number FP32, other processes account for 107, and the matrix multiplication process accounts for 155. As can be seen, the processing time of floating-point number FP32 is longer.

[0053] In terms of ease of use, the ease of use of dynamic quantization is greater than that of static quantization.

[0054] Static quantization: Although the static quantization schemes of various inference frameworks are similar, the tool chains of each party are different. Static quantization needs to prepare the corresponding verification set according to the settings of each inference framework. This engineering landing is very troublesome. And the selection of the verification set is also very particular. The verification set needs to be "representative", otherwise the accuracy of the final actual scene will be greatly affected.

[0055] Dynamic quantization: HIE dynamic quantization does not need calibration data to generate a quantization table. It only needs to call an executable program to automatically convert the input data of the model into low-precision quantization calculation mode. This process can be done without providing any verification set.

[0056] The current conventional quantization is symmetric quantization + PerTensor granularity. This scheme is simple to implement, so it is used by a large number of frameworks. However, this scheme has some problems:

[0057] Problem 1: Symmetric quantization in some scenarios can cause 8Bit data to only express 7Bit range. For example, after the activation function operation, the negative half-axis is all 0. Generally, UINT8 can be used to solve the problem. However, if the activation function is a non-elementary function form of activation function (Gaussian, referred to as GELU) and exponential linear unit (Exponential, referred to as ELU), obviously, UINT8 is not a general solution.

[0058] Problem 2: The PerTensor method is easy to cause a part of data to be abnormal, resulting in global data abnormality. Taking matrix multiplication as an example, once the data of an intermediate layer is abnormal, the global quantization parameter will be abnormal, eventually causing the end-to-end accuracy to decrease.

[0059] Currently, deep learning inference hardware usually uses low precision, especially quantized (int) calculation format for acceleration, and most of the quantized inference engines use calibration data to generate quantization tables, which causes difficulties in use and differences between users in calibration data and functional data, making it difficult for quantized acceleration technology to be truly used on a large scale. The present application is mainly used for deep learning model inference acceleration, can be used in many deep learning system services, improves the efficiency of system services, reduces service costs, and improves market competitiveness with little loss of precision.

[0060] Embodiment 1

[0061] According to the embodiments of the present application, a method for accelerating the processing of a deep learning model is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0062] The method embodiments provided by the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 2 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for accelerating the processing of a deep learning model according to the embodiments of the present application. As shown in Figure 2 , the computer terminal 10 (or mobile device) can include one or more (in the figure, 102a, 102b, …, 102n are used to show) processors 102 (the processor 102 can include but not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication function. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 2 The structure shown is only schematic, which does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can also include more or fewer components than those shown in Figure 2 , or have a different configuration from that shown in Figure 2 .

[0063] It should be noted that the one or more processors 102 and / or other data processing circuitry described above can be referred to herein generically as "data processing circuitry". The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuitry can be a single standalone network layer, or any one of the other elements incorporated in whole or in part into a computer terminal 10 (or mobile device). As referred to in embodiments of the present application, the data processing circuitry acts as a processor to control, for example, the selection of variable resistance terminal paths connected to the interface.

[0064] The memory 104 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the method for accelerating deep learning model in embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the method for accelerating deep learning model described above. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0065] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet in a wireless manner.

[0066] The display can be, for example, a touch screen type liquid crystal display (LCD), which can enable a user to interact with the user interface of the computer terminal 10 (or mobile device).

[0067] It should be noted that in some optional embodiments, the above-mentioned Figure 2 The computer device (or mobile device) shown can include hardware elements (including circuitry), software elements (including computer code stored on a computer readable medium), or combinations of both hardware and software elements. It should be noted that Figure 2is merely one instance of a particular, specific example, and is intended to show the types of components that can be present in the above-described computer device (or mobile device).

[0068] In the above operating environment, the present application provides a method for accelerating a deep learning model as shown in Figure 3 Figure 3 is a flowchart of a method for accelerating a deep learning model according to Embodiment 1 of the present application.

[0069] Step S302, a quantization acceleration instruction is monitored, triggering a deep learning model to be executed for acceleration processing.

[0070] The deep learning model to be executed for acceleration processing can be a neural network model that has been pre-trained.

[0071] The above-mentioned deep learning model to be executed for acceleration processing can be a deep learning model such as a Transformer, but is not limited thereto, and can also be other types of deep models.

[0072] The above-mentioned quantization acceleration instruction includes but is not limited to Advanced Matrix Extension (AMX) and processing instruction (sdoit8). The quantization acceleration instruction can be an instruction that the user wishes to quantize and accelerate the deep learning model, and can be a specific button that determines whether to accelerate the deep learning model.

[0073] In an alternative embodiment, it can be determined whether the deep learning model needs to be accelerated according to the location where the deep learning model is deployed. If the deep learning model is deployed on a server, it does not need to be accelerated. If the deep learning model is deployed on a client, in order to improve processing speed and avoid space occupation of the deep learning model on the client, the deep learning model needs to be accelerated. The user can generate a quantization acceleration instruction each time the deep learning model is used, so as to accelerate the deep learning model according to the quantization acceleration instruction.

[0074] In an alternative embodiment, upon monitoring the quantization acceleration instruction, the deep learning model to be executed for acceleration processing can be triggered. Generally, lower precision or lower quantization bit width can be triggered at a faster speed. In addition, since the model is trained in full precision and there is no acceleration, the deep learning model to be accelerated can be a pre-trained neural network model.

[0075] Step S304, acquiring target data collected in a real environment.

[0076] ​The real environment described above can be a real running environment of the deep learning model, and the specific scene of the real environment is not limited here.

[0077] The target data collected in the real environment described above can be an image to be processed input to the deep learning model by a user actually requiring data processing, which is provided by the user according to actual processing requirements, and the type of the target data is not limited here, which can be an image, text, or voice, etc.

[0078] In step S306, the deep learning model is dynamically quantized based on the target data, and target quantization parameters of the deep learning model are obtained.

[0079] The target quantization parameters described above can reduce the representation range of the deep learning model to the corresponding quantization parameters when the representation range of the deep learning model is scaled to other representation ranges. For example, the input of FP32 can be scaled from the representation range of FP32 to the quantization parameters stored when the representation range of int8 or uint8 is scaled.

[0080] In an optional embodiment, the target data collected in the real environment can be used to dynamically quantize the deep learning model to generate a quantization table online, and the target quantization parameters of the deep learning model can be obtained through the quantization table. Through the design of software, the quantization process can not consume too much additional consumption. It should be noted that the quantization table is the quantization parameter, and since the quantization process is performed during the model inference process, the quantization process can be dynamically quantized.

[0081] Regarding the dynamic quantization part, the performance update principle is to update according to the time ratio, and the part with a large time ratio is updated first. In the deep learning model, there are two important operations that need to be updated. Dynamic quantization is generally divided into three parts: quantization, calculation, and dequantization. The quantization process can refer to scaling the input of FP32 from the representation range of FP32 to the representation range of int8 or uint8, and storing some quantization parameters. The calculation process can be a calculation-intensive calculation for low-precision input (input) and weight (weight). The dequantization operation can be the conversion of the int8 or uint8 output to the representation range of FP32 through the quantization parameters calculated in the front.

[0082] Figure 4is a flow chart of a dynamic quantization process according to an embodiment of the present application. First, high-precision input input and high-precision weight weight are taken as inputs. In the quantization process, the quantization parameters of input and the quantization parameters of weight are determined. The full-precision input and weight are quantized according to the quantization parameters of input and the quantization parameters of weight, to obtain a quantized result. The target quantization parameters are dequantized to obtain a target result. Optionally, the target quantization parameters are dequantized according to the input quantization parameters (Iquantparams) and the weight quantization parameters (Wquantparams).

[0083] The above dequantization refers to converting other representation ranges of the deep learning model into the original representation range of the deep learning model. For example, the output of the deep learning model in int8 or uint8 can be converted into the original FP32 representation range of the deep learning model through the target quantization parameters calculated in advance.

[0084] The quantization process is as follows:

[0085] Scaling + translation:

[0086] Rounding + truncation:

[0087] Round (Q f32 ) = Q Xbit = Q f32 + Δ

[0088] Clip: omitted;

[0089] The dequantization process is as follows:

[0090] D' f32 = (Q Xbit -ZeroPoint) * Scale

[0091] Wherein, D f32 is the original data; is the maximum value of Xbit; is the minimum value of Xbit; 8bit: [0, 255] or [-128, 127]; 4bit: [0, 16] or [-8, 7];

[0092] From the perspective of hardware, more and more quantization acceleration instructions are supported, such as int8 tensorCore, AMX, sdoit8. Moreover, the lower the precision and the lower the quantization bit width, the faster the speed. However, there is a difference that model training is usually trained with full precision, and inference will have a mapping relationship if low precision or quantization is used. This process is called quantization, and there is a key data called mapping table. How to generate the mapping table and what data to generate the table are the key designs of the quantization system.

[0093] In step S308, the target result of the target data is obtained by performing acceleration processing on the deep learning model by the acceleration processor based on the target quantization parameter.

[0094] The target result of the target data described above can be a processing result of the deep learning model. The target result described above is not limited to a recognition result or a calculation result.

[0095] The acceleration processor described above includes but is not limited to TRT, CUDNN.

[0096] In an optional embodiment, the input target data can be quantized according to the target quantization parameter to obtain quantized data corresponding to the input data. The quantized data can be accelerated by the acceleration processor according to the model weight of the deep learning model to obtain a quantization result corresponding to the quantized data. The quantization result can be dequantized to obtain a dequantization result of the target data, and the dequantization result output by the last network layer of the deep learning model is determined as the target result of the target data.

[0097] Through the above steps, the quantization acceleration instruction can be monitored to trigger the deep learning model to be executed for acceleration processing, wherein the deep learning model is a pre-trained neural network model. The target data collected in the real environment is obtained. The target quantization parameter of the deep learning model is obtained by dynamically quantizing the deep learning model based on the target data. The target result of the target data is obtained by performing acceleration processing on the deep learning model by the acceleration processor based on the target quantization parameter. The purpose of improving the processing efficiency of the deep learning model is achieved. It is easy to note that by dynamically quantizing the deep learning model based on the target data, the quantization table can be generated online, so that the process of quantization does not generate additional overhead, and the speed can be improved while ensuring the quantization accuracy, thereby solving the technical problem of low processing efficiency of the deep learning model in the related art.

[0098] The application is a way of online generating a quantization table, and through the design of software, the quantization process does not generate too much additional overhead, thereby being an acceleration system between quantization and full-precision (fp32) operation, balancing accuracy and speed, and enabling better deployment of quantized hardware.

[0099] In the above embodiment of the application, the deep learning model includes a plurality of network layers and model weights, the network layers in the deep learning model are dynamically quantized based on target data to obtain target quantization parameters of the network layers, including: determining input data input to the network layers based on the target data; determining quantization parameters of the network layers based on a numerical range of the input data and a preset numerical range, wherein the quantization parameters include a scaling parameter, a translation parameter and a dimension reduction parameter, wherein the scaling parameter is used for scaling operation on the input data, the translation parameter is used for translation operation on the input data, and the dimension reduction parameter is used for dimension reduction operation on the input data; and obtaining the target quantization parameters based on the quantization parameters of the plurality of network layers.

[0100] In the process of implementing quantization calculation, HIE introduces a dimension reduction parameter (ReduceSum) in addition to the standard scaling parameter (Scale) and translation parameter (ZeroPoint) parameters. By introducing the dimension reduction parameter, one additional memory access overhead can be eliminated, and performance can be further improved.

[0101] To solve the problem of conventional quantization, HIE adopts an asymmetric quantization + neural network quantization (PerChannel) granularity quantization scheme, and introduces an additional ReduceSum parameter. By introducing the additional parameter, one additional memory access overhead can be eliminated. This quantization scheme can be used for quantization of both GEMM and Conv (i.e., asymmetric + PerChannel granularity), but here only quantization for GEMM calculation is emphasized. Conv quantization can be performed using PerTensor quantization.

[0102] The calculation formula of asymmetric quantization is:

[0103] X_{int8}=Scale*X_{fp32}+ZeroPointX int8 =Scale*X fp32 +ZeroPoint

[0104] Figure 5 It is a schematic diagram of symmetric quantization and asymmetric quantization according to an embodiment of the application. Symmetric quantization cannot fully utilize the 8-bit representation range, while asymmetric quantization can make data shift after introducing a zero point parameter (ZeroPoint), so that the 8-bit representation range can be fully utilized.

[0105] In the above embodiment of the application, the network layer comprises a plurality of processing channels, and determining the quantization parameter of the network layer based on the numerical range of the input data and the preset numerical range comprises: determining channel data input to the processing channel based on the input data; determining the sub-quantization parameter of the processing channel based on the numerical range of the channel data and the preset numerical range; and obtaining the quantization parameter of the network layer based on the sub-quantization parameters of the plurality of processing channels.

[0106] The network layer described above can be a convolution layer.

[0107] The processing channel described above is PerChannel.

[0108] In an alternative embodiment, the numerical range of the channel data and the preset numerical range can be calculated to obtain the sub-quantization parameter of the processing channel, and the quantization parameter of the entire network layer can be obtained based on the sub-quantization parameters of the plurality of channels. The pre-calculated quantization parameter is mainly for two processes, one is matrix multiplication, the formula is input*weight=output, and the other is the convolution formula, which has two parts, one is input and the other is kernel, wherein the kernel is obtained by training.

[0109] The numerical range of the channel data described above can be a weight range.

[0110] The preset numerical range described above can be a fixed-point quantization range.

[0111] For weight quantization, since the data distribution of the weight is static, the minimum and maximum linear mapping can be directly found. However, for inference activation, the data distribution is dynamic, and in order to obtain the data distribution of the activation value, a so-called calibration set is often used to sample the distribution. After obtaining the sampling distribution, some quantization algorithms are used to select the quantization threshold.

[0112] For example, after model training, the weight or activation value is often distributed within a limited range, such as a weight range of [-2.0, 6.0], i.e. Tmax=6.0, Tmin=-0.2, and then using int8 to quantize the model, the fixed-point quantization value range is [-128, 127], i.e. Qmax=127, Qmin=-127. The sub-quantization parameters of the plurality of processing channels can be calculated based on the weight range and the fixed-point quantization value range, and the quantization parameter of the network layer can be obtained by superimposing the sub-quantization parameters of the plurality of processing channels.

[0113] Figure 6is a matrix multiplication schematic diagram according to an embodiment of the present application. Taking matrix multiplication as an example, the PerChannel quantization mode is to calculate a set of quantization parameters for each row and each column of the matrix respectively. In this way, it can be avoided that the global error is too large due to any abnormal row or column.

[0114] Taking matrix multiplication as an example, the size of the matrix multiplication is [M, N, K], and the calculation process of the asymmetric quantization is as follows:

[0115] Y = [S L · (X L -Z L ) - Z R ] R R ]

[0116] Wherein, X represents an 8Bit matrix, S represents a Scale of a left and right matrix, Z represents a ZeroPoint, and Y represents a final floating point output. Since the PerChannel quantization is adopted, the S and Z scalars in the formula need to be expanded into matrices, as follows:

[0117] Y = [S L · (X L -Z L ×E L ) - Z R · (X R -E R ×Z R )

[0118] Wherein, S L corresponds to (M, 1), X L corresponds to (M, k), Z L corresponds to (M, 1), E L corresponds to (1, K), S R corresponds to (1, N), X R corresponds to (K, N), E R corresponds to (K, 1), Z R corresponds to (1, N), X represents an 8Bit matrix, S represents a Scale matrix of a left and right matrix, Z represents a ZeroPoint matrix, and E L is a unit matrix with a size of (1, K), E R is a unit matrix with a size of (K, 1), and the corresponding sizes of various matrices can be obtained in turn.

[0119] According to the above formula, when performing matrix multiplication, the INT8 data needs to be dequantized into floating point data first, and then floating point matrix multiplication is performed. The above formula can be transformed as follows:

[0120] Y = [S​L • (X L -Z L ×E L )] × [S R • (X R -E R × Z R )] ;

[0121] Y = [S L × S R ] • [(X L -Z L × E L ) • (X R -E R × Z R )] ;

[0122] Y = [S L × S R ] • [X L × X R - X L × (E R × Z R ) - (Z L × E L ) × X R + (Z L × E L ) × ER × ZR;

[0123] Y = [S L × S R ] • [X L × X R - (X L × E R ) × Z R - Z L (E L × X R ) + Z L × (E L × E R ) × ZR;

[0124] Y = [S L × S R ] • [X L × X R - RS L × Z R - Z L × RS R + K • Z L × Z R ] ;

[0125] where RS L represents the reflection of X LThe result of ReduceSum of each row, RS R Indicates the X R ReduceSum of each column, X L X X R It is a pure 8Bit matrix multiplication, which is the most core and intensive calculation part, and the remaining part is some additional addition, subtraction and multiplication operations.

[0126] The last formula is the final quantization calculation formula. In order to separately calculate X L X X R , ReduceSum operation needs to be performed on the quantized matrix X L and X R . This will introduce an additional memory access operation. HIE completely embeds the ReduceSum operation in the quantization process, thereby eliminating the redundant memory access operation.

[0127] In the above embodiment of the application, before triggering the deep learning model to be executed for acceleration processing, the method further includes: performing static quantization processing on original weights in the deep learning model based on sample data to obtain model weights.

[0128] The static quantization described above can use an existing static quantization method to calculate the quantization parameters of the weights. Specifically, sample data collected in a real environment can be used to perform static quantization processing on the original weights in the deep learning model to obtain the model weights. Calibration data, test data, etc. can also be used to perform static quantization processing on the original weights in the deep learning model to obtain the model weights.

[0129] It should be noted that static quantization uses a set of verification sets to "estimate" the quantization parameters of each layer through statistical methods; and dynamic quantization calculates the quantization parameters according to the current input in real time, so as to achieve more accurate simulation calculation results.

[0130] In the above embodiment of the application, based on the target quantization parameter, the deep learning model is accelerated by the acceleration processor to obtain the target result of the target data, including: quantizing the input data input to the network layer based on the target quantization parameter to obtain quantization data corresponding to the input data; accelerating the quantization data based on the model weight by the acceleration processor to obtain a quantization result corresponding to the quantization data; and dequantizing the quantization result based on the target quantization parameter to obtain a dequantization result of the network layer, wherein the dequantization result of the last network layer in the plurality of network layers is the target result.

[0131] The above dequantization operation is the output of int8 or uint8 converted to the representation range of FP32 through the previously calculated quantization parameter.

[0132] In the above embodiment of the present application, the deep learning model further includes: a running algorithm adopted by the network layer running in the acceleration processor, and in a case where the running algorithm corresponding to the network layer is a matrix multiplication algorithm, the deep learning model is accelerated by the acceleration processor based on the target quantization parameter to obtain a target result of the target data, including: quantizing input data input to the network layer based on the target quantization parameter to obtain quantization data corresponding to the input data; performing acceleration processing on the quantization data based on the model weight by the acceleration processor using the fused quantization operator to obtain a quantization result corresponding to the quantization data; performing dequantization processing on the quantization result based on the target quantization parameter using the fused quantization operator to obtain a dequantization result corresponding to the network layer; and performing processing on the dequantization result based on the activation function using the fused quantization operator to obtain a processing result corresponding to the network layer.

[0133] The quantized s8GEMM update can be divided into three parts and two kernel implementations. First, the per-channel-quant op, then the s8GEMM-activation-fused op, and finally the post-processing part of the s8GEMM serial fusion dequant, bias add, and activation based on the Tensor Core, outputting the fp32 / fp16 result. The pre-processing (quantization) can be directly completed by inserting a dynamic quantization operator, or it can be executed by fusing it in the previous operator, or static quantization can also be used. HIE uses dynamic quantization + operator hiding. The following mainly explains the matrix multiplication fusion post-processing part, and develops and optimizes the performance based on the optimization logic of s8GEMM.

[0134] For the fused s8GEMM kernel, for performance consideration, the memory access instruction (mma.sync.m8n8k16) introduced in the parallel thread execution (ptx ver6.5) can be selected instead of the more flexible memory access instruction (wmma inline ptx or mma.sync.m8n8k4). The fused operator can quickly complete a set of 8x8x16 multiplication and accumulation operations in the warp view using the memory access instruction (mma.sync.m8n8k16). In order to meet the requirements of mma for register indexes of different threads in the same warp, for the T4 computing card supporting ldmatrix, we use ldmatrix.sync.aligned.m8n8 to complete the high-performance memory access from the shared memory to the register before mma. This instruction can allow each thread to complete the memory access with the equivalent bandwidth of lds128 while completing the partial transpose operation required by the B matrix block, but this also puts higher requirements on the reading and placing of the shared memory. Based on the use of tensor core, preliminary estimation of the calculation and memory ratio of the s8GEMM part can find that the memory of the dram and the shared memory on the T4 computing card constitutes a bottleneck.

[0135] The Gemm operator is a typical compute-intensive operator. Considering the need to accelerate the GPU in the task, within the error allowable range, inserting a quantization operator, while cooperating with the high-performance low-precision matrix multiplication and accumulation operation capability of the TensorCore, is a regular choice. After the quantization Gemm calculation is completed, a dequantization operation is required, and an activation function is often executed in series after the Gemm operator. Therefore, this group of operators can naturally be fused with the dequantization logic of the post-processing part of the s8GEMM to eliminate redundant access and further improve performance.

[0136] In the above embodiments of the present application, the acceleration processor uses the fused quantization operator to perform accelerated processing on the quantization data based on the model weight to obtain a quantization result corresponding to the quantization data, including: monitoring a plurality of different memory access instructions to obtain quantization data from a multi-level cache; monitoring a preset calling instruction to call the acceleration processor to use the fused quantization operator to perform accelerated processing on the quantization data based on the model weight to obtain a quantization result.

[0137] The plurality of different memory access instructions described above include but are not limited to ld.global.v4.u32 and ld.global.v2.u32.

[0138] The multi-level cache includes, but is not limited to, shared memory and thread register.

[0139] The preset calling instruction includes, but is not limited to, an mma / ldmatrix instruction.

[0140] The vectorized memory access can improve the memory access efficiency by using ldg128 / ldg64; the ordered index thread block improves the L2 cache hit rate, and a simple method can improve the effective memory bandwidth, that is, the vectorized memory access is performed under the premise of ensuring the memory access continuity between threads. In the current S8GEMM fusion kernel, global addresses (ld.global.v4.u32 and ld.global.v2.u32) can be used for A / B matrix blocks to reduce the number of instructions while meeting the thread block design size. The optimization of the L2 hit rate can further improve the equivalent bandwidth of the dram and the memory access efficiency. The basic idea of this part of logic is to recalculate the index (that is, m / n head index) of each thread block in the output C matrix, and try to ensure that the thread blocks running at the same time (in the same wave) are adjacent in the m / n dimension, thereby improving the memory locality of the A / B matrix block.

[0141] In the above embodiments of the application, the acceleration processor accelerates the processing of quantized data by starting a plurality of thread blocks, wherein the indexes of the thread blocks belonging to the same thread bundle are adjacent in the quantization result, the thread bundle is obtained by dividing at least one thread, and the thread is obtained by dividing a plurality of thread blocks; the thread blocks belonging to the same thread bundle read the same data block in the quantized data; and the thread blocks belonging to the same thread cache the quantized data read in the same register.

[0142] The shared memory and thread register are used to cache the A / B matrix blocks; and the shared memory is reasonably arranged to reduce the bank conflict. This part of work mainly focuses on sharing the shared memory. As mentioned in the foregoing, there is a memory access bottleneck on the T4 computing card, so it is necessary to use different methods to alleviate the bottleneck problem caused by the large difference in calculation memory access in the global to shared and shared to register two stages.

[0143] The first stage, thread block is indexed by warp, 256 threads are divided into 4x2 blocks by 8 warps, and the A / B corresponding matrix block is cached to the shared memory. The same row thread bundle (warp) reads the same A matrix block, and the same column warp reads the same B matrix block, thereby effectively reducing the dram access and relieving the calculation access ratio. In this stage, since the T4 computing card does not have a direct access channel from the global memory to the shared memory, it needs to be completed through the thread register. This requires the sts stage to maximize the storage bandwidth of the shared memory by using st.shared.v4.b32, and the existence of this stage also provides convenience for the necessary placement and conversion of sts128 before. The second stage uses a method similar to the matrix outer product, and uses the register to cache the A / B blocks required for calculation on each thread. This also matches the ldmatrix instruction, effectively reducing the repeated access of the shared memory.

[0144] In practice, in order to avoid the occurrence of shared memory bank conflict, the shared memory can be used to stagger the index.

[0145] In the above embodiments of the present application, the fused quantization operator includes double buffering and multi-stage pipelining.

[0146] Considering the existing register consumption, lower occupancy is not enough to meet the scheduling requirements of the warp scheduler. This requires us to introduce double buffering and multi-stage pipelining in the kernel to cover the warp level delay and improve hardware utilization. According to the current pipelining design and register usage, ldg128 / 64+sts128 and ldmatrix+mma can be covered by double buffer on the shared memory. Figure 7 FIG. 1 is a schematic diagram of a fused quantization operator according to an embodiment of the present application.

[0147] Based on the existing s8GEMM design, after each thread completes its own mma calculation, we need to complete the dequantization, bias and activation calculation step by step. But limited by the number of registers, low occupancy and difficult to parallel logic, combined with the characteristics of the actual model post-processing time-consuming proportion is not high, the post-processing part is simply to complete the calculation step by step through the shared memory transpose of the int32GEMM result, and vectorized write back to the dynamic memory (dram). In the actual test, we also found that introducing complex pipeline design in post-processing is a futile negative optimization, which cannot solve the problem of warp stall caused by the difficulty of instruction overlap.

[0148] Figure 8 is a performance test comparison diagram according to an embodiment of the application. The performance test comparison of single-precision matrix operation (cublassgemm) and dynamic quantization single-precision (s8gemm) (including dequantization + activation function) can be seen from the figure. Using s8gemm with dynamic quantization can achieve significant performance improvement under various typical sizes. Cublasfp32: GEMM performance of cublasfp32; s8GEMM-activation-fused-fp32 is accumulated to fp32 after dynamic quantization; s8GEMM-activation-fused-fp16 is accumulated to fp16 after dynamic quantization.

[0149] In the above embodiment of the application, the deep learning model further includes: a running algorithm adopted by the network layer running in the acceleration processor, and in a case where the running algorithm corresponding to the network layer is a convolution algorithm, the acceleration processor is used to accelerate processing of the deep learning model based on the target quantization parameter to obtain a target result of the target data, including: converting the convolution algorithm into a matrix multiplication algorithm; quantizing input data input to the network layer based on the target quantization parameter to obtain quantized data corresponding to the input data; using the acceleration processor to accelerate processing of the quantized data based on the model weight by using the fused quantization operator to obtain a quantization result corresponding to the quantized data; using the fused quantization operator to perform dequantization processing on the quantization result based on the target quantization parameter to obtain a dequantization result corresponding to the network layer; and using the fused quantization operator to process the dequantization result based on the activation function to obtain a processing result corresponding to the network layer.

[0150] Similar to GEMM, Conv is also a compute-intensive OP, and this OP has many implementation methods and update methods, so generally a library provided by a platform is called to implement it, such as on a GPU, an acceleration library (CuDNN) is called to implement it. CuDNN and the previous quantization line of Nvidia do not use dynamic quantization for acceleration, and in actual use, according to the current version support, there is no way to do dynamic quantization acceleration. This also has the innovation point of the system on Conv.

[0151] Figure 9 is a flowchart of another dynamic quantization processing process in the embodiment of the application, as shown in Figure 9 The implementation method here is to convert the Conv into a matrix multiplication algorithm by using the method of calculating the convolution through the matrix multiplication algorithm (Img2Col), which is a relatively common method, and then perform low-precision calculation through the dynamic quantization fusion described above, and output Fp32. That is, the input data input to the network layer is quantized based on the target quantization parameter to obtain quantized data corresponding to the input data, the acceleration processor uses the fusion quantization operator to perform acceleration processing on the quantized data based on the model weight to obtain a quantization result corresponding to the quantized data, the fusion quantization operator is used to perform dequantization processing on the quantization result based on the target quantization parameter to obtain a dequantization result corresponding to the network layer, and finally the fusion quantization operator is used to process the dequantization result based on the activation function to obtain a processing result corresponding to the network layer. The innovation point of this and the usual method is that the dynamic quantization fusion process described in the GEMM part is used, so that the dequantization memory access overhead is saved, and efficient calculation is realized.

[0152] In the above embodiment of the application, the deep learning model to be executed is triggered to perform acceleration processing, including: obtaining a model file of the deep learning model, wherein the model file is used to store original weights and a plurality of network layers; and analyzing the model file to obtain the deep learning model.

[0153] Since the weights do not change, the deep learning model can be obtained by analyzing the original weights and the plurality of network layers in the model file. In order to pre-calculate part of the quantization parameters, thereby improving the efficiency of the runtime.

[0154] In the above embodiments of the present application, the model file is parsed to obtain a deep learning model, including: parsing the model file to obtain initial model information of the deep learning model, wherein the initial model information at least includes: original weights, network structure of a network layer; performing graph optimization on the initial model information to obtain optimized model information; determining a running algorithm adopted by the network layer in the optimized model information, a data format of input data input to the network layer, and a data format of output data output from the network layer, to obtain the deep learning model.

[0155] The running algorithm described above can be a specific implementation algorithm of different network structures, for example, the implementation methods of the fast convolution algorithm include Img2Col, Winograd, ImpIPreCompute-GEMM, etc.

[0156] The data format described above can be a data arrangement format, for example, NCHW, NHWC, NCHW32, etc.

[0157] The graph optimization described above can be optimization of data of the initial model information to obtain optimized model information that can be used for replacement.

[0158] Figure 10 According to an embodiment of the present application, a block diagram for accelerating a deep learning model is as shown in Figure 10 , first, the model file is parsed, then an operator calculation graph is obtained, the static graph can be updated, then static quantization is performed to obtain a dynamic operator adjustment (OPTUNE), an updated running operator calculation graph can be obtained, and the running operator calculation graph is calculated. Wherein, the model file can be parsed, the parsed model generates a directed acyclic graph (Directed Acyclic Graph, abbreviated as DAG) composed of a calculation operator; static quantization can be the calculation of quantization parameters of weights; dynamic quantization operator adjustment can be the selection of the fastest implementation of the same OP; the updated running operator calculation graph can be the calculation graph with the fastest running speed on this hardware; dynamic quantization calculation can use the weights quantized in the front and the input of OP to perform quantization calculation.

[0159] Figure 11 According to an embodiment of the present application, a flow chart of a process for accelerating a deep learning model is as shown in Figure 11

[0160] In the first step, the deep learning model can be taken as input, wherein the deep learning model includes but is not limited to onnx.

[0161] ​Second step, after the static graph step, get an updated graph, this process is static graph selection (StaticGraphOPT).

[0162] Third step, static quantization (StaticQuant) part here is mainly the processing of static quantization, and the weight processing of the dynamic quantization process. Since the weight does not change, a part of the quantization parameters can be calculated in advance here, thereby improving the running efficiency. The pre-calculated quantization parameters here are mainly for two processes. One is matrix multiplication, the formula is input * weight = output. The other is the convolution formula, which has two parts, one is input and the other is kernel (also known as weight). The kernel is obtained by training. The pre-computation here is the quantization method mentioned in the present application.

[0163] Fourth step, after dynamic quantization adjustment (DynamicOPTune), each OP in the previous model, such as Conv, may have many implementations, such as Conv, fast convolution algorithm (Img2Col, Winograd, ImpIPreCompute-GEMM), and different data arrangement formats (NCHW, NHWC, NCHW32). The running time of each format and each running algorithm combination is different, which is a common mechanism. It should be noted that since the model itself is a calculation graph composed of a bunch of Operators, each OP here refers to each OP in the calculation graph. In addition, since an Operator has many implementations, such as avx2 assembly implementation on X86 CPU, and avx512 assembly implementation, different implementations run at different speeds or have different support levels on different hardware. Therefore, in the running time, each input is first run with a random input, and then the fastest implementation is selected.

[0164] Fifth step, in the fourth step, dynamic quantization implementation for some key OPs is added to the system, such as DynamicQuantConv for Conv. This implementation uses dynamic quantization.

[0165] Sixth step, after quantization adjustment, the fastest combination of the entire model is selected. The combination here refers to the fastest combination of different implementations of each OP in the calculation graph, including the quantization implementation of the OP and other implementations such as non-quantization but low-precision implementation, such as fp16.

[0166] It should be noted that the above calculation process can run on various systems at present, but the update focus and direction of each system is different.

[0167] The natural language processing library (HuggingFace) referring to the standard visual model (Swin) speed in the large visual database (ImageNet) is as follows: model: swin-base-patch4-window7-224 (ONNX model exported from HuggingFace), test set: ImageNet, and the accuracy test data is shown in Table 1:

[0168] Table 1

[0169]

[0170]

[0171] Note: The TensorRT version is 8.4.1.5

[0172] HIE-Mix is the mixed precision of HIE, which mixes float16 and int8. From the table, we can see that under the mixed precision mode of HIE, the algorithm accuracy only drops by 4 per million, which should meet most functional requirements. However, the test result of TensorRT in FP16 precision is a random value, and the algorithm accuracy loss is serious. We guess that the reason may be that TensorRT uses FP16 registers as accumulation registers, causing overflow, while HIE uses FP32 as accumulation registers, so there is no similar problem.

[0173] Figure 12 is a performance test comparison diagram according to an embodiment of the present application. The mixed precision of HIE-Mix has an acceleration ratio of 1.8 times compared to TensorRT-FP32. Compared with Pytorch-FP32, the acceleration ratio is as high as 7.3 times. In the test, the QPS of TensorRT-FP16 reached 263, which is much higher than that of HIE-Mix, but due to the serious problem of algorithm accuracy, it is not listed as a comparison object.

[0174] Conv dynamic quantization:

[0175] Compared with ORT, it can achieve a speed advantage of 1.6-3.3, and maintain much less precision loss than static quantization. The accuracy is shown in Table 2 as follows:

[0176] Table 2

[0177] Precision FP32 74.97% Dequantization (HIE-DyQuant) 74.44%(-0.5%) Static Quantization (HIE-StaticQuant) 74.01%(-0.96%)

[0178] The speed is shown in Table 3 as follows:

[0179] Table 3

[0180]

[0181]

[0182] Design and architecture of dynamic quantization and efficient implementation: the application is a way to generate a quantization table online, and through the design of software, the quantization process will not generate too much additional overhead, so it is an acceleration system between quantization and full precision (fp32) operation.

[0183] Weight preprocessing part of the calculation part of dynamic quantization: the static quant part here is mainly the processing of static quantization and the weight processing of the dynamic quantization process. Because the weight does not change, a part of the quantization parameters can be calculated in advance here, so as to improve the runtime efficiency.

[0184] Dynamic quantization operator, dynamic quantization of Conv and GEMM, dynamic quantization implementation for some key OPs is added to the system, such as dynamic quantization implementation of Conv. This implementation uses dynamic quantization technology.

[0185] Innovation in quantization method, PreChannel + asymmetric quantization + ReduceSum parameter introduction, high precision and low calculation amount quantization scheme is achieved.

[0186] Implementation method of GEMM fusion operator, matrix multiplication fusion post-processing: s8GEMMactivation-fused op. Conv based on image2Col + fusion GEMM calculation method.

[0187] Through the above dynamic quantization method of the application, the artificial intelligence system can be fast and accurate, improve the efficiency of system service, reduce service cost, and improve market competitiveness.

[0188] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the application is not limited by the action sequence described, because according to the application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the application.

[0189] Those skilled in the art can clearly understand that the method according to the above-mentioned embodiments can be realized by means of software and necessary general hardware platforms, and of course can also be realized by hardware based on the above description of the embodiments. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, or network device) execute the method of each embodiment of the present application.

[0190] Embodiment 2

[0191] According to the embodiments of the present application, a data processing apparatus for implementing the above-mentioned method of accelerating the deep learning model is further provided, Figure 13 is a schematic diagram of a data processing apparatus according to Embodiment 2 of the present application, as Figure 13 shown, the apparatus 1300 includes a trigger model 1302, an acquisition module 1304, a dynamic quantization module 1306, and an acceleration network layer 1308.

[0192] The trigger module is configured to monitor a quantization acceleration instruction and trigger a deep learning model to be executed for acceleration processing, wherein the deep learning model is a pre-trained neural network model; the acquisition module is configured to acquire target data collected in a real environment; the dynamic quantization module is configured to perform dynamic quantization processing on the deep learning model based on the target data to obtain target quantization parameters of the deep learning model; and the acceleration network layer is configured to perform acceleration processing on the deep learning model based on the target quantization parameters by using an acceleration processor to obtain a target result of the target data.

[0193] It should be noted that the trigger model 1302, the acquisition module 1304, the dynamic quantization module 1306, and the acceleration network layer 1308 correspond to steps S202 to S208 in Embodiment 1, and the four modules have the same instances and application scenarios as the corresponding steps, but are not limited to the contents disclosed in Embodiment 1. It should be noted that the above modules as part of the apparatus can run in the computer terminal provided in Embodiment 1.

[0194] In the above embodiments of the present application, the deep learning model includes a multi-layer network layer and a model weight, and the dynamic quantization module includes a determination unit and a generation unit.

[0195] The determining unit is configured to determine input data input to the network layer based on target data; and the determining unit is further configured to determine a quantization parameter of the network layer based on a numerical range of the input data and a preset numerical range, wherein the quantization parameter comprises a scaling parameter, a translation parameter and a dimension reduction parameter, the scaling parameter is used for scaling operation on the input data, the translation parameter is used for translation operation on the input data, and the dimension reduction parameter is used for dimension reduction operation on the input data; and the generating unit is configured to obtain a target quantization parameter based on the quantization parameters of the multiple network layers.

[0196] In the above embodiments of the present application, the network layer comprises a plurality of processing channels, the determining unit is configured to determine channel data input to the processing channels based on the input data; the determining unit is further configured to determine a sub-quantization parameter of the processing channel based on a numerical range of the channel data and a preset numerical range; and the first determining unit is configured to obtain the quantization parameter of the network layer based on the sub-quantization parameters of the multiple processing channels.

[0197] In the above embodiments of the present application, the device further comprises a static quantization processing module.

[0198] The static quantization processing module is configured to perform static quantization processing on original weights in the deep learning model based on sample data to obtain model weights.

[0199] In the above embodiments of the present application, the accelerated network layer comprises a quantization unit, an acceleration processing unit and a dequantization processing unit.

[0200] The quantization unit is configured to quantize input data input to the network layer based on a target quantization parameter to obtain quantization data corresponding to the input data; the acceleration processing unit is configured to perform acceleration processing on the quantization data based on the model weights by using an acceleration processor to obtain a quantization result corresponding to the quantization data; and the dequantization processing unit is configured to perform dequantization processing on the quantization result based on the target quantization parameter to obtain a dequantization result of the network layer, wherein the dequantization result of the last network layer in the multiple network layers is a target result.

[0201] In the above embodiments of the present application, the deep learning model further comprises a running algorithm adopted by the network layer when running in the acceleration processor, the dequantization processing unit is further configured to quantize input data input to the network layer based on a target quantization parameter to obtain quantization data corresponding to the input data; the dequantization processing unit is further configured to perform acceleration processing on the quantization data based on the model weights by using a fusion quantization operator of the acceleration processor to obtain a quantization result corresponding to the quantization data; the dequantization processing unit is further configured to perform dequantization processing on the quantization result based on the target quantization parameter by using the fusion quantization operator to obtain a dequantization result corresponding to the network layer; and the dequantization processing unit is further configured to perform processing on the dequantization result based on an activation function by using the fusion quantization operator to obtain a processing result corresponding to the network layer.

[0202] In the above embodiments of the present application, the inverse quantization processing unit is further configured to monitor a plurality of different memory access instructions to obtain quantized data from the multi-level cache; and the inverse quantization processing unit is further configured to monitor a preset calling instruction to call the acceleration processor to perform acceleration processing on the quantized data based on the model weight using the fused quantization operator to obtain a quantization result.

[0203] In the above embodiments of the present application, the acceleration processor performs acceleration processing on the quantized data by starting a plurality of thread blocks, wherein the indexes of the thread blocks belonging to the same thread bundle are adjacent in the quantization result, the thread bundle is obtained by dividing at least one thread, and the thread is obtained by dividing a plurality of thread blocks; the thread blocks belonging to the same thread bundle read the same data block in the quantized data; and the thread blocks belonging to the same thread cache the quantized data read in the same register.

[0204] In the above embodiments of the present application, the fused quantization operator includes double buffering and multi-stage pipelining.

[0205] In the above embodiments of the present application, the deep learning model further includes a running algorithm used by a network layer to run in the acceleration processor, and the acceleration network layer includes a conversion unit, a quantization unit, an acceleration processing unit, and a processing unit.

[0206] The conversion unit is configured to convert the convolution algorithm into a matrix multiplication algorithm; the quantization unit is configured to quantize input data input to the network layer based on a target quantization parameter to obtain quantized data corresponding to the input data; the acceleration processing unit is configured to perform acceleration processing on the quantized data based on the model weight using the fused quantization operator to obtain a quantization result corresponding to the quantized data; the quantization unit is configured to perform inverse quantization processing on the quantization result based on the target quantization parameter using the fused quantization operator to obtain an inverse quantization result corresponding to the network layer; and the processing unit is configured to process the inverse quantization result based on an activation function using the fused quantization operator to obtain a processing result corresponding to the network layer.

[0207] In the above embodiments of the present application, the triggering module includes an obtaining unit and an analyzing unit.

[0208] The obtaining unit is configured to obtain a model file of the deep learning model, wherein the model file is used to store original weights and a plurality of network layers; and the analyzing unit is configured to analyze the model file to obtain the deep learning model.

[0209] It should be noted that the above modules or units can be hardware components or software components stored in a memory (for example, the memory 104) and processed by one or more processors (for example, the processors 102a, 102b, …, 102n), and the above modules can also be run in the computer terminal 10 provided in Embodiment One as part of the device.

[0210] Embodiment 3

[0211] Embodiments of the present application can provide an electronic device, which can be any one of the computer terminals in the computer terminal group. Alternatively, in the embodiments, the electronic device can be replaced by a terminal device such as a mobile terminal.

[0212] Alternatively, in the embodiments, the electronic device can be located in at least one of the network devices in the computer network.

[0213] In the embodiments, the electronic device can execute program codes for the following steps in a method for accelerating a deep learning model: monitoring a quantization acceleration instruction, triggering a deep learning model to be executed for acceleration, wherein the deep learning model is a pre-trained neural network model; obtaining target data collected in a real environment; performing dynamic quantization processing on the deep learning model based on the target data to obtain target quantization parameters of the deep learning model; and performing acceleration processing on the deep learning model based on the target quantization parameters by using an acceleration processor to obtain a target result of the target data.

[0214] Alternatively, Figure 14 is a structural block diagram of a computer terminal according to an embodiment of the present application. As shown in Figure 14 the computer terminal A can include one or more (only one is shown in the figure) processors 102, a memory 104, a storage controller, and a peripheral interface, wherein the peripheral interface is connected with a radio frequency module, an audio module, and a display.

[0215] The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the method and device for accelerating a deep learning model in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the method for accelerating a deep learning model described above. The memory can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the terminal A through a network. Examples of the network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0216] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: monitoring the quantization acceleration instruction, triggering the deep learning model to be executed for acceleration processing, wherein the deep learning model is a pre-trained neural network model; acquiring target data collected in a real environment; performing dynamic quantization processing on the deep learning model based on the target data to obtain target quantization parameters of the deep learning model; and performing acceleration processing on the deep learning model based on the target quantization parameters through the acceleration processor to obtain a target result of the target data.

[0217] Optionally, the processor can further execute program codes of the following steps: determining input data input into the network layer based on the target data; determining the quantization parameters of the network layer based on the numerical range of the input data and the preset numerical range, wherein the quantization parameters include: a scaling parameter, a translation parameter and a dimension reduction parameter, wherein the scaling parameter is used for scaling operation on the input data, the translation parameter is used for translation operation on the input data, and the dimension reduction parameter is used for dimension reduction operation on the input data; and obtaining the target quantization parameters based on the quantization parameters of the multiple network layers.

[0218] Optionally, the processor can further execute program codes of the following steps: determining channel data input into the processing channel based on the input data; determining the sub-quantization parameters of the processing channel based on the numerical range of the channel data and the preset numerical range; and obtaining the quantization parameters of the network layer based on the sub-quantization parameters of the multiple processing channels.

[0219] Optionally, the processor can further execute program codes of the following steps: performing static quantization processing on the original weights in the deep learning model based on the sample data to obtain model weights.

[0220] Optionally, the processor can further execute program codes of the following steps: quantizing the input data input into the network layer based on the target quantization parameters to obtain quantization data corresponding to the input data; performing acceleration processing on the quantization data based on the model weights through the acceleration processor to obtain a quantization result corresponding to the quantization data; and performing dequantization processing on the quantization result based on the target quantization parameters to obtain a dequantization result of the network layer, wherein the dequantization result of the last network layer in the multiple network layers is the target result.

[0221] Optionally, the processor can further execute program codes of the following steps: quantizing input data input into the network layer based on the target quantization parameter to obtain quantization data corresponding to the input data; performing, by the acceleration processor, accelerated processing on the quantization data based on the model weight by using the fused quantization operator to obtain a quantization result corresponding to the quantization data; performing, by the fused quantization operator, dequantization processing on the quantization result based on the target quantization parameter to obtain a dequantization result corresponding to the network layer; and performing, by the fused quantization operator, processing on the dequantization result based on the activation function to obtain a processing result corresponding to the network layer.

[0222] Optionally, the processor can further execute program codes of the following steps: monitoring a plurality of different memory access instructions to obtain quantization data from the multi-level cache; and monitoring a preset calling instruction to call the acceleration processor to perform accelerated processing on the quantization data based on the model weight by using the fused quantization operator to obtain a quantization result.

[0223] Optionally, the processor can further execute program codes of the following steps: performing, by the acceleration processor, accelerated processing on the quantization data by starting a plurality of thread blocks, wherein indexes of thread blocks belonging to a same thread bundle are adjacent in the quantization result, the thread bundle is obtained by dividing at least one thread, and the thread is obtained by dividing a plurality of thread blocks; thread blocks belonging to the same thread bundle read a same data block in the quantization data; and thread blocks belonging to the same thread cache the quantization data read in a same register.

[0224] Optionally, the processor can further execute program codes of the following steps: the fused quantization operator comprises double buffering and multi-stage pipelining.

[0225] Optionally, the processor can further execute program codes of the following steps: converting a convolution algorithm into a matrix multiplication algorithm; quantizing input data input into the network layer based on the target quantization parameter to obtain quantization data corresponding to the input data; performing, by the acceleration processor, accelerated processing on the quantization data based on the model weight by using the fused quantization operator to obtain a quantization result corresponding to the quantization data; performing, by the fused quantization operator, dequantization processing on the quantization result based on the target quantization parameter to obtain a dequantization result corresponding to the network layer; and performing, by the fused quantization operator, processing on the dequantization result based on the activation function to obtain a processing result corresponding to the network layer.

[0226] Optionally, the processor can further execute program codes of the following steps: obtaining a model file of the deep learning model, wherein the model file is used to store original weights and a plurality of network layers; and parsing the model file to obtain the deep learning model.

[0227] Optionally, the processor can further execute program codes of the following steps: parsing the model file to obtain initial model information of the deep learning model, wherein the initial model information at least includes original weights and network structures of network layers; performing graph optimization on the initial model information to obtain optimized model information; determining running algorithms adopted by the network layers in the acceleration processor, data formats of input data input to the network layers, and data formats of output data output from the network layers in the optimized model information, to obtain the deep learning model.

[0228] By adopting the embodiment of the application, the quantization acceleration instruction can be monitored to trigger the deep learning model to be executed for acceleration processing, the target data collected in a real environment is acquired, the deep learning model is dynamically quantized based on the target data to obtain target quantization parameters of the deep learning model, and the deep learning model is processed by the acceleration processor based on the target quantization parameters to obtain a target result of the target data, thereby achieving the purpose of improving the processing efficiency of the deep learning model. It is easy to note that the quantization table can be generated online by dynamically quantizing the deep learning model based on the target data, so that the process of quantization does not generate additional overhead, the speed is improved while the quantization accuracy is ensured, and thus the technical problem of low processing efficiency of the deep learning model in the related art is solved.

[0229] Those skilled in the art can understand that Figure 14 The structure shown is only schematic, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, or other terminal devices. Figure 14 It does not limit the structure of the electronic device. For example, the computer terminal A can further include more or fewer components (such as a network interface, a display device, etc.) than Figure 14 or have a different configuration from Figure 14 The structure shown.

[0230] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the related hardware of the terminal device through a program, and the program can be stored in a computer readable storage medium, which can include a flash disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, etc.

[0231] Embodiment 4

[0232] The embodiment of the present application also provides a storage medium. Optionally, in the embodiment, the storage medium can be used to save the program code executed by the method for accelerating the deep learning model provided in the first embodiment.

[0233] Optionally, in the embodiment, the storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0234] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: monitoring the quantization acceleration instruction, triggering the deep learning model to be executed for acceleration processing, wherein the deep learning model is a pre-trained neural network model; obtaining target data collected in a real environment; performing dynamic quantization processing on the deep learning model based on the target data to obtain target quantization parameters of the deep learning model; and performing acceleration processing on the deep learning model by using an acceleration processor based on the target quantization parameters to obtain a target result of the target data.

[0235] Optionally, the storage medium is further configured to store program code for performing the following steps: determining input data input to the network layer based on the target data; determining the quantization parameters of the network layer based on the numerical range of the input data and the preset numerical range, wherein the quantization parameters include a scaling parameter, a translation parameter and a dimension reduction parameter, wherein the scaling parameter is used for scaling operation on the input data, the translation parameter is used for translation operation on the input data, and the dimension reduction parameter is used for dimension reduction operation on the input data; and obtaining the target quantization parameters based on the quantization parameters of the multiple network layers.

[0236] Optionally, the storage medium is further configured to store program code for performing the following steps: determining channel data input to the processing channel based on the input data; determining the sub-quantization parameters of the processing channel based on the numerical range of the channel data and the preset numerical range; and obtaining the quantization parameters of the network layer based on the sub-quantization parameters of the multiple processing channels.

[0237] Optionally, the storage medium is further configured to store program code for performing the following steps: performing static quantization processing on the original weights in the deep learning model based on the sample data to obtain model weights.

[0238] Optionally, the storage medium is further configured to store program code for performing the following steps: quantizing input data input into the network layer based on the target quantization parameter to obtain quantization data corresponding to the input data; performing, by the acceleration processor, acceleration processing on the quantization data based on the model weight to obtain a quantization result corresponding to the quantization data; and performing dequantization processing on the quantization result based on the target quantization parameter to obtain a dequantization result of the network layer, wherein the dequantization result of the last network layer in the plurality of network layers is the target result.

[0239] Optionally, the storage medium is further configured to store program code for performing the following steps: quantizing input data input into the network layer based on the target quantization parameter to obtain quantization data corresponding to the input data; performing, by the acceleration processor, acceleration processing on the quantization data based on the model weight to obtain a quantization result corresponding to the quantization data; performing dequantization processing on the quantization result based on the target quantization parameter to obtain a dequantization result of the network layer; and performing processing on the dequantization result based on the activation function to obtain a processing result of the network layer.

[0240] Optionally, the storage medium is further configured to store program code for performing the following steps: monitoring a plurality of different memory access instructions to obtain quantization data from the multi-level cache; and monitoring a preset calling instruction to call the acceleration processor to perform acceleration processing on the quantization data based on the model weight by using the fused quantization operator to obtain a quantization result.

[0241] Optionally, the storage medium is further configured to store program code for performing the following steps: performing, by the acceleration processor, acceleration processing on the quantization data by starting a plurality of thread blocks, wherein indexes of thread blocks belonging to a same thread bundle are adjacent in the quantization result, the thread bundle is obtained by dividing at least one thread, and the thread is obtained by dividing a plurality of thread blocks; the thread blocks belonging to the same thread bundle read a same data block in the quantization data; and the thread blocks belonging to the same thread cache the quantization data read in a same register.

[0242] Optionally, the storage medium is further configured to store program code for performing the following steps: the fused quantization operator comprises double buffering and multi-stage pipelining.

[0243] Optionally, the storage medium is further configured to store program code for performing the following steps: converting the convolution algorithm into a matrix multiplication algorithm; quantizing input data input into the network layer based on a target quantization parameter to obtain quantized data corresponding to the input data; performing, by the acceleration processor, accelerated processing on the quantized data based on the model weight by using the fused quantization operator to obtain a quantization result corresponding to the quantized data; performing, by the fused quantization operator, dequantization processing on the quantization result based on the target quantization parameter to obtain a dequantization result corresponding to the network layer; and performing, by the fused quantization operator, processing on the dequantization result based on the activation function to obtain a processing result corresponding to the network layer.

[0244] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining a model file of the deep learning model, wherein the model file is used to store original weights and a plurality of network layers; and parsing the model file to obtain the deep learning model.

[0245] Optionally, the storage medium is further configured to store program code for performing the following steps: parsing the model file to obtain initial model information of the deep learning model, wherein the initial model information at least includes the original weights, and a network structure of the network layer; performing graph optimization on the initial model information to obtain optimized model information; determining a running algorithm adopted by the network layer in the optimized model information for running in the acceleration processor, a data format of input data input into the network layer, and a data format of output data output from the network layer to obtain the deep learning model.

[0246] By adopting the embodiments of the present application, the quantization acceleration instruction can be monitored to trigger the deep learning model to be executed for accelerated processing, wherein the deep learning model is a pre-trained neural network model; target data collected in a real environment is obtained; the deep learning model is dynamically quantized based on the target data to obtain a target quantization parameter of the deep learning model; and the deep learning model is processed by the acceleration processor based on the target quantization parameter to obtain a target result of the target data, thereby achieving the purpose of improving the processing efficiency of the deep learning model. It is easy to note that, by dynamically quantizing the deep learning model based on the target data, a quantization table can be generated online, so that the process of quantization does not generate additional overhead, and the speed can be improved while ensuring the quantization accuracy, thereby solving the technical problem of low processing efficiency of the deep learning model in the related art.

[0247] The serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0248] In the above-described embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0249] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented in other manners. The above described apparatus embodiments are merely exemplary, and the units can be divided into other manners, or some features can be ignored or not performed. In addition, the display or discussion of relative coupling or direct coupling or communication connection between the units can be indirect coupling or communication connection through some interfaces, and can be electrical or other forms.

[0250] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place or distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments.

[0251] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0252] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the embodiments of the present application. The foregoing storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, and various program codes that can be stored in the medium.

[0253] The above are only the preferred embodiments of the present application, and it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, some improvements and refinements can be made, and these improvements and refinements should be regarded as the protection scope of the present application.

Claims

1. A method for accelerating processing of a deep learning model, the method comprising: The method comprises: monitoring a quantization acceleration instruction, triggering a deep learning model to be executed for acceleration processing, wherein the deep learning model is a pre-trained neural network model; acquiring target data collected in a real environment; performing dynamic quantization processing on the deep learning model based on the target data to obtain target quantization parameters of the deep learning model; performing acceleration processing on the deep learning model based on the target quantization parameters by using an acceleration processor to obtain a target result of the target data, wherein the target result is a dequantization result of a last network layer in a plurality of network layers included in the deep learning model, the dequantization result is determined based on quantization data corresponding to the network layer and the acceleration processor, and the quantization data is obtained by quantizing input data of the network layer based on the target quantization parameters.

2. The method of claim 1, wherein, The deep learning model comprises a plurality of network layers and model weights, and the network layers in the deep learning model are dynamically quantized based on the target data to obtain target quantization parameters of the network layers, comprising: determining input data input to the network layer based on the target data; determining quantization parameters of the network layer based on a numerical range of the input data and a preset numerical range, wherein the quantization parameters comprise a scaling parameter, a translation parameter, and a dimension reduction parameter, wherein the scaling parameter is used for scaling operation on the input data, the translation parameter is used for translation operation on the input data, and the dimension reduction parameter is used for dimension reduction operation on the input data; obtaining the target quantization parameters based on quantization parameters of a plurality of network layers.

3. The method of claim 2, wherein, The network layer comprises a plurality of processing channels, and the quantization parameters of the network layer are determined based on a numerical range of the input data and a preset numerical range, comprising: determining channel data input to the processing channel based on the input data; determining sub-quantization parameters of the processing channel based on a numerical range of the channel data and the preset numerical range; obtaining the quantization parameters of the network layer based on sub-quantization parameters of a plurality of processing channels.

4. The method of claim 2, wherein, Before triggering the deep learning model to be executed for acceleration processing, the method further comprises: performing static quantization processing on original weights in the deep learning model based on sample data to obtain the model weights.

5. The method of claim 2, wherein, Performing acceleration processing on the deep learning model based on the target quantization parameters by using an acceleration processor to obtain a target result of the target data, comprising: quantizing input data input to the network layer based on the target quantization parameters to obtain quantization data corresponding to the input data; performing acceleration processing on the quantization data based on the model weights by using the acceleration processor to obtain a quantization result corresponding to the quantization data; performing dequantization processing on the quantization result based on the target quantization parameters to obtain a dequantization result of the network layer, wherein the dequantization result of a last network layer in a plurality of network layers is the target result.

6. The method of claim 2, wherein, The deep learning model further comprises: a running algorithm adopted by the network layer running in the acceleration processor, and in a case where the running algorithm corresponding to the network layer is a matrix multiplication algorithm, the target data is obtained by performing acceleration processing on the deep learning model by the acceleration processor based on the target quantization parameter, comprising: quantizing input data input to the network layer based on the target quantization parameter to obtain quantized data corresponding to the input data; performing acceleration processing on the quantized data based on the model weight by the acceleration processor using a fusion quantization operator to obtain a quantization result corresponding to the quantized data; performing dequantization processing on the quantization result based on the target quantization parameter using the fusion quantization operator to obtain a dequantization result corresponding to the network layer; performing processing on the dequantization result based on an activation function using the fusion quantization operator to obtain a processing result corresponding to the network layer.

7. The method of claim 6, wherein, performing acceleration processing on the quantized data based on the model weight by the acceleration processor using a fusion quantization operator to obtain a quantization result corresponding to the quantized data, comprising: monitoring a plurality of different memory access instructions to obtain the quantized data from a multi-level cache; monitoring a preset calling instruction to call the acceleration processor to perform acceleration processing on the quantized data based on the model weight using the fusion quantization operator to obtain the quantization result.

8. The method of claim 7, wherein, The acceleration processor performs acceleration processing on the quantized data by starting a plurality of thread blocks, wherein the indexes of thread blocks belonging to the same thread bundle are adjacent in the quantization result, the thread bundle is obtained by dividing at least one thread, and the thread is obtained by dividing the plurality of thread blocks; the thread blocks belonging to the same thread bundle read the same data block in the quantized data; and the thread blocks belonging to the same thread read the quantized data cached in the same register.

9. The method of claim 6, wherein, The fusion quantization operator comprises double buffering and multi-stage pipelining.

10. The method of claim 2, wherein, Triggering a deep learning model to be executed for acceleration processing, comprising: obtaining a model file of the deep learning model, wherein the model file is used to store original weights and a plurality of network layers; parsing the model file to obtain the deep learning model.

11. The method of claim 10, wherein, Parsing the model file to obtain the deep learning model, comprising: parsing the model file to obtain initial model information of the deep learning model, wherein the initial model information at least comprises: the original weights, and network structures of the network layers; performing graph optimization on the initial model information to obtain optimized model information; determining the running algorithm adopted by the network layer in the acceleration processor, the data format of input data input to the network layer, and the data format of output data output from the network layer in the optimized model information to obtain the deep learning model.

12. A data processing apparatus, characterized by comprising: a triggering module configured to monitor a quantization acceleration instruction and trigger a deep learning model to be executed for acceleration processing, wherein the deep learning model is a pre-trained neural network model; an obtaining module configured to obtain target data collected in a real environment; a dynamic quantization module, configured to perform dynamic quantization processing on the deep learning model based on the target data, to obtain target quantization parameters of the deep learning model; an acceleration network layer, configured to perform acceleration processing on the deep learning model by using an acceleration processor based on the target quantization parameters, to obtain a target result of the target data, wherein the target result is a dequantization result of a last network layer in a plurality of network layers included in the deep learning model, and the dequantization result is determined based on quantization data corresponding to the network layer and the acceleration processor, and the quantization data is obtained by quantizing input data of the network layer based on the target quantization parameters.

13. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a stored program, wherein the program, when executed, controls a device in which the computer readable storage medium is located to perform the method of any one of claims 1 to 11.

14. An electronic device, comprising: comprise: a memory storing an executable program; a processor configured to execute the program, wherein the program, when executed, performs the method of any one of claims 1 to 11.

Citation Information

Patent Citations

  • Deep learning model tuning method and computing device

    CN113139650A