A neural network acceleration system based on logarithmic block floating-point quantization

By designing a neural network acceleration system based on logarithmic block floating point quantization, the problem of low computing efficiency of general hardware architecture in the field is solved, and efficient end-to-end deployment of deep neural network models is achieved.

CN114626516BActive Publication Date: 2025-05-30NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210300275.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-24
Publication Date
2025-05-30
Estimated Expiration
2042-03-24

AI Technical Summary

Technical Problem

In the prior art, the computing efficiency of the field general hardware architecture is low, making it more difficult to deploy the deep neural network model end-to-end.

Method used

Design a neural network acceleration system based on logarithmic block floating point quantization, including compilers, runtimes and neural network accelerators. The compiler blocks the deployment model data and converts it into hardware instructions. The neural network accelerator uses tensor DMA, on-chip cache unit and computing unit to carry and quantize data, and finally performs neural network operations.

Benefits of technology

Through logarithmic block floating point quantization and adaptation of hardware architecture, computation redundancy is reduced, computing efficiency is improved, and end-to-end deployment of deep neural network models is supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114626516B_ABST
    Figure CN114626516B_ABST
Patent Text Reader

Abstract

The present application provides a neural network acceleration system based on logarithmic block floating-point quantization. The system includes a compiler, a runtime, and a neural network accelerator. When in use, the compiler divides the model data to be deployed according to the quantization block granularity, and converts all the model data to be deployed into hardware instructions. Through the runtime, it interacts with the neural network accelerator. The neural network accelerator, according to the instructions, transfers the data from off-chip in chunks to on-chip for loading according to the transfer block granularity, and performs logarithmic block floating-point quantization on each data quantization block. Finally, it executes the corresponding neural network operations using the quantization results. The entire system converts the model into instructions recognizable by the hardware through the compiler, sends instructions and data to the hardware by the runtime and communicates efficiently with the hardware. At the same time, it adopts a hardware architecture fully adapted to the logarithmic block floating-point quantization method, with less computational redundancy and higher computational efficiency, and can effectively support the end-to-end deployment of deep neural network models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a neural network acceleration system based on logarithmic block floating-point quantization. Background Art

[0002] Deep neural network models have currently been widely applied to various tasks, such as natural language processing, image processing, etc. However, deep neural network models contain a large number of layers, and most of the calculations are large-scale matrix multiplication operations and vector operations. Therefore, the amount of calculation and the number of parameters are extremely large, which will cause deep neural network models to occupy a large amount of storage space during operation and have high requirements for computing power. Edge devices are difficult to meet the above two requirements, thus restricting the deployment of deep neural network models on edge devices.

[0003] In order to achieve the deployment of deep neural network models on edge devices, currently, the method of performing logarithmic block floating-point quantization compression on the data of deep neural network models and using a general-purpose hardware architecture in the field to replace the CPU (or GPU) on the edge device to execute the inference process of the deep neural network model can be adopted.

[0004] Although this method can reduce the size of model parameters and has a low loss of quantization accuracy, due to the lack of a hardware architecture adapted to this quantization method, there are still many redundancies in the calculations of the general-purpose hardware architecture in the field, so the computing efficiency is low, and thus the end-to-end deployment of deep neural network models is still relatively difficult. Summary of the Invention

[0005] This application provides a neural network acceleration system based on logarithmic block floating-point quantization, which can be used to solve the technical problem that the computing efficiency of the general-purpose hardware architecture in the field in the prior art is low and the end-to-end deployment of deep neural network models is still relatively difficult.

[0006] To solve the above technical problem, the embodiments of this application disclose the following technical solutions:

[0007] A neural network acceleration system based on logarithmic block floating-point quantization includes a compiler, a runtime, and a neural network accelerator connected in sequence. The neural network accelerator includes a control unit, a conversion unit, and a tensor DMA, an on-chip cache unit, and a computing unit connected in sequence, where:

[0008] The compiler is configured to perform the following steps:

[0009] Partition the model data to be deployed according to a preset quantization block granularity to obtain a plurality of data quantization blocks. The model data to be deployed includes the weight values and current activation values of the model to be deployed, and the current activation values include the current input values and current output activation values;

[0010] Convert the to-be-deployed model into multiple hardware instructions recognizable by the neural network accelerator. The multiple hardware instructions include memory access instructions and computation instructions. The memory access instructions are used to instruct the tensor DMA to block-transfer each data quantization block from off-chip storage to the on-chip cache unit for loading and from the on-chip cache unit to off-chip storage for storage according to the transfer block granularity during runtime. The transfer block granularity is an integer multiple of the quantization block granularity. The computation instructions are used to instruct the control unit to allocate computation data and data conversion methods to the computation unit and the conversion unit;

[0011] The control unit is configured to perform the following steps:

[0012] Control the tensor DMA to block-transfer each data quantization block from off-chip storage to the on-chip cache unit for loading and from the on-chip cache unit to off-chip storage for storage according to the memory access instructions;

[0013] Control the conversion unit to perform logarithmic block floating-point quantization on each data quantization block in the on-chip cache unit according to the block floating-point shared exponent of each data quantization block. Among them, the block floating-point shared exponent of the weight value quantization block in each data quantization block is pre-determined by the compiler according to all weight elements in the weight value quantization block, and the block floating-point shared exponent of the current activation value quantization block in each data quantization block is pre-determined by the compiler offline according to all elements in the pre-acquired activation value sample set, or is determined online by the conversion unit according to all elements in the current activation value quantization block;

[0014] Control the computation unit to perform the computation of computation-intensive operators and the computation of memory access-intensive operators according to the logarithmic block floating-point quantization results of each data quantization block.

[0015] In one implementable manner, the quantization block granularity is set in the following way:

[0016] Determine the basic block granularity according to a preset maximum quantization error or a preset quantization signal-to-noise ratio;

[0017] Determine the quantization block granularity according to the basic block granularity and a preset block multiple.

[0018] In one implementable manner, the transfer block granularity is set in the following way:

[0019] Determine the total off-chip data transfer volume based on the number of transfers of the weight value transfer block, the number of transfers of the current input value transfer block, the number of transfers of the current output activation value transfer block, the on-chip storage occupied by the weight value transfer block, the on-chip storage occupied by the current input value transfer block, and the on-chip storage occupied by the current output activation value transfer block. The weight value transfer block is determined according to the weight value quantization block and the first integer multiple. The current input value transfer block is determined according to the current input value quantization block in each data quantization block and the second integer multiple. The current output activation value transfer block is determined according to the current output activation value quantization block in each data quantization block and the third integer multiple;

[0020] Search and determine the minimum total off-chip data transfer volume from each total off-chip data transfer volume according to the constraint condition that the on-chip storage occupied by the weight value transfer block, the current input value transfer block, and the current output activation value transfer block is less than or equal to their respective allowed total on-chip cache volumes;

[0021] Obtain the sizes of the target weight value transfer block, the target current input value transfer block, and the target current output activation value transfer block corresponding to the minimum total off-chip data transfer volume;

[0022] Determine the sizes of the target weight value transfer block, the target current input value transfer block, and the target current output activation value transfer block as the transfer block granularity.

[0023] In one implementable manner, the block floating-point shared exponent of the weight value quantization block in each data quantization block is determined as follows:

[0024] Convert each weight element in the weight value quantization block in each data quantization block into a floating-point form;

[0025] For each weight value quantization block, obtain the exponent value corresponding to the weight element with the largest absolute value;

[0026] Determine the exponent value corresponding to the weight element with the largest absolute value as the block floating-point shared exponent of the weight value quantization block.

[0027] In one implementable manner, the block floating-point shared exponent of the current activation value quantization block in each data quantization block is determined as follows:

[0028] Determine the quantization execution mode, where the quantization execution mode includes an offline quantization mode and an online quantization mode;

[0029] In the offline quantization mode, obtain the original probability distribution corresponding to all elements in the activation value sample set;

[0030] Obtain the respective quantized probability distributions corresponding to all elements in the activation value sample set under quantization block schemes with different sharing exponents;

[0031] Determine the KL divergence between the original probability distribution and each quantized probability distribution;

[0032] Determine the sharing exponent corresponding to the minimum KL divergence as the block floating-point sharing exponent for the current activation value quantization block;

[0033] Alternatively, in the online quantization mode, convert all elements in the current activation value quantization block into floating-point form;

[0034] For each current activation value quantization block, obtain the exponent value corresponding to the current activation value element with the largest absolute value;

[0035] Determine the exponent value corresponding to the current activation value element with the largest absolute value as the block floating-point sharing exponent for the current activation value quantization block.

[0036] In one implementable manner, the determining the KL divergence between the original probability distribution and each quantized probability distribution includes:

[0037] Determine the KL divergence between the original probability distribution and each quantized probability distribution through the following formula:

[0038]

[0039] Where KL(p||q) is the KL divergence between the original probability distribution and any quantized probability distribution, p(x) is the original probability distribution, and q(x) is any quantized probability distribution.

[0040] In one implementable manner, the performing logarithmic block floating-point quantization on each data quantization block in the on-chip cache unit according to the block floating-point sharing exponents of each data quantization block includes:

[0041] Determine the final block floating-point representation of each data quantization block in the on-chip cache unit according to the block floating-point sharing exponents of each data quantization block;

[0042] Convert the mantissa of each element in each data quantization block into a logarithmic representation.

[0043] In one implementable manner, the determining the final block floating-point representation of each data quantization block in the on-chip cache unit according to the block floating-point sharing exponents of each data quantization block includes:

[0044] Determine the final block floating-point representation of each data quantization block in the on-chip cache unit through the following formula:

[0045]

[0046] Among them, V b is any data quantization block in the on-chip cache unit, and v bi is the i-th element in the data quantization block, M bv is the data block composed of the mantissa representations of the elements in the data quantization block, ∈ v is the block floating-point shared exponent of the data quantization block, s i is the positive or negative sign of a single element in the data quantization block, m bi is the mantissa of a single element in the data quantization block.

[0047] In an implementable manner, the computing unit includes an input preprocessing module, an output postprocessing module, a logarithmic block floating-point matrix multiplication computing module, a vector computing module, and a memory access index generator. The logarithmic block floating-point matrix multiplication computing module and the vector computing module are respectively connected to the input preprocessing module, the output postprocessing module, and the memory access index generator;

[0048] The input preprocessing module is used to obtain the logarithmic block floating-point quantization results of each data quantization block online;

[0049] The logarithmic block floating-point matrix multiplication computing module is used to perform the calculation of the compute-intensive operator according to the calculation order generated by the memory access index generator based on the logarithmic block floating-point quantization results of each data quantization block, and output the result to the output postprocessing module;

[0050] The vector computing module is used to perform the calculation of the memory access-intensive operator according to the calculation order generated by the memory access index generator based on the logarithmic block floating-point quantization results of each data quantization block, and output the result to the output postprocessing module;

[0051] The output postprocessing module is used to complete the activation calculation and output the result to the on-chip cache unit.

[0052] In an implementable manner, the neural network accelerator further includes an interrupt control unit and a register bank connected to the runtime. The register bank includes a control register, a configuration register, an address register, and a status register;

[0053] The interrupt control unit is used to notify the runtime that the calculation is completed or an abnormal situation occurs;

[0054] The control register is used to control the neural network accelerator to start the calculation or perform a reset;

[0055] The configuration register is used to store the configurable function information of each module;

[0056] The address register is used to determine the base address of memory access;

[0057] The status register is used to count the running status of the neural network accelerator and send the statistical results to the runtime.

[0058] Thus, the neural network acceleration system based on logarithmic block floating-point quantization provided by the embodiments of the present application includes a compiler, a runtime, and a neural network accelerator. When in use, the compiler divides the model data to be deployed according to the quantization block granularity, and converts all the model data to be deployed into hardware instructions. Through the interaction between the runtime and the neural network accelerator, the neural network accelerator blocks and transports the data from off-chip storage to on-chip for loading according to the transfer block granularity, and performs logarithmic block floating-point quantization on each data quantization block. Finally, using the logarithmic block floating-point quantization results, the corresponding neural network operations are executed. The entire system converts the model into instructions recognizable by the hardware through the compiler, issues instructions and data to the hardware and communicates efficiently with the hardware by the runtime, and adopts a hardware architecture fully adapted to the logarithmic block floating-point quantization method, with less redundancy in calculations, so the calculation efficiency is relatively high, and it can effectively support the end-to-end deployment of deep neural network models. Description of the Drawings

[0059] Figure 1 It is a schematic structural diagram of a neural network acceleration system based on logarithmic block floating-point quantization provided by the embodiments of the present application;

[0060] Figure 2 It is a schematic diagram of the overall representation form of logarithmic block floating-point provided by the embodiments of the present application. Detailed Embodiments

[0061] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0062] To improve the calculation efficiency of the acceleration hardware architecture and achieve the end-to-end deployment of deep neural network models, the embodiments of the present application provide a neural network acceleration system based on logarithmic block floating-point quantization. The specific implementation mainly includes the hardware level (such as a neural network accelerator) and the software level (such as a model quantization algorithm, a compiler, a runtime, and an instruction set definition). Figure 1 Exemplarily shown is a schematic structural diagram of a neural network acceleration system based on logarithmic block floating-point quantization provided by the embodiments of the present application, as Figure 1As shown in the figure, the neural network acceleration system provided by the embodiment of the present application specifically includes a compiler 1, a runtime 2, and a neural network accelerator 3 that are connected in sequence. The neural network accelerator 3 includes a control unit 31, a conversion unit 32, and a tensor DMA (Direct Memory Access) 33, an on-chip cache unit 34, and a computing unit 35 that are connected in sequence.

[0063] The compiler 1 provided by the embodiment of the present application will be described below.

[0064] The compiler 1 is configured to execute the following Step 1 and Step 2:

[0065] Step 1: Block the model data to be deployed according to a preset quantization block granularity to obtain a plurality of data quantization blocks.

[0066] Among them, the model to be deployed is a deep neural network model, and the model data to be deployed includes the weight values and current activation values of the model to be deployed. The current activation values include the current input values and the current output activation values.

[0067] Specifically, the current input values are the current input values of each network layer in the model to be deployed. That is to say, the initial input of the model is also included in the current input values. The current output activation values are the current output activation values of each network layer in the model to be deployed.

[0068] In some embodiments, the quantization block granularity can be set in the following manner:

[0069] Determine the basic block granularity according to a preset maximum quantization error or a preset quantization signal-to-noise ratio.

[0070] Determine the quantization block granularity according to the basic block granularity and a preset block multiple.

[0071] Among them, the basic block granularity includes the weight value block granularity, the current input value block granularity, and the current output activation value block granularity. Correspondingly, the quantization block granularity includes the weight value quantization block granularity, the current input value quantization block granularity, and the current output activation value quantization block granularity. The block multiples corresponding to each block granularity can be the same or different.

[0072] In this way, by using the above-mentioned blocking method, the impact of quantization on the model accuracy can be reduced as much as possible, and it is ensured that the on-chip cache loading and storage time of the compute-intensive operator is shorter than the computing time, while the on-chip cache utilization rate of the memory-intensive operator is as high as possible, and as large a block as possible is maintained to maximize the external storage access bandwidth efficiency.

[0073] In other possible embodiments, the quantization block granularity can also be arbitrarily specified. For example, the weight value block granularity is 3×3, the current input value block granularity is 4×4, and the current output activation value block granularity is 2×2, which is not specifically limited.

[0074] After the compiler 1 chunks the data of the model to be deployed, it quantizes each quantized data chunk. The quantization method uses logarithmic block floating-point quantization. The compiler 1 first needs to determine the block floating-point shared exponent of each quantized data chunk.

[0075] For the weight value quantization chunks in each quantized data chunk, the block floating-point shared exponent is determined through the following method:

[0076] First, convert each weight element in the weight value quantization chunks of each quantized data chunk into floating-point form.

[0077] A regular floating-point number consists of a sign bit, an exponent bit, and a mantissa bit.

[0078] Specifically, the floating-point form of all weight elements in the weight value quantization chunk can be represented by formula (1):

[0079]

[0080] In formula (1), V is the weight value quantization chunk, v i is the i-th weight element in the weight value quantization chunk, N is the number of all weight elements, s i is the positive or negative sign of the i-th weight element, m i is the mantissa of the i-th weight element, e i is the exponent value of the i-th weight element.

[0081] Then, for each weight value quantization chunk, obtain the exponent value corresponding to the weight element with the largest absolute value.

[0082] Finally, determine the exponent value corresponding to the weight element with the largest absolute value as the block floating-point shared exponent of the weight value quantization chunk.

[0083] For the current activation value quantization chunks in each quantized data chunk, the block floating-point shared exponent is determined through the following steps:

[0084] The first step is to determine the execution mode of quantization. Among them, the execution mode of quantization includes an offline quantization mode and an online quantization mode. If the execution mode of quantization is the offline quantization mode, the compiler 1 continues to execute the second step. If the execution mode of quantization is the online quantization mode, it is transferred to the conversion unit 32 to execute the sixth step.

[0085] The second step is to obtain the original probability distribution corresponding to all elements in the activation value sample set in the offline quantization mode.

[0086] Among them, the activation value sample set consists of the initial input value and the output activation values of each layer in the forward propagation process, and can be obtained from the pre-stored sample set. The original probability distribution is obtained by histogram statistics based on the data in the activation value sample set.

[0087] In particular, logarithmic block floating-point quantization does not require a training process, that is, it does not require backpropagation to optimize and adjust quantization parameters. Only the maximum value, minimum value, and histogram statistics in the forward propagation process need to be collected to determine the exponential value e of the maximum absolute value. max 。

[0088] Step 3: Obtain the respective quantized probability distributions corresponding to all elements in the activation value sample set under the quantization block scheme with different shared exponents.

[0089] Among them, each shared exponent can be searched within a range less than e max . The quantized probability distribution is obtained by histogram statistics based on the quantized data.

[0090] Step 4: Determine the KL divergence between the original probability distribution and each quantized probability distribution.

[0091] Specifically, the KL divergence between the original probability distribution and each quantized probability distribution can be determined by formula (2):

[0092]

[0093] In formula (2), KL(p||q) is the KL divergence between the original probability distribution and any quantized probability distribution, p(x) is the original probability distribution, and q(x) is any quantized probability distribution.

[0094] Step 5: Determine the shared exponent corresponding to the minimum KL divergence as the block floating-point shared exponent of the current activation value quantization block.

[0095] It should be noted that the first step to the fifth step of determining the block floating-point shared exponent of the current activation value quantization block are executed by the compiler 1.

[0096] Step 6: In the online quantization mode, convert all elements in the current activation value quantization block into floating-point form.

[0097] Step 7: For each current activation value quantization block, obtain the exponential value corresponding to the current activation value element with the largest absolute value.

[0098] Step 8: Determine the exponential value corresponding to the current activation value element with the largest absolute value as the block floating-point shared exponent of the current activation value quantization block.

[0099] The solutions corresponding to the sixth to eighth steps are the same as the solutions for determining the block floating-point shared exponent of the weighted value quantization block, except that the processing object changes from all the weighted elements in the weighted value quantization block to all the elements in the current activation value quantization block. For the specific solution, please refer to the solution for determining the block floating-point shared exponent of the weighted value quantization block, which will not be elaborated here.

[0100] It should be noted that the sixth to eighth steps for determining the block floating-point shared exponent of the current activation value quantization block are executed by the conversion unit 32.

[0101] It should also be noted that since the current activation value quantization block includes the current input value quantization block and the current output activation value quantization block, determining the block floating-point shared exponent of the current activation value quantization block actually means determining the block floating-point shared exponents of the current input value quantization block and the current output activation value quantization block respectively. Since the methods for determining the block floating-point shared exponents of the current input value quantization block and the current output activation value quantization block are the same, they are not described separately but are introduced together as determining the block floating-point shared exponent of the current activation value quantization block. That is to say, the method for determining the block floating-point shared exponent of the current activation value quantization block can be applied to determining the block floating-point shared exponent of the current input value quantization block and to determining the block floating-point shared exponent of the current output activation value quantization block.

[0102] Step 2: Convert the model to be deployed into multiple hardware instructions that can be recognized by the neural network accelerator 3.

[0103] Among them, the multiple hardware instructions include memory access instructions and computing instructions. The memory access instructions are used to instruct the tensor DMA 33 to block-transfer each data quantization block from off-chip storage to the on-chip cache unit 34 for loading and from the on-chip cache unit 34 to off-chip storage for storage according to the transfer block granularity during runtime 2. The transfer block granularity is an integer multiple of the quantization block granularity. The computing instructions are used to instruct the control unit 31 to allocate computing data and data conversion methods to the computing unit 35 and the conversion unit 32.

[0104] Specifically, the model to be deployed is a computational graph represented in the form of a directed acyclic graph. Each node in this computational graph is an individual operator, such as a convolution, fully connected layer, Softmax layer, etc. The object to be converted into hardware instructions also includes the calculation between the block floating-point shared exponents of the data quantization blocks corresponding to logarithmic block floating-point quantization. All the parameter loading and storage processes, activation value loading and storage processes, calculations involved in the quantization process, and other calculation operations that need to instruct the subsequent neural network accelerator 3 to execute in the model to be deployed will be converted into corresponding hardware instructions.

[0105] Hardware instructions refer to binary instructions that can be recognized by a neural network accelerator, such as addition, matrix multiplication, etc.

[0106] In the implementation of compiler 1, first, a computational graph is represented based on a high-level intermediate representation, where each node in this intermediate representation is a separate operator. The high-level intermediate representation will ultimately be lowered to a low-level intermediate representation. The low-level intermediate representation is close to hardware instructions and is mainly divided into memory access instructions and computational instructions. Among them, the memory access instructions specify the source address and destination address of data transfer, as well as the data length. The computational instructions specify the type of computational operation, as well as the source operand sources and the result storage locations. Operator fusion and quantization parameter calculation are performed based on the high-level intermediate representation. In addition, the calculation results of quantization parameters are implemented by the low-level intermediate representation.

[0107] In some embodiments, the data transfer block granularity can be set in the following manner:

[0108] In the first step, based on the transfer times of the weight value transfer block, the transfer times of the current input value transfer block, the transfer times of the current output activation value transfer block, the on-chip storage occupied by the weight value transfer block, the on-chip storage occupied by the current input value transfer block, and the on-chip storage occupied by the current output activation value transfer block, the total off-chip data transfer volume is determined.

[0109] Among them, the transfer times of the weight value transfer block are determined according to the reuse times required by the weight values of the model to be deployed and the size of the weight value transfer block. The transfer times of the current input value transfer block are determined according to the reuse times required by the current input values of the model to be deployed and the size of the current input value transfer block. The transfer times of the current output activation value transfer block are determined according to the reuse times required by the current output activation values of the model to be deployed and the size of the current output activation value transfer block.

[0110] The weight value transfer block is determined according to the weight value quantization block and the first integer multiple. The current input value transfer block is determined according to the current input value quantization block in each data quantization block and the second integer multiple. The current output activation value transfer block is determined according to the current output activation value quantization block in each data quantization block and the third integer multiple.

[0111] The weight value quantization block is obtained by partitioning the weight values of the model to be deployed according to the weight value block granularity. The current input value quantization block is obtained by partitioning the current input values of the model to be deployed according to the current input value block granularity. The current output activation value quantization block is obtained by partitioning the current output activation values of the model to be deployed according to the current output activation value block granularity.

[0112] The first integer multiple, the second integer multiple, and the third integer multiple can be the same or different. That is to say, the block granularity of the transfer block can be constructed or merged based on the quantization block granularity, but splitting into smaller block granularities is not allowed. Therefore, the final transfer block needs to be an integer multiple of the quantization block.

[0113] Specifically, the total off-chip data transfer volume can be determined by formula (3):

[0114] M total = M w + M i + M a = N w S w + N i S i + N a S a Formula (3)

[0115] In formula (3), M total is the total off-chip data transfer volume, M w , M i , M a are respectively the off-chip data transfer volume of the weight value, the off-chip data transfer volume of the current input value, and the off-chip data transfer volume of the current output activation value. S w , S i , S a are respectively the on-chip storage occupied by the weight value transfer block, the on-chip storage occupied by the current input value transfer block, and the on-chip storage occupied by the current output activation value transfer block. N w , N i , N a are respectively the transfer times of the weight value transfer block, the transfer times of the current input value transfer block, and the transfer times of the current output activation value transfer block.

[0116] Second step, search and determine the minimum total off-chip data transfer volume from each total off-chip data transfer volume according to the constraint conditions.

[0117] Among them, the constraint condition is that the on-chip storage occupied by the weight value transfer block, the current input value transfer block, and the current output activation value transfer block is less than or equal to their respective allowed total on-chip caches.

[0118] Specifically, the minimum total off-chip data transfer volume can be determined by formula (4)

[0119] min M total = min M w + M i + M a = min N w S w + Ni S i +N a S a , formula (4)

[0120] S w ≤C w , S i ≤C i , S a ≤C a

[0121] In formula (4), M total is the total off-chip data transfer volume, M w , M i , M a are respectively the off-chip data transfer volumes of the weight value, the off-chip data transfer volume of the current input value, and the off-chip data transfer volume of the current output activation value. S w , S i , S a are respectively the on-chip storage occupied by the weight value transfer block, the on-chip storage occupied by the current input value transfer block, and the on-chip storage occupied by the current output activation value transfer block. N w , N i , N a are respectively the transfer times of the weight value transfer block, the transfer times of the current input value transfer block, and the transfer times of the current output activation value transfer block. C w , C i , C a are respectively the allowable total on-chip cache corresponding to the weight value transfer block, the allowable total on-chip cache corresponding to the current input value transfer block, and the allowable total on-chip cache corresponding to the current output activation value transfer block.

[0122] The search method can be carried out by means of traversal search, greedy search or heuristic search, and is not specifically limited.

[0123] Step 3: Obtain the sizes of the target weight value transfer block, the target current input value transfer block, and the target current output activation value transfer block corresponding to the minimum total off-chip data transfer volume.

[0124] That is to say, obtain the sizes of the target weight value transfer block, the target current input value transfer block, and the target current output activation value transfer block when the value on the left side of formula (4) is the smallest.

[0125] Step 4: Determine the sizes of the target weight value transfer block, the target current input value transfer block, and the target current output activation value transfer block as the transfer block granularity.

[0126] The method for determining the data transfer block granularity in the above embodiments takes into account the on-chip storage space size of the accelerator and minimizes off-chip storage data transfer, enabling effective overlap of memory access and computing, hiding memory access latency, and thus achieving a relatively high data transfer efficiency.

[0127] In other possible embodiments, the data transfer block granularity can also be arbitrarily specified without specific limitation.

[0128] After the compiler 1 converts the model to be deployed into multiple hardware instructions recognizable by the neural network accelerator 3, these hardware instructions are packaged in a loadable file.

[0129] Using the above compiler to perform quantization-oriented fusion on operators can not only reduce the number of operators, thereby reducing additional DDR (Double Data Rate Synchronous Dynamic Random Access Memory) access and redundant computations, ensuring that as much computation as possible depends on the on-chip cache, but also reduce the number of quantization-sensitive layers, converting more layers into efficient log-block floating-point computations. Since quantization-oriented operator fusion integrates linear computation layers into computation-intensive layers, it avoids the loss of computation accuracy after quantization of high-precision linear layer systems. In addition, the compiler can also complete the conversion process from the model to hardware instructions end-to-end and form an instruction sequence that can interact with the hardware.

[0130] The following describes the runtime 2 provided by the embodiments of the present application.

[0131] The runtime 2 can parse the instruction sequence and storage partition in the loadable file generated by the compiler 1, complete communication and data transfer with the neural network accelerator 3, and offload tasks to the neural network accelerator 3. Therefore, the runtime 2 is responsible for interacting with the hardware. In the scheduling and planning of the runtime 2, the neural network accelerator 3 supports double buffering, which allows a part to be used for data interaction with external storage and another part to be used for the execution of operations by each computing module inside the neural network accelerator 3, achieving overlap of computing and storage to hide memory access latency.

[0132] Specifically, the runtime 2 is used to send computation instructions and memory access instructions to the control unit 31 through a PCIe (Peripheral Component Interconnect Express, high-speed serial computer expansion bus standard) interface.

[0133] The following describes the neural network accelerator 3 provided by the embodiments of the present application.

[0134] The neural network accelerator 3 includes a control unit 31, a conversion unit 32, and a tensor DMA 33, an on-chip cache unit 34, and a computing unit 35 connected in sequence.

[0135] Specifically, the control unit 31 includes an instruction storage module 311, an instruction decoding module 312, an instruction issuing module 313, and an execution control module 314 that are connected in sequence. Among them, the instruction storage module 311 is connected to the runtime 2 through a PCIe interface and is used to specifically store the on-chip instructions sent to the neural network accelerator 3. The instruction decoding module 312 is responsible for decoding the on-chip instructions and transmitting the decoded information to the execution control module 314 through the instruction issuing module 313. The execution control module 314 provides control signals for the remaining units and modules in the neural network accelerator 3 to maintain the data dependency relationships between these units and modules.

[0136] Specifically, the control unit 31 is configured to execute the following steps 1 to 3:

[0137] Step 1, control the tensor DMA 33 to block-transfer each data quantization block from off-chip storage to the on-chip cache unit 34 for loading according to the memory access instruction, and transfer it from the on-chip cache unit 34 to off-chip storage for storage, where the block-transfer granularity has been described above and will not be elaborated here.

[0138] Among them, the block-transfer granularity has been described above and will not be elaborated here.

[0139] Specifically, when the tensor DMA 33 transfers data to the on-chip cache unit 34, it can access and read a view of the complete tensor and arrange it in the on-chip cache as required. Therefore, the computing unit 35 can directly read the multi-dimensional tensor after being divided into blocks and arranged in order. The tensor DMA 33 can also complete the filling of constant values in a certain dimension of the tensor during the data transfer process. The calculation results of the computing unit 35 are also stored in the on-chip cache unit 34 and are transferred by the tensor DMA 33 to external storage or after transforming the storage arrangement for use by other computing units.

[0140] The on-chip cache unit 34 can be divided into three parts, which are respectively responsible for the loading and storage of weight values, current activation values, and intermediate data or calculation results.

[0141] Step 2, control the conversion unit 32 to perform logarithmic block floating-point quantization on each data quantization block in the on-chip cache unit 34 according to the block floating-point shared exponent of each data quantization block.

[0142] Among them, the block floating-point shared exponent of the weight value quantization block in each data quantization block is pre-determined by the compiler 1 according to all weight elements in the weight value quantization block. The block floating-point shared exponent of the current activation value quantization block in each data quantization block is determined offline by the compiler 1 according to all elements in the pre-acquired activation value sample set, or is determined online by the conversion unit 32 according to all elements in the current activation value quantization block.

[0143] The determination of the block floating-point shared exponent for each data quantization block has been described above and will not be elaborated here.

[0144] Figure 2 An exemplary schematic diagram of the overall representation form of the logarithmic block floating-point provided by the embodiment of the present application is shown, as Figure 2 shown. The logarithmic block floating-point consists of four parts, namely the shared exponent bit, the sign bit, the exponent difference bit, and the logarithmic mantissa bit. Among them, the shared exponent of the shared exponent bit represents the common exponent within the entire data block and is the same for each element in the block. The sign of the sign bit represents the positive or negative of each element in the block, the exponent difference of the exponent difference bit represents the exponent difference between the actual exponent of each element in the block and the shared exponent, and the logarithmic mantissa of the logarithmic mantissa bit is the conversion of the mantissa bit in the conventional floating-point number to the representation in the logarithmic domain. Figure 2 The logarithmic block floating-point shown in

[0145] contains four elements.

[0146] Specifically, the conversion unit 32 performs logarithmic block floating-point quantization on each data quantization block in the on-chip cache unit 34 according to the block floating-point shared exponent of each data quantization block, which can be specifically implemented through the following two steps:

[0147] In the first step, according to the block floating-point shared exponent of each data quantization block, determine the final block floating-point representation of each data quantization block in the on-chip cache unit 34.

[0148]

[0149] In formula (5), V b is any data quantization block in the on-chip cache unit 34, v bi is the i-th element in the data quantization block, M bv is the data block composed of the mantissa representations of each element in the data quantization block, ∈ v is the block floating-point shared exponent of the data quantization block, s i is the positive or negative sign of a single element in the data quantization block, m bi is the mantissa of a single element in the data quantization block. The mantissa of each element can be obtained by shifting through the exponent difference d i =∈ V -e i shift.

[0150] Using the above block floating-point representation, the original exponent representation space is reduced to Moreover, the mantissa bit width is compressed. Assuming that the bit width of m N before compression is b1 , after compression has a bit width of b 2 , thus reducing the mantissa representation space to In addition, complex floating-point operations are also converted to fixed-point operations, saving the processes of partial exponent calculation, alignment, normalization, rounding, and denormalization calculation in floating-point operations, and reducing the calculation overhead.

[0151] In the second step, the mantissa of each element in each data quantization block is converted to a logarithmic representation.

[0152] Exemplarily, the mantissa is converted to a logarithmic representation, which is

[0153] In this way, the mantissa multiplication operation between two block floating-point numbers can be converted to an addition operation in the logarithmic domain, that is Further reducing the calculation overhead.

[0154] In addition, the conversion unit 32 is also used to convert between the logarithmic block floating-point representation and the conventional floating-point representation, or to unify different shared exponent blocks to complete the quantization process, and to complete the conversion of the data quantization block size between the input and the output. For the conversion of the data quantization block size between the input and the output, exemplarily, the size of the current input value quantization block is 2×2, and the size of the current output activation value quantization block is 1×1. The current output activation value of any network layer is the current input value of the next network layer. Therefore, a conversion between 1×1 and 2×2 is required.

[0155] In this way, adopting the above logarithmic block floating-point quantization method can have both the high dynamic representation range of floating-point numbers and the low calculation complexity of fixed-point numbers, and convert multiplication to addition operations, further reducing the calculation overhead.

[0156] In step three, the control calculation unit 35 executes the calculation of the computation-intensive operator and the calculation of the memory access-intensive operator according to the logarithmic block floating-point quantization results of each data quantization block.

[0157] The calculation unit 35 includes an input preprocessing module 351, an output postprocessing module 352, a logarithmic block floating-point matrix multiplication calculation module 353, a vector calculation module 354, and a memory access index generator 355. The logarithmic block floating-point matrix multiplication calculation module 353 and the vector calculation module 354 are respectively connected to the input preprocessing module 351, the output postprocessing module 352, and the memory access index generator 355. Among them:

[0158] The input preprocessing module 351 is used to obtain the logarithmic block floating-point quantization results of each data quantization block online.

[0159] The logarithmic block floating-point matrix multiplication calculation module 353 is used to perform the calculation of computationally intensive operators according to the calculation order generated by the memory access index generator 355, based on the logarithmic block floating-point quantization results of each data quantization block, and output the results to the output post-processing module 352.

[0160] Among them, computationally intensive means that the calculation speed is slower than the data supply speed.

[0161] Specifically, the logarithmic block floating-point matrix multiplication calculation module 353 is responsible for the calculation of matrix multiplication with intensive computational workloads in the model to be deployed. Among them, multiplication is calculated by addition in the logarithmic domain, and there is a two-level accumulation mechanism. Data blocks sharing the same exponent use fixed-point accumulation, and floating-point accumulation is used when overflow occurs or the shared exponent changes.

[0162] The vector calculation module 354 is used to perform the calculation of memory access intensive operators according to the calculation order generated by the memory access index generator 355, based on the logarithmic block floating-point quantization results of each data quantization block, and output the results to the output post-processing module 352.

[0163] Among them, memory access intensive means that the calculation speed is faster than the data supply speed.

[0164] Specifically, the vector calculation module 354 is responsible for the calculation of memory access intensive operators, such as activation functions, batch normalization layers, etc. The calculation orders of both the logarithmic block floating-point matrix multiplication calculation module 353 and the vector calculation module 354 are generated by the memory access index generator 355.

[0165] The output post-processing module 352 is used to complete the activation calculation and output the results to the on-chip cache unit 34.

[0166] Among them, the input pre-processing module 351 and the output post-processing module 352 perform dynamic collection of quantization statistics to online determine the shared exponent, and handshake with the conversion unit 32 in a timely manner to complete the unification of the shared exponent. In addition, the output post-processing module 352 also completes simple activation function calculations, such as ReLU.

[0167] In addition, the neural network accelerator 3 also includes an interrupt control unit 36 and a register bank 37 connected to the runtime 2. The register bank 37 includes control registers, configuration registers, address registers, and status registers.

[0168] The interrupt control unit 36 is used to notify the runtime 2 that the calculation is completed or an abnormal situation has occurred.

[0169] Among them, when notifying the runtime 2 that an abnormal situation has occurred, the runtime 2 will read the interrupt value to judge the specific abnormal event.

[0170] A control register for controlling the neural network accelerator 3 to start calculations or perform a reset.

[0171] A configuration register for storing configurable function information of each module.

[0172] An address register for determining the base address of memory access.

[0173] A status register for counting the operating status of the neural network accelerator 3 and sending the statistical results to the runtime 2.

[0174] With the above neural network accelerator, the quantization process is completely based on on-chip cache data for online conversion, and the calculations are relatively continuous. Moreover, the quantization process, including the conversion from floating-point to logarithmic block floating-point, the conversion from logarithmic block floating-point to floating-point, and the conversion between logarithmic block floating-point numbers with different shared exponents, is completed within the neural network accelerator without the need for other computing devices such as a CPU to assist in the quantization calculation, reducing the CPU's participation in the calculation process. Therefore, the calculation efficiency is greatly improved. In addition, the architecture precisely adapts to the logarithmic block floating-point quantization method with less calculation redundancy, further improving the calculation efficiency of the neural network and being able to effectively support the deployment of deep neural network models on edge devices.

[0175] Thus, the neural network acceleration system based on logarithmic block floating-point quantization provided by the embodiments of the present application includes a compiler, a runtime, and a neural network accelerator. When in use, the compiler divides the model data to be deployed according to the quantization block granularity and converts all the model data to be deployed into hardware instructions, interacts with the neural network accelerator through the runtime. The neural network accelerator blocks and transports the data from off-chip storage to on-chip for loading according to the transfer block granularity, and performs logarithmic block floating-point quantization on each data quantization block. Finally, according to the logarithmic block floating-point quantization results of the data quantization blocks, the corresponding neural network operations are executed. The entire neural network acceleration system converts the model into instructions recognizable by the hardware through the compiler, issues instructions and data to the hardware and communicates with the hardware efficiently through the runtime, and adopts a hardware architecture that is fully adapted to the logarithmic block floating-point quantization method with less redundancy in calculations, so the calculation efficiency is relatively high, and it can effectively support the end-to-end deployment of deep neural network models.

[0176] The above has described the present application in detail in combination with specific implementation manners and exemplary examples, but these descriptions should not be construed as limiting the present application. Those skilled in the art understand that without departing from the spirit and scope of the present application, various equivalent replacements, modifications, or improvements can be made to the technical solutions and their implementation manners of the present application, and these all fall within the scope of the present application. The protection scope of the present application is subject to the appended claims.

Claims

1. A neural network acceleration system based on logarithmic block floating-point quantization, characterized in that, it includes a compiler, a runtime, and a neural network accelerator connected in sequence. The neural network accelerator includes a control unit, a conversion unit, and a tensor DMA, an on-chip cache unit, and a computing unit connected in sequence, where: The compiler is configured to perform the following steps: Chunk the model data to be deployed according to a preset quantization chunk granularity to obtain multiple data quantization chunks. The model data to be deployed includes the weight values and current activation values of the model to be deployed. The current activation values include the current input values and current output activation values; Convert the model to be deployed into multiple hardware instructions recognizable by the neural network accelerator. The multiple hardware instructions include memory access instructions and computing instructions. The memory access instructions are used to instruct the tensor DMA to chunk and transfer each data quantization chunk from off-chip storage to the on-chip cache unit for loading according to the transfer chunk granularity through the runtime, and to transfer from the on-chip cache unit to off-chip storage for storage. The transfer chunk granularity is an integer multiple of the quantization chunk granularity. The computing instructions are used to instruct the control unit to allocate computing data and data conversion methods to the computing unit and the conversion unit; The control unit is configured to perform the following steps: Control the tensor DMA to chunk and transfer each data quantization chunk from off-chip storage to the on-chip cache unit for loading according to the memory access instructions, and to transfer from the on-chip cache unit to off-chip storage for storage according to the transfer chunk granularity; Control the conversion unit to perform logarithmic block floating-point quantization on each data quantization chunk in the on-chip cache unit according to the block floating-point shared exponent of each data quantization chunk. Among them, the block floating-point shared exponent of the weight value quantization chunk in each data quantization chunk is pre-determined by the compiler according to all weight elements in the weight value quantization chunk. The block floating-point shared exponent of the current activation value quantization chunk in each data quantization chunk is pre-determined by the compiler offline according to all elements in the pre-acquired activation value sample set, or is determined online by the conversion unit according to all elements in the current activation value quantization chunk; Control the computing unit to perform calculations of compute-intensive operators and memory access-intensive operators according to the logarithmic block floating-point quantization results of each data quantization chunk; The transfer chunk granularity is set in the following way: Determine the total off-chip data transfer volume according to the number of transfer times of the weight value transfer block, the number of transfer times of the current input value transfer block, the number of transfer times of the current output activation value transfer block, the on-chip storage occupied by the weight value transfer block, the on-chip storage occupied by the current input value transfer block, and the on-chip storage occupied by the current output activation value transfer block. The weight value transfer block is determined according to the weight value quantization block and the first integer multiple. The current input value transfer block is determined according to the current input value quantization block in each data quantization block and the second integer multiple. The current output activation value transfer block is determined according to the current output activation value quantization block in each data quantization block and the third integer multiple; Search and determine the minimum total off-chip data transfer volume from each total off-chip data transfer volume according to the constraint condition that the on-chip storage occupied by the weight value transfer block, the current input value transfer block, and the current output activation value transfer block is less than or equal to their respective corresponding allowable total on-chip cache volumes; Obtain the sizes of the target weight value transfer block, the target current input value transfer block, and the target current output activation value transfer block corresponding to the minimum total off-chip data transfer volume; Determine the sizes of the target weight value transfer block, the target current input value transfer block, and the target current output activation value transfer block as the transfer block granularity.

2. The neural network acceleration system according to claim 1, wherein, the quantization block granularity is set by the following method: Determine the basic block granularity according to a preset maximum quantization error or a preset quantization signal-to-noise ratio; Determine the quantization block granularity according to the basic block granularity and a preset block multiple.

3. The neural network acceleration system according to claim 1, wherein, the block floating-point shared exponent of the weight value quantization block in each data quantization block is determined by the following method: Convert each weight element in the weight value quantization block in each data quantization block into a floating-point form; For each weight value quantization block, obtain the exponent value corresponding to the weight element with the largest absolute value; Determine the exponent value corresponding to the weight element with the largest absolute value as the block floating-point shared exponent of the weight value quantization block.

4. The neural network acceleration system according to claim 1, wherein, the block floating-point shared exponent of the current activation value quantization block in each data quantization block is determined by the following method: Determine the quantization execution mode, and the quantization execution mode includes an offline quantization mode and an online quantization mode; In the offline quantization mode, obtain the original probability distribution corresponding to all elements in the activation value sample set; Obtain the respective quantized probability distributions corresponding to all elements in the activation value sample set under the quantization block schemes with different shared exponents; Determine the KL divergence between the original probability distribution and each quantized probability distribution; Determine the shared exponent corresponding to the minimum KL divergence as the block floating-point shared exponent of the current activation value quantization block; Alternatively, in the online quantization mode, convert all elements in the current activation value quantization block into floating-point form; For each current activation value quantization block, obtain the exponent value corresponding to the current activation value element with the largest absolute value; Determine the exponent value corresponding to the current activation value element with the largest absolute value as the block floating-point shared exponent of the current activation value quantization block.

5. The neural network acceleration system according to claim 4, wherein, the determination of the KL divergence between the original probability distribution and each quantized probability distribution includes: determine the KL divergence between the original probability distribution and each quantized probability distribution through the following formula: where KL(p||q) is the KL divergence between the original probability distribution and any quantized probability distribution, p(x) is the original probability distribution, and q(x) is any quantized probability distribution.

6. The neural network acceleration system according to claim 1, wherein, the logarithmic block floating-point quantization of each data quantization block in the on-chip cache unit according to the block floating-point shared exponent of each data quantization block includes: determine the final block floating-point representation of each data quantization block in the on-chip cache unit according to the block floating-point shared exponent of each data quantization block; convert the mantissa of each element in each data quantization block into a logarithmic representation.

7. The neural network acceleration system according to claim 6, wherein, the determination of the final block floating-point representation of each data quantization block in the on-chip cache unit according to the block floating-point shared exponent of each data quantization block includes: determine the final block floating-point representation of each data quantization block in the on-chip cache unit through the following formula: Among them, V b is any data quantization block in the on-chip cache unit, v bi is the i-th element in the data quantization block, M bv is the data block composed of the mantissa representations of the elements in the data quantization block, ∈ v is the block floating-point shared exponent of the data quantization block, s i is the positive or negative sign of a single element in the data quantization block, m bi is the mantissa of a single element in the data quantization block.

8. The neural network acceleration system according to claim 1, wherein, the calculation unit includes an input preprocessing module, an output postprocessing module, a logarithmic block floating-point matrix multiplication calculation module, a vector calculation module, and a memory access index generator, and both the logarithmic block floating-point matrix multiplication calculation module and the vector calculation module are respectively connected to the input preprocessing module, the output postprocessing module, and the memory access index generator; the input preprocessing module is used to obtain the logarithmic block floating-point quantization results of each data quantization block online; the logarithmic block floating-point matrix multiplication calculation module is used to execute the calculation of the computation-intensive operator according to the logarithmic block floating-point quantization results of each data quantization block in the calculation order generated by the memory access index generator, and output the result to the output postprocessing module; the vector calculation module is used to execute the calculation of the memory access-intensive operator according to the logarithmic block floating-point quantization results of each data quantization block in the calculation order generated by the memory access index generator, and output the result to the output postprocessing module; the output postprocessing module is used to complete the activation calculation and output the result to the on-chip cache unit.

9. The neural network acceleration system according to claim 1, wherein, The neural network accelerator further includes an interrupt control unit and a register bank connected to the runtime, and the register bank includes a control register, a configuration register, an address register, and a status register; The interrupt control unit is used to notify the runtime that the calculation is completed or an abnormal situation occurs; The control register is used to control the neural network accelerator to start calculation or perform a reset; The configuration register is used to store the configurable function information of each module; The address register is used to determine the base address of memory access; The status register is used to count the running status of the neural network accelerator and send the statistical result to the runtime.

Citation Information

Patent Citations

  • Hardware neural network conversion method, computing device, compiling method and neural network software and hardware collaboration system

    CN106650922A

  • Neural network compiler architecture and compiling method

    CN110766147A