Memristor accelerator for deploying quantitative neural network and quantitative neural network deployment method
By introducing additional zero columns into the memristor array and designing complete weighted hardware, the problem of degradation of inference accuracy during deployment of quantized neural networks in the prior art is solved, achieving higher inference accuracy and lower accuracy losses.
Patent Information
- Application Number
- CN202510225854.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Existing memristor accelerators lead to a decrease inference accuracy when deploying quantized neural networks, especially in low-precision quantized neural networks, where the accuracy loss is significant when the data bit width is low.
By introducing additional zero columns into the memristor array, asymmetric quantization is supported and quantized activated zero calculations are fused into the bias calculations, reducing accuracy losses during data conversion. At the same time, weighting hardware including shifters, adders and limiting hardware was designed to achieve complete weighting operations and reduce the accuracy loss of the weighting process.
The inference accuracy of the quantized neural network deployed in the memristor accelerator is improved, especially in low-precision scenarios, where the inference accuracy is improved more significantly, reducing the accuracy loss in data mapping and weighting.
Smart Images

Figure CN120087431A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of information storage and computer architecture, and more particularly, to a memristor accelerator for deploying a quantized neural network and a method for deploying a quantized neural network. Background Art
[0002] The forward propagation in the neural network inference process is calculated as follows:
[0003] a l = σ(W l a l-1 + b l )
[0004] Wherein, a l-1 represents the activation vector of the (l-1)-th layer, a l , W l and b l represent the activation vector, weight matrix and bias vector of the l-th layer respectively, and σ(·) is an activation function, which is a non-linear function. The key operation of this calculation (i.e., W l a l-1 ) is matrix-vector multiplication (MVM).
[0005] A memristor crossbar array (XB) has the ability to perform in-situ matrix-vector multiplication in the analog domain with high energy efficiency and high speed. A memristor-based accelerator has been proposed to accelerate neural networks. In such an accelerator, digital-to-analog converters (DACs) and analog-to-digital converters (ADCs) are used to complete the conversion between digital signals and analog signals. Due to the limited ADC / DAC accuracy and the number of bits of memristor cells, vector elements usually need multiple cycles for input, and matrix elements also need to be stored in multiple memristor cells located in multiple columns. Therefore, a shift-and-add (S&A) circuit is used to combine the calculation results from different input cycles and different columns.
[0006] A quantized neural network refers to a neural network whose parameters have been quantized. In order to use a memristor-based accelerator to accelerate a neural network, the corresponding quantized neural network needs to be deployed into the memristor-based accelerator, and the entire deployment process is as Figure 1 shown. Each layer of the neural network performs the three steps in the dashed box. The first step is the MVM of quantized weights and quantized activations (the first layer is quantized input), and this matrix-vector multiplication is mapped to a processing element (PE) for execution The quantized weights come from the offline quantization of pre-trained weights and remain unchanged during runtime. Except for the first layer, the quantized activations come from the re-quantization (referred to as weight quantization) of intermediate data during runtime. In particular, the quantized input of the first layer is obtained through offline quantization before participating in the MVM. In addition, the calculation results from multiple PEs are merged by the adder within the tile. During the MVM process, the data bitwidth increases. In the following steps, the MVM results will be further processed through other intermediate data processing operations such as activation and pooling. These operations are executed in the form of quantized representation with the help of auxiliary hardware within the tile, during which the data bitwidth remains unchanged. The last step is weight quantization. During this process, the weight quantization hardware within the tile reduces the data bitwidth to the required low bitwidth. Then, the low-bitwidth output is passed to the next layer. The subsequent neural network layers also repeatedly execute these three steps until the final output is obtained.
[0007] When performing matrix-vector multiplication operations, the elements in the operation results may have larger values than the elements of the matrix and vector participating in the MVM. Therefore, more bits are required to store the operation results. Neural networks happen to need to perform consecutive matrix-vector multiplications, and modern neural networks usually contain dozens or even hundreds of layers, which will lead to a serious problem of bitwidth inflation in quantized neural networks. Figure 2 Shown is an example of bitwidth inflation. In this figure, when the weight and activation bitwidths of the original input are both 8 bits, the bitwidth of the final activation may be as high as hundreds. It is impractical to support the calculation and storage of data with such high and irregular bitwidths. Therefore, the bitwidth of the activation must be reduced before it is passed to the next layer. The indispensable process of converting the data in the quantized representation from high bitwidth to low bitwidth is called weight quantization.
[0008] Deploying a quantized neural network on a memristor-based accelerator helps achieve lower power consumption, higher computational density, higher energy efficiency, and higher acceleration ratio. However, at the same time, the inference accuracy of the neural network will decrease due to limited data mapping capabilities and incomplete weight quantization implementation. Specifically, for both activation and weight mapping, existing memristor accelerators only support symmetric quantization. In practical applications, the data ranges of activations and weights can be extremely asymmetric, which makes the data range represented by the mapping method not match the actual data range, thereby increasing the precision loss during the quantization data conversion process. In low-precision quantized neural networks, the data bitwidth is low (weights, activations, etc. are quantized into fewer bits, such as less than 8 bits), and this precision loss is particularly obvious. On the other hand, in existing memristor accelerators, either the host computing device (such as CPU, GPU, and FPGA) is used to perform the weight quantization operation, or the weight quantization operation is simplified. The former offloading the weight quantization operation to other computing devices will weaken the benefits of reducing data movement in the memristor accelerator, greatly reducing its performance and energy efficiency, and is also not applicable to scenarios such as edge computing that lack additional computing devices; the latter only uses a shifter to implement weight quantization, which will result in serious precision loss and even overflow problems. Figure 3 Shown is an example of overflow caused by using this kind of weight quantization hardware. Overflow will cause the sign bit of the integer with overflow to flip, and this serious overflow error introduces a great deal of precision loss. In low-precision quantized neural networks, the lower data bitwidth increases the likelihood of overflow.
[0009] Generally speaking, existing memristor accelerators will lead to a decrease in inference accuracy during the process of accelerating the inference of quantized neural networks (especially low-precision quantized neural networks). Summary of the Invention
[0010] In view of the defects and improvement requirements of the prior art, the present invention provides a memristor accelerator for deploying a quantized neural network and a method for deploying a quantized neural network. The purpose is to improve the inference accuracy of the quantized neural network deployed on the memristor accelerator.
[0011] To achieve the above object, according to one aspect of the present invention, there is provided a memristor accelerator for deploying a quantized neural network, including: a plurality of processing units interconnected by a bus, a Tile-level cache connected to the bus, and a first adder, a pooling hardware, an activation hardware, and a weight quantization hardware connected to the Tile-level cache;
[0012] The processing unit includes multiple memristor arrays. Each row of each memristor array is connected to a digital-to-analog converter, and each column is sequentially connected to an analog-to-digital converter and a first shifter; the columns in the memristor array include regular columns and zero columns. The regular columns are used to map the quantized weight matrix Q(W) of the target layer to be deployed in the quantized neural network, and the zero columns are used to map the zero point z W ; the digital-to-analog converter is used to input the quantized activation vector Q(a) of the target layer to the memristor array, so that the memristor array performs the matrix-vector multiplication operation between the weight matrix and the activation vector of the target layer; the analog-to-digital converter and the first shifter are sequentially used to perform analog-to-digital conversion and shifting on the operation results of each column;
[0013] The Tile-level cache is used to cache the operation results output by each processing unit and the fused constant z a (Q(W)-z W J W )J a The quantized bias Q(b)′ = Q(b)-z a (Q(W)-z W J W )J a ;
[0014] The first adder is used to combine the operation results of the target layer on each processing unit and the quantized bias Q(b)′ to obtain the calculation result of the target layer and cache it to the Tile-level cache;
[0015] The pooling hardware and the activation hardware are sequentially used to perform pooling operation and activation operation on the calculation result of the target layer and cache the operation results to the Tile-level cache;
[0016] The requantization hardware is used to requantize the calculation result of the target layer after the pooling operation and the activation operation to reduce its data bit width to the specified bit width, so as to obtain the activation vector of the next layer;
[0017] where z W is the zero point in the quantization function for quantizing the weight matrix, z a is the zero point in the quantization function for quantizing the activation vector, Q(b) is the quantized bias of the target layer, J W and J a respectively represent all-1 matrices with the same shape as the weight matrix and the activation vector.
[0018] In some alternative embodiments, the requantization hardware includes:
[0019] A second shifter, whose input end is connected to the Tile-level cache, and is used to shift the calculation result Q(x) of the target layer after the pooling operation and the activation operation to the right by (e a -e x) bits to obtain the operation result R 1 ; e x represents the exponent of Q(x), and e a represents the exponent of Q(a);
[0020] A second adder, whose first input terminal is connected to the output terminal of the second shifter, whose second input terminal is used to input the rounding bit, and which is used to perform a rounding operation on the operation result R 1 to obtain the operation result R 2 ;
[0021] A third adder, whose first input terminal is connected to the output terminal of the second adder, whose second input terminal is used to input the zero point z a , and which is used to perform R 2 +z a to obtain the operation result R 3 ;
[0022] And a clipping hardware, whose input terminal is connected to the output terminal of the third adder, and which is used to clip the operation result R 3 within the range of [N min , N max ;
[0023] wherein, z a is the zero point in the quantization function for quantifying the activation vector, and N max and N min are respectively the upper and lower bounds of the quantization value.
[0024] In some alternative embodiments, the requantization hardware includes: a second shifter, a second adder, and a clipping hardware;
[0025] A second shifter, whose input terminal is connected to the Tile-level cache, and which is used to shift the calculation result Q(x) of the target layer after the pooling operation and the activation operation to the right by (e a -e x ) bits to obtain the operation result R 1 ; e x represents the exponent of Q(x), and e a represents the exponent of Q(a);
[0026] The second adder includes calculation mode one and calculation mode two executed successively; in calculation mode one, the first input terminal of the second adder is connected to the output terminal of the second shifter, the second input terminal is used to input the rounding bit, and the second adder is used to perform a rounding operation on the operation result R 1 to obtain the operation result R 2 ; in calculation mode two, the first input terminal of the second adder is used to input the operation result R 2 , and the second input terminal is used to input the zero point z a, the second adder is used to perform r 2 +Z a , to obtain an operation result R 3 ;
[0027] Clipping hardware, whose input end is connected to the output end of the second adder, and is used to limit the operation result R 3 within the range of [N min , N max ;
[0028] wherein, z a is the zero point in the quantization function for quantizing the activation vector, and N max and N min are respectively the upper and lower bounds of the quantization values.
[0029] Furthermore, the clipping hardware includes:
[0030] A first comparator, whose positive input end is used to input N min , and whose negative input end is used to input the operation result R 3 ;
[0031] A first multiplexer, whose input end 0 is used to input the operation result R 3 , whose input end 1 is used to input N min , and whose address input end is connected to the output end of the first comparator;
[0032] A second comparator, whose positive input end is used to input N max , and whose negative input end is connected to the output end of the first multiplexer;
[0033] And a second multiplexer, whose input end 0 is used to input N min , whose input end 1 is connected to the output end of the first multiplexer, and whose address input end is connected to the output end of the second comparator.
[0034] Furthermore, in the same memristor array, the number of zero point columns is 1.
[0035] According to another aspect of the present invention, there is provided a method for deploying a quantized neural network based on a memristor accelerator, including:
[0036] Deploying each layer in the quantized neural network to the memristor accelerator in sequence; the memristor accelerator is the above-mentioned memristor accelerator for deploying the quantized neural network provided by the present invention;
[0037] For the l-th layer in the quantized neural network, deploying it to the memristor accelerator includes:
[0038] Take the l-th layer as the target layer, map the quantized weight matrix Q(W) of the target layer to the regular columns in the memristor array, and map the corresponding zeros of the weight matrix to the zero column z in the memristor array. W ;
[0039] Input the quantized activation vector Q(a) of the target layer into the memristor array through a digital-to-analog converter.
[0040] Fuse the constant z a (Q(W) - z W J W )J a The quantized bias Q(b)' = Q(b) - z a (Q(W) - z W J W )J a Cache it to the Tile-level cache.
[0041] where l ∈ {1, 2, …, L}, and L represents the number of layers of the quantized neural network.
[0042] Furthermore, in the quantization function, the calculation formula for the zero point is:
[0043]
[0044] where z represents the zero point, B represents the bit width of the quantization value, [α, β] represents the quantization range, s represents the scaling factor, represents the rounding operator.
[0045] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0046] (1) For the memristor accelerator for deploying a quantized neural network provided by the present invention, an additional zero column is introduced into the memristor array to map the zero points of the quantization parameters, and the zero point calculation of the quantized activation is fused into the bias calculation. Thus, it can support asymmetric quantization, so that the data range mapped by the memristor array matches the actual data range, reducing the accuracy loss in the quantization data conversion process and improving the inference accuracy of the quantized neural network deployed on the memristor accelerator. For a low-precision quantized neural network, the improvement in its inference accuracy is more significant.
[0047] (2) For the memristor accelerator for deploying a quantized neural network provided by the present invention, in its preferred solution, the re-quantization hardware includes a shifter, an adder, and a clipping hardware, and can perform a complete re-quantization operation where Q(x) and Q(a) respectively represent the data before and after re-quantization, e x represents the exponent of Q(x), e a represents the exponent of Q(a), za Represents the zero point in the quantization function for quantizing the activation vector, N max and N min are the upper and lower bounds of the quantization values respectively. In the present invention, since the re-quantization hardware can implement complete re-quantization operations, it can effectively reduce the precision loss during the re-quantization process and further improve the inference accuracy of the quantized neural network deployed on the memristor accelerator. For the low-precision quantized neural network, the improvement in its inference accuracy is more significant. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Shows a schematic diagram of the existing memristor accelerator architecture and the deployment of the quantized neural network on the memristor accelerator;
[0049] Figure 2 Shows an example diagram of the problem of bit-width expansion of the existing quantized neural network;
[0050] Figure 3 Shows an example diagram of data overflow during the existing re-quantization process;
[0051] Figure 4 Shows a schematic diagram of the structure of the memristor and the mapping of the weight matrix in the memristor accelerator provided by the embodiment of the present invention;
[0052] Figure 5 Shows a schematic diagram of the re-quantization hardware provided by the embodiment of the present invention;
[0053] Figure 6 Shows an example diagram of the re-quantization performed by the re-quantization hardware in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0055] In the present invention, the terms "first", "second", etc. (if any) in the present invention and the accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0056] Before explaining the technical solution of the present invention in detail, the calculations related to neural network quantization will be introduced first.
[0057] Quantization can map high-precision floating-point numbers into low-precision integers. A commonly used quantization function is as follows:
[0058]
[0059] Where x is the real number to be quantized, s is the scaling factor, z is the zero point (an integer, the value is not necessarily 0, actually it is the quantization value corresponding to the real number 0), N max and N min are the upper and lower bounds of the quantization value respectively, is the rounding operator, and the clamping function clamp(x; a, b) = min(max(x, a), b). After quantization, the quantization representation of x consists of the integer Q(x) and the quantization parameters s and z. Restoring the quantization representation to a real value can be calculated through the following dequantization operation:
[0060]
[0061] Obviously, due to the non-injectivity of the clamping function and rounding, the restored real value is not exactly equal to x.
[0062] Fixed-point numbers are widely used to represent data in existing memristor-based neural network accelerators, which means that the scaling factor during data quantization is an integer power of 2 (usually a negative integer). In this way, some calculations can be efficiently performed in a shift manner, which can provide further acceleration. Fixed-point quantization introduces a new quantization parameter - the exponent e, which can be calculated by the following formula:
[0063]
[0064] where [α, β] is the quantization range, B is the bit width of the quantization value, is the ceiling function. For fixed-point quantization, the scaling factor can be calculated by Equation (4) to ensure that it is an integer power of 2. Finally, the fixed-point representation of x consists of the integer Q(x) and the quantization parameters s and z.
[0065] s = 2 e (4)
[0066] The calculation of the forward propagation during the neural network inference process is shown in Equation (5). Where a l-1 , a l , W l and b l represent the activation of the (l - 1)-th layer, the activation, weight, and bias of the l-th layer respectively. σ(·) is the activation function, which is a non-linear function. The key operation of this calculation (i.e., W l a l-1 ) is the matrix-vector multiplication (MVM).
[0067] a l = σ(Wl a l-1 +b l ) (5)
[0068] In a quantized neural network, since elements are quantized to integers, MVM is performed in the form of multiplying an integer matrix by an integer vector. Taking the integer matrix-vector multiplication given by Equation (6) as an example, the original matrix and vector only require 8 bits to store elements, but the resulting vector requires 16 bits for storage. That is to say, the elements in the MVM result may have larger values and thus require more bits for storage. Neural networks happen to need to perform consecutive matrix-vector multiplications. Modern neural networks usually consist of dozens or even hundreds of layers, which exacerbates the problem of bit-width inflation as shown in Figure 2 the figure. As can be seen from the figure, the bit-width of the activations can be as high as hundreds. It is impractical to support the computation and storage of data with such high and irregular bit-widths. Therefore, the bit-width of the activations must be reduced before being passed to the next layer. The essential process of converting data in the quantized representation from a high bit-width to a low bit-width is called requantization.
[0069]
[0070] The deployment methods used by existing memristor-based accelerators are not designed specifically for low-precision scenarios, which may lead to a serious decline in the final classification accuracy. By analyzing the deployment of quantized neural networks on existing accelerators, the accuracy decline comes not only from the use of low-precision representations, but also from limited data mapping capabilities and incomplete requantization implementations.
[0071] Limited data mapping ability: Existing data mapping methods for memristor-based accelerators only support symmetric data ranges, which may not match the actual data range, resulting in additional precision loss. For weight mapping, there are mainly two methods used in existing accelerators: one uses the difference between positive and negative arrays, and the other represents the offset of integer powers of 2 (essentially the quantization parameter of zero point) through an additional unit column (i.e., the column storing the value "1"). For the former (referred to as the positive-negative array method for short), it uses the difference between two unsigned integers to represent a signed integer with one higher bit, and the represented data range is symmetric. For the latter (referred to as the unit column method for short), it uses the difference between an unsigned integer and an integer power of 2 to represent a signed integer. The value of this integer power of 2 is approximately half of the maximum representable value of the unsigned integer, so the unit column method also represents a symmetric data range. For activation mapping, the most commonly used is the two's complement representation method. Obviously, the range it represents is also symmetric. In short, both activation and weight mapping methods only support symmetric quantization. Unfortunately, the actual data ranges of activations and weights can be extremely asymmetric. Since the data range represented by the mapping method does not match the actual data range, this will increase the precision loss during the quantization data conversion process.
[0072] Incomplete weight quantization implementation: In existing memristor-based neural network accelerators, weight quantization hardware is either not mentioned or is too simple. Most previous accelerators do not mention weight quantization hardware in their architectures. If the host computing device (such as CPU, GPU, and FPGA) is used to perform weight quantization operations. Offloading these operations to other computing devices will weaken the benefits of reducing data movement in memristor-based neural network accelerators, greatly reducing their performance and energy efficiency. In addition, it is not suitable for scenarios such as edge computing that lack additional computing devices. A state-of-the-art work uses simplified weight quantization hardware, and this re-quantization hardware consists only of shifters. Using this weight quantization hardware may face problems of precision loss or overflow because it does not support rounding and clipping in Equation (1). Figure 3 Shows an example of overflow caused by using this kind of weight quantization hardware, and 3 integers overflowed. The overflow caused the sign bits of all of them to flip, and this serious overflow error introduced a great deal of precision loss.
[0073] Aiming at the problem of accuracy decline in the deployment of low-precision quantized neural networks on existing memristor-based accelerators, the present invention provides a memristor accelerator for deploying quantized neural networks and a method for deploying quantized neural networks. The overall idea is to improve the memristor array in the memristor accelerator for deploying quantized neural networks based on the observation that the data ranges of actual neural network activations and weights may be extremely asymmetric, so as to support asymmetric data to reduce the precision loss during the data conversion process and improve the inference accuracy of the quantized neural network after deployment. On this basis, further improve the weight quantization hardware in the memristor accelerator for deploying quantized neural networks to achieve a complete weight quantization function, so as to reduce the precision loss during the weight quantization process and further improve the inference accuracy of the quantized neural network.
[0074] The following are embodiments.
[0075] Embodiment 1:
[0076] A memristor accelerator for deploying quantized neural networks, as Figure 1 shown, includes: a plurality of processing units interconnected by a bus, a Tile-level cache connected to the bus, and a first adder, pooling hardware, activation hardware, and weight quantization hardware connected to the Tile-level cache.
[0077] In this embodiment, the processing unit includes a plurality of memristor arrays. Each row of each memristor array is connected to a digital-to-analog converter, and each column is sequentially connected to an analog-to-digital converter and a first shifter. The columns in the memristor array include regular columns and zero columns. The regular columns are used to map the quantized weight matrix Q(W) of the target layer to be deployed in the quantized neural network, and the zero columns are used to map the zero point z W ; The digital-to-analog converter is used to input the quantized activation vector Q(a) of the target layer to the memristor array, so that the memristor array performs the matrix-vector multiplication operation between the weight matrix and the activation vector of the target layer. The analog-to-digital converter and the first shifter are sequentially used to perform analog-to-digital conversion and shifting on the operation results of each column;
[0078] The Tile-level cache is used to cache the operation results output by each processing unit and the quantized bias Q(b)′ = Q(b) - z a (Q(W) - z W J W )J a fused with the constant z a (Q(W) - z W J W )J a ;
[0079] The first adder is used to combine the operation results of the target layer on each processing unit and the quantization bias Q(b)', obtain the calculation result of the target layer, and cache it in the Tile-level cache;
[0080] The pooling hardware and the activation hardware are respectively used to perform pooling operations and activation operations on the calculation result of the target layer, and cache the operation results in the Tile-level cache;
[0081] The requantization hardware is used to requantize the calculation result of the target layer after the pooling operation and the activation operation, reduce its data bit width to the specified bit width, so as to obtain the activation vector of the next layer;
[0082] where, z W is the zero point in the quantization function for quantizing the weight matrix, z a is the zero point in the quantization function for quantizing the activation vector, Q(b) is the bias after quantization of the target layer, J W and J a respectively represent all-ones matrices with the same shape as the weight matrix and the activation vector.
[0083] The memristor accelerator provided in this embodiment has the same overall architecture as the traditional memristor accelerator. In this embodiment, the data is quantized into unsigned integers, and its zero point is set to The relevant parameter definitions are as follows:
[0084] B represents the bit width of the quantization value, [α, β] represents the quantization range, s represents the scaling factor, represents the rounding operator.
[0085] Considering that the weight and activation data ranges in the neural network are often asymmetric, it is expected to support asymmetric data representation by improving the memristor array. Specifically, in this embodiment, an additional zero column is introduced in each memristor array to map the zero point z W . Figure 4 As shown, an example of implementing asymmetric weight mapping in this embodiment, where the quantized weights are stored in different units of multiple arrays in a bit-slice manner, Figure 4 The memristor units shown store 2 bits, and the weights are quantized to 8 bits. It should be noted that the number of bits that the memristor units can store and the number of bits occupied by the quantized weights here are only exemplary descriptions and should not be understood as the only limitation of the present invention. In practical applications, the number of bits that the memristor units can store and the number of bits occupied by the quantized weights may take other values. The elements in the quantized weight matrix Q(W) or the zero point z Ware mapped onto four memristor cells distributed across four arrays. An additional column is added to each array (see the memristor cells in the dashed box in the figure) to store the zeros introduced by the asymmetric weight quantization. The storage overhead introduced by adding a column is extremely small. For the most common 128×128 memristor array, the additional overhead is less than 1%. As for the calculation introduced by the zero terms, it is completed using the computing power of the array itself. The complete MVM calculation for asymmetric quantization y = Wa is given by Equation (7).
[0086]
[0087] In Equation (7), Q(W) and Q(a) represent the quantized weight matrix W and activation vector a respectively, and Q(y) represents the result y of the quantized matrix-vector multiplication operation; z W and z a are the zeros of the weight matrix W and activation vector a respectively. J is a matrix of all 1s with all elements being 1, and J W and J a are matrices of all 1s with the same shape as W and a respectively. The additional calculation term z W J W Q(a) is a vector with all elements being z W sum(Q(a)). Here, the function sum(·) represents summing all elements of a matrix (a vector is also a matrix). The calculations of z W sum(Q(a)) and Q(W)Q(a) are completed by the additional zero column and the regular column respectively with the help of the auxiliary peripheral circuit. By subtracting z W sum(Q(a)) element-wise from the vector Q(W)Q(a), the calculation result of part 1 in Equation (7) is completed.
[0088] In this embodiment, in order to implement arbitrary zeros to support asymmetric activation, the calculation of the zeros of the activation vector, i.e., part 2 in Equation (7), is fused into only the bias calculation. z a , Q(W) and z W are all pre-computed during offline quantization and remain unchanged during inference. Therefore, the calculation result of part 2 is a constant vector, and this constant vector can be absorbed by the bias. During the calculation of the bias involved in Equation (5), the additional calculation introduced by the activation zeros is naturally completed. In this embodiment, by fusing the constant z a (Q(W) - z W J W )J a into the original quantized bias Q(b), a new quantized bias Q(b)' is obtained and cached in the Tile-level cache. Through this preprocessing method, the zeros introduced by the activation do not introduce any additional overhead during runtime.
[0089] Generally speaking, in this embodiment, an additional zero column is introduced into the memristor array, introducing an additional quantization parameter zero, which can support asymmetric quantization and realize a new low-overhead mapping method. Specifically, by adding an additional zero column and directly utilizing the computing characteristics of the memristor array, the additional calculations caused by the weight zero are completed in-situ. For the zero of the activation vector, the calculation terms introduced by it are fused into the bias in advance, which can support the representation of asymmetric data ranges with almost no additional hardware overhead, effectively reducing the precision loss in the data mapping process and improving the inference accuracy of the quantized neural network deployed on the memristor accelerator.
[0090] To further reduce the precision loss and improve the inference accuracy, this embodiment further proposes a weight quantization hardware that can support the complete weight quantization function. Before introducing the weight quantization hardware, the quantization representation of the intermediate data, which is the input of the weight quantization hardware, is first explained. The intermediate data needs to perform various signed arithmetic operations, so the intermediate data is usually represented as a signed fixed-point number with zero as the zero point to avoid repeatedly adding / subtracting a non-zero zero point when performing various operations. Assume that the fixed-point representation of the output of the intermediate data processing (i.e., the input of the weight quantization) is composed of an integer tensor Q(x) with zero as the zero point and e x as the exponent, and the fixed-point representation of the desired activation vector is composed of an integer tensor Q(a) with z a as the zero point and e a as the exponent (the quantization parameters such as the zero point and the exponent are obtained in advance offline). According to formulas (1), (2), and (4), the calculations involved in weight quantization can be converted to:
[0091]
[0092] where ">>" represents a shift operation for scaling, the rounding operator is used to reduce the precision loss when converting to an integer, and the clipping function can mitigate the impact of overflow. The weight quantization hardware circuit designed in this embodiment can complete the above complete weight quantization process.
[0093] To simplify the hardware circuit as much as possible to reduce the hardware overhead while ensuring the realization of the corresponding functions, in this embodiment, for the first step Q(x) >> (e a - e x ), a shifter is used to complete it; when rounding the shifted result, the simplest rounding scheme (i.e., rounding down for 0 and rounding up for 1) is adopted. If the rounding bit is "1", 1 needs to be added; if the rounding bit is "0", no processing is performed. This rounding process is specifically implemented as the addition of the elements of the integer tensor and their corresponding rounding bits. Therefore, in this embodiment, the rounding operation is completed using an adder; after rounding, the zero point z aIt needs to be added to the integer tensor, and this embodiment will reuse this adder to implement this addition operation; the last step is to use hardware to implement the clamping operation. As described above, the clamping function clamp(x; a, b) = min(max(x, a), b). Therefore, this embodiment uses a maximum and minimum circuit to implement the clamping hardware to complete this calculation. Both of these circuits are composed of a comparator and a multiplexer.
[0094] It is easy to understand that the rounding bit is determined by the number of bits shifted to the right, specifically the next bit of the bits retained, that is, the highest bit among the bits discarded by the shift is the rounding bit.
[0095] Based on the above design idea, as Figure 5 shown, in this embodiment, the weight quantization hardware includes:
[0096] A second shifter, a second adder, and a clamping hardware;
[0097] The second shifter, whose input terminal is connected to the Tile-level cache, is used to shift the calculation result Q(x) of the target layer after the pooling operation and the activation operation to the right by (e a -e x ) bits to obtain the operation result R 1 ; e x represents the exponent of Q(x), and e a represents the exponent of Q(a);
[0098] The second adder includes calculation mode one and calculation mode two executed successively; in calculation mode one, the first input terminal of the second adder is connected to the output terminal of the second shifter, the second input terminal is used to input the rounding bit, and the second adder is used to perform a rounding operation on the operation result R 1 to obtain the operation result R 2 ; in calculation mode two, the first input terminal of the second adder is used to input the operation result R 2 , the second input terminal is used to input the zero point z a , and the second adder is used to perform r 2 +Z a to obtain the operation result R 3 ;
[0099] The clamping hardware, whose input terminal is connected to the output terminal of the second adder, is used to limit the operation result R 3 within the range of [N min , N max ;
[0100] where z a is the zero point in the quantization function used to quantize the activation vector, and N max and N min are the upper and lower bounds of the quantization values respectively.
[0101] As Figure 5 shown, in this embodiment, the clipping hardware specifically includes:
[0102] A first comparator, whose positive input terminal is used to input N min , and whose negative input terminal is used to input the operation result R 3 ;
[0103] A first multiplexer, whose input terminal 0 is used to input the operation result R 3 , whose input terminal 1 is used to input N min , and whose address input terminal is connected to the output terminal of the first comparator;
[0104] A second comparator, whose positive input terminal is used to input N max , and whose negative input terminal is connected to the output terminal of the first multiplexer;
[0105] And a second multiplexer, whose input terminal 0 is used to input N min , whose input terminal 1 is connected to the output terminal of the first multiplexer, and whose address input terminal is connected to the output terminal of the second comparator.
[0106] The components of the weight quantization hardware implemented in this embodiment are only composed of a shifter, an adder, and clipping hardware. The hardware logic of each component is simple enough and will not introduce significant hardware overhead.
[0107] It should be noted that in some other embodiments of the present invention, the weight quantization hardware for implementing the rounding operation and the operation of adding the rounding operation result and the activation zero point z a may not reuse the same adder, but instead use different adders respectively. Correspondingly, the weight quantization hardware specifically includes:
[0108] A second shifter, whose input terminal is connected to the Tile-level cache, and which is used to shift the calculation result Q(x) of the target layer after the pooling operation and the activation operation to the right by (e a -e x ) bits to obtain the operation result R 1 ; e x represents the exponent of Q(x), and e a represents the exponent of Q(a);
[0109] A second adder, whose first input terminal is connected to the output terminal of the second shifter, whose second input terminal is used to input the rounding bit, and which is used to perform a rounding operation on the operation result R 1 to obtain the operation result R 2 ;
[0110] A third adder, whose first input terminal is connected to the output terminal of the second adder, and whose second input terminal is used to input the zero point za , which is used to execute R 2 +z a , and obtain the operation result R 3 ;
[0111] And a clipping hardware, whose input end is connected to the output end of the third adder, and which is used to limit the operation result R 3 within the range of [N min , N max ; The specific implementation of the clipping hardware is as shown in Figure 5 ;
[0112] Wherein, z a is the zero point in the quantization function for quantifying the activation vector, and N max and N min are respectively the upper and lower bounds of the quantization value.
[0113] It should be noted that in some other embodiments of the present invention, the adder used to implement the rounding operation and the operation of adding the rounding operation result and the activation zero point z a in the re-quantization hardware may not share the same adder, but instead use different adders respectively. Accordingly, the re-quantization hardware specifically includes:
[0114] A second shifter, whose input end is connected to the Tile-level cache, and which is used to shift the calculation result Q(x) of the target layer after the pooling operation and the activation operation to the right by (e a -e x ) bits to obtain the operation result R 1 ; e x represents the exponent of Q(x), and e a represents the exponent of Q(a);
[0115] A second adder, whose first input end is connected to the output end of the second shifter, and whose second input end is used to input the rounding bit, and which is used to perform a rounding operation on the operation result R 1 to obtain the operation result R 2 ;
[0116] A third adder, whose first input end is connected to the output end of the second adder, and whose second input end is used to input the zero point z a , and which is used to execute R 2 +z a to obtain the operation result R 3 ;
[0117] And a clipping hardware, whose input end is connected to the output end of the third adder, and which is used to limit the operation result R 3 within the range of [N min , N max ; The specific implementation of the clipping hardware is as shown inFigure 5 as shown;
[0118] where z a is the zero point in the quantization function for quantizing the activation vector, and N max and N min are the upper and lower bounds of the quantization values respectively.
[0119] Compared with Figure 5 the weight quantization hardware shown, this implementation will use one more adder, but the overall hardware overhead is still extremely small.
[0120] Figure 6 shows the process of weight quantization of a specific integer tensor by the weight quantization hardware proposed in this embodiment. This integer tensor is consistent with Figure 3 the integer tensor to be weight quantized in, initially 8 bits, and the data bit width needs to be reduced to 4 bits through the weight quantization hardware. The upper and lower bounds of the quantization values correspond to the maximum and minimum values of 4-bit signed numbers, i.e., 7(0111b) and -8(1000b), where b is the suffix indicating a binary number.
[0121] Comparing Figure 6 and Figure 3 for the weight quantization process, it can be seen that Figure 3 the existing weight quantization hardware shown in realizes weight quantization through shifting and direct truncation. The right shift operation that directly loses the low bits without processing will make the value smaller, and the expected deviation is negative, while the direct truncation operation will cause an overflow problem. In contrast, this embodiment realizes complete weight quantization. This embodiment realizes the rounding operation in the weight quantization process, which may increase or decrease the value, and the expected deviation is close to 0, thereby avoiding the single error direction of directly discarding the low bits after shifting, and reducing the precision loss as a whole. In addition, this embodiment realizes the clipping operation, which greatly alleviates the impact brought by overflow. Figure 6 In the weight quantization result of, none of the four integers has the sign flip problem caused by overflow. Comparing Figure 3 with the weight quantization result of, it can be seen that this embodiment can effectively reduce the precision loss in the data weight quantization process.
[0122] Generally speaking, this embodiment uses simple devices such as shifters, adders, multiplexers, and comparators to realize a weight quantization hardware that can support the complete weight quantization function, which can effectively reduce the precision loss in the data weight quantization process without introducing significant hardware overhead, and further improve the inference accuracy of the quantized neural network deployed on the memristor accelerator.
[0123] Embodiment 2:
[0124] A method for deploying a quantized neural network based on a memristor accelerator, including:
[0125] Deploy each layer in the quantized neural network to the memristor accelerator in sequence; the memristor accelerator is the memristor accelerator for deploying the quantized neural network provided in the above-mentioned Embodiment 1;
[0126] For the l-th layer in the quantized neural network, deploying it to the memristor accelerator includes:
[0127] Take the l-th layer as the target layer, map the quantized weight matrix Q(W) of the target layer to the regular columns in the memristor array, and map the corresponding zero points of the weight matrix to the zero point column z in the memristor array W ;
[0128] Input the quantized activation vector Q(a) of the target layer into the memristor array through a digital-to-analog converter;
[0129] Cache the quantized bias Q(b)′ = Q(b) - z a (Q(W) - z W J W )J a incorporated with the constant z a (Q(W) - z W J W )J a to the Tile-level cache;
[0130] where l ∈ {1, 2, …, L}, and L represents the number of layers of the quantized neural network.
[0131] In the quantization function, the calculation formula for the zero point is:
[0132]
[0133] It is easy to understand that mapping the weight matrix Q(W) to the regular columns in the memristor array is the same as the existing method for deploying the quantized neural network; the calculation of the zero point of the activation vector is a constant vector and will be incorporated into the bias.
[0134] An accelerator usually contains multiple Tiles. If the arrays in the existing Tiles can store the weights of all layers, then the quantized neural network can be completely deployed to the accelerator. If the weights of the entire network cannot be completely stored, then the deployment of the entire quantized neural network can be achieved through time-division multiplexing.
[0135] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A memristor accelerator for deploying a quantized neural network, characterized in that: include: A plurality of processing units interconnected by a bus, a tile-level cache connected to the bus, and a first adder, pooling hardware, activation hardware, and re-quantization hardware connected to the tile-level cache; The processing unit includes a plurality of memristor arrays, each row of each memristor array is connected to a digital-to-analog converter, and each column is sequentially connected to an analog-to-digital converter and a first shifter; the columns in the memristor array include regular columns and zero-point columns, the regular columns are used to map the quantized weight matrix Q(W) of the target layer to be deployed in the quantized neural network, and the zero-point columns are used to map the zero-point z W ; The digital-to-analog converter is used to input the quantized activation vector Q(a) of the target layer to the memristor array, so that the memristor array performs a matrix-vector multiplication operation between the weight matrix and the activation vector of the target layer; The analog-to-digital converter and the first shifter are used to perform analog-to-digital conversion and shift on the operation results of each column in sequence; The Tile-level cache is used to cache the calculation results output by each processing unit and the constant z a (Q(W)-z W J W )J a The quantization bias Q(b)′=Q(b)-z a (Q(W)-z W J W )J a ; The first adder is used to merge the operation results of the target layer on each processing unit and the quantization bias Q(b)′ to obtain the target layer calculation result, and cache it in the tile level cache; The pooling hardware and the activation hardware are used to perform pooling operations and activation operations on the target layer calculation results in sequence, and cache the operation results in the tile-level cache; The re-quantization hardware is used to re-quantize the target layer calculation result after the pooling operation and the activation operation, so that the data bit width is reduced to the specified bit width, thereby obtaining the activation vector of the next layer; Among them, z W is the zero point in the quantization function used to quantize the weight matrix, z a is the zero point in the quantization function used to quantize the activation vector, Q(b) is the quantized bias of the target layer, and J W and J a Represents an all-1 matrix with the same shape as the weight matrix and activation vector, respectively.
2. The memristor accelerator for deploying a quantized neural network according to claim 1, characterized in that: The requantization hardware includes: The second shifter, whose input end is connected to the Tile-level cache, is used to shift the target layer calculation result Q(x) after the pooling operation and the activation operation to the right (e a -e x ) bit, and get the operation result R1; e x represents the order code of Q(x), e a represents the order of Q(a); A second adder, a first input terminal of which is connected to the output terminal of the second shifter, and a second input terminal of which is used to input a rounding bit, and is used to perform a rounding operation on the operation result R1 to obtain an operation result R2; A third adder, whose first input terminal is connected to the output terminal of the second adder, and whose second input terminal is used to input a zero point z a , which is used to perform R2+z a , and obtain the operation result R3; and a limiting hardware, whose input end is connected to the output end of the third adder, which is used to limit the operation result R3 to [N min ,N max ] within the scope; Among them, z a is the zero point in the quantization function used to quantize the activation vector, N max and N min They are the upper and lower bounds of the quantized value respectively.
3. The memristor accelerator for deploying a quantized neural network according to claim 1, characterized in that: The re-quantization hardware includes: a second shifter, a second adder and limiting hardware; The second shifter, whose input end is connected to the tile-level cache, is used to shift the target layer calculation result Q(x) after the pooling operation and the activation operation to the right (e a -e x ) bit, and get the operation result R1; e x represents the order code of Q(x), e a represents the order of Q(a); The second adder includes a calculation mode 1 and a calculation mode 2 which are executed successively; in the calculation mode 1, the first input end of the second adder is connected to the output end of the second shifter, and the second input end is used to input the rounding bit, and the second adder is used to perform a rounding operation on the operation result R1 to obtain the operation result R2; in the calculation mode 2, the first input end of the second adder is used to input the operation result R2, and the second input end is used to input the zero point z a The second adder is used to perform R2+z a , and obtain the operation result R3; The limiting hardware, whose input end is connected to the output end of the second adder, is used to limit the operation result R3 to [N min ,N max ] within the scope; Among them, z a is the zero point in the quantization function used to quantize the activation vector, N max and N min They are the upper and lower bounds of the quantized value respectively.
4. The memristor accelerator for deploying a quantized neural network according to claim 2 or 3, characterized in that: The limiting hardware includes: The first comparator has its positive input terminal used to input N min , its reverse input terminal is used to input the operation result R3; The first multiplexer has its input terminal 0 for inputting the operation result R3, and its input terminal 1 for inputting N min , whose address input terminal is connected to the output terminal of the first comparator; The second comparator has its positive input terminal used to input N max , whose inverting input terminal is connected to the output terminal of the first multiplexer; And the second multiplexer, whose input terminal 0 is used to input N min , its input terminal No. 1 is connected to the output terminal of the first multiplexer, and its address input terminal is connected to the output terminal of the second comparator.
5. The memristor accelerator for deploying a quantized neural network according to claim 1, characterized in that: In the same memristor array, the number of zero-point columns is 1.
6. A quantized neural network deployment method based on a memristor accelerator, characterized in that: include: Deploy each layer in the quantized neural network into a memristor accelerator in sequence; the memristor accelerator is a memristor accelerator for deploying a quantized neural network as claimed in any one of claims 1 to 5; For the lth layer in the quantized neural network, deploying it to the memristor accelerator includes: The lth layer is taken as the target layer, the quantized weight matrix Q(W) of the target layer is mapped to the regular columns in the memristor array, and the corresponding zero points of the weight matrix are mapped to the zero point columns z in the memristor array W ; Inputting the quantized activation vector Q(a) of the target layer into the memristor array through a digital-to-analog converter; The constant z is incorporated a (Q(W)-z W J W )J a The quantization bias Q(b)′=Q(b)-z a (Q(W)-z W J W )J a Cache to the Tile-level cache; Among them, l∈{1,2,…,L}, L represents the number of layers of the quantized neural network.
7. The method for deploying a quantized neural network based on a memristor accelerator according to claim 6, characterized in that: In the quantization function, the calculation formula for the zero point is: Among them, z represents the zero point, B represents the bit width of the quantized value, [α, β] represents the quantization range, and s represents the scaling factor. Represents the rounding operator.
Citation Information
Patent Citations
Fixed-point quantitative convolutional neural network accelerator calculation circuit
CN111832719A
RRAM in-memory computing system array structure optimization-oriented method
CN115879530A
U8 quantized data matrix multiplication acceleration convolution operation method and device
CN118211013A
U8 quantization convolution acceleration device, method, equipment and medium
CN118966292A
Compiling asymmetrically-quantized neural network models for deep learning acceleration
US20230401420A1