Memristor accelerator for deploying quantized neural networks and quantized neural network deployment method

By introducing zero-point columns and weighted hardware into the memristor array, the accuracy degradation caused by data mapping mismatch in memristor accelerators is solved, achieving higher inference accuracy.

CN120087431BActive Publication Date: 2025-11-28HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510225854.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-11-28
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

Existing memristor accelerators suffer from decreased inference accuracy when deploying quantized neural networks, especially low-precision quantized neural networks. This is mainly due to data loss caused by limited data mapping capabilities and unsuitable data mapping methods, as well as the fact that existing data processing methods do not support asymmetric quantization.

Method used

By introducing an additional zero column into the memristor array, asymmetric quantization is supported, and shifters, adders, and limiters are introduced into the weighting hardware to achieve complete weighting operations and reduce accuracy loss during data conversion.

Benefits of technology

It improves the inference accuracy of quantized neural networks on memristor accelerators, especially the accuracy of low-precision quantized neural networks, and reduces the accuracy loss during data mapping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087431B_ABST
    Figure CN120087431B_ABST
Patent Text Reader

Abstract

The application discloses a memristor accelerator for deploying a quantized neural network and a quantized neural network deployment method, and belongs to the field of information storage and computer architecture. A column in a processing unit memristor array includes a regular column and a zero point column. The regular column is used for mapping a quantized weight matrix of a target layer to be deployed in the quantized neural network. The zero point column is used for mapping a weight zero point. A digital-to-analog converter is used for inputting a quantized activation vector of the target layer to the memristor array, so that the memristor array performs a matrix vector multiplication operation between the weight matrix and the activation vector. A Tile-level cache is used for caching operation results output by each processing unit and a bias fused with a zero point of a constant activation vector. An adder is used for merging operation results of the target layer on each processing unit and the fused bias. The application can improve the inference accuracy of the quantized neural network deployed on the memristor accelerator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information storage and computer architecture, and more specifically, relates to a memristor accelerator for deploying quantized neural networks and a method for deploying quantized neural networks. Background Technology

[0002] The forward propagation calculation in the neural network inference process is as follows:

[0003]

[0004] in, Indicates the first The activation vector of the layer, , and They represent the first The activation vector, weight matrix, and bias vector of the layer. The activation function is a non-linear function. The key operation of this computation (i.e.) ) is Matrix Vector Multiplication (MVM).

[0005] Memristor crossbar (XB) arrays possess the ability to perform analog domain matrix-vector multiplication in situ with high energy efficiency and speed. Memristor-based accelerators have been proposed to accelerate neural networks. In such accelerators, digital-to-analog converters (DACs) and analog-to-digital converters (ADCs) are used to perform the conversion between digital and analog signals. Due to the limited accuracy of ADCs / DACs and the bit depth of memristor cells, vector elements typically require multiple cycles of input, and matrix elements also need to be stored in multiple memristor cells located across multiple columns. Therefore, shift-and-add (S&A) circuits are used to combine computation results from different input cycles and different columns.

[0006] A quantized neural network refers to a neural network whose parameters have been quantized. To accelerate this neural network using a memristor-based accelerator, the quantized neural network needs to be deployed within the memristor-based accelerator. The entire deployment process is as follows: Figure 1 As shown. Each layer of the neural network performs the three steps shown in the dashed box. The first step is an MVM that quantizes the weights and activations (the first layer quantizes the input), and this matrix-vector multiplication is mapped to the processing element (PE) for execution. ). The quantized weights come from offline quantization of the pretrained weights, which remain unchanged at runtime. Except for the first layer, quantized activations come from re-quantization of intermediate data at runtime (referred to as re-quantization). In particular, the quantized input of the first layer is obtained through offline quantization before participating in MVM. In addition, the computation results from multiple PEs are merged by adders within a tile ). During MVM, the data bit-width increases. In the following steps, the MVM results will be further processed by other intermediate data processing operations such as activations, pooling, etc. ). These operations are performed in quantized representation with the help of auxiliary hardware within a tile, during which the data bit-width remains unchanged. The last step is re-quantization ). During this process, the re-quantization hardware within a tile reduces the data bit-width to the required low bit-width. Then, the low bit-width output is passed to the next layer. The three steps are repeatedly performed for subsequent neural network layers until the final output is obtained.

[0007] When performing matrix vector multiplication operations, the elements in the operation results can have larger values than the elements of the matrix and vector participating in MVM, so more bits are needed to store the operation results. Neural networks need to perform consecutive matrix vector multiplications, and modern neural networks usually contain tens of layers or even hundreds of layers, which can cause serious bit-width expansion problems in quantized neural networks. Figure 2 An example of bit-width expansion is shown, in which the bit-width of the final activation can be as high as hundreds under the condition that the bit-width of the original input weight and activation is 8 bits. It is impractical to support the computation and storage of data with such high and irregular bit-width. Therefore, the bit-width of the activation must be reduced before it is passed to the next layer. This indispensable process of converting data in quantized representation from high bit-width to low bit-width is called re-quantization.

[0008] Deploying quantized neural networks on memristor-based accelerators helps to achieve lower power consumption, higher computing density, higher energy efficiency and higher speedup, but at the same time, the inference accuracy of the neural network is reduced due to limited data mapping capability and incomplete weight quantization implementation. Specifically, regardless of the mapping of activations or weights, existing memristor accelerators only support symmetric quantization, while in actual applications, the data range of activations and weights can be extremely asymmetric, which makes the data range represented by the mapping method not match the actual data range, thereby increasing the precision loss in the quantization data conversion process. In low-precision quantized neural networks, the data bit width is low (the weight, activation, etc. quantized data is quantized into a smaller number of bits, such as 8 bits or less), and this precision loss is particularly evident. On the other hand, in existing memristor accelerators, either the host computing device (such as CPU, GPU and FPGA) is used to perform the weight quantization operation, or the weight quantization operation is simplified, the former offloads the weight quantization operation to other computing devices, which will weaken the benefits of reducing data movement in the memristor accelerator, greatly reducing its performance and energy efficiency, and is not suitable for edge computing and other scenarios that lack additional computing devices; the latter only uses a shifter to implement weight quantization, which will cause serious precision loss and even overflow problems, Figure 3 An example of using such weight quantization hardware to cause overflow is shown, and overflow will cause the sign bit of the integer that appears to overflow to flip, which introduces a large precision loss due to this serious overflow error. In low-precision quantized neural networks, the lower data bit width increases the likelihood of overflow.

[0009] Overall, existing memristor accelerators will cause a decrease in inference accuracy when accelerating the inference process of quantized neural networks (especially low-precision quantized neural networks). SUMMARY

[0010] In view of the defects of the prior art and the need for improvement, the present application provides a memristor accelerator for deploying a quantized neural network and a method for deploying a quantized neural network, which aims to improve the inference accuracy of a quantized neural network deployed on a memristor accelerator.

[0011] To achieve the above-mentioned purpose, according to one aspect of the present application, a memristor accelerator for deploying a quantized neural network is provided, comprising: a plurality of processing units interconnected by a bus, a Tile-level cache connected to the bus, and a first adder, pooling hardware, activation hardware and weight quantization hardware connected to the Tile-level cache;

[0012] The processing unit includes multiple memristor arrays. Each row of each memristor array is connected to a digital-to-analog converter (DAC), and each column is connected to an analog-to-digital converter (ADC) and a first shifter. The columns in the memristor arrays include regular columns and zero-point columns. The regular columns are used to map and quantize the quantized weight matrix of the target layer to be deployed in the neural network. The zero-point column is used to map zero points. The digital-to-analog converter is used to input the quantized activation vector of the target layer into the memristor array. This enables the memristor array to perform matrix-vector multiplication between the weight matrix and activation vector of the target layer; the analog-to-digital converter and the first shifter are used sequentially to perform analog-to-digital conversion and shifting on the results of each column of operations;

[0013] Tile-level cache is used to cache the calculation results output by each processing unit and incorporates constants. Quantization bias ;

[0014] The first adder is used to combine the computation results of the target layer across each processing unit and the quantization bias. The target layer calculation results are obtained and cached in the Tile-level cache;

[0015] The pooling hardware and activation hardware are used to perform pooling and activation operations on the target layer computation results in turn, and the operation results are cached in the Tile-level cache.

[0016] The weighting hardware is used to weight the target layer computation results after pooling and activation operations, reducing its data bit width to a specified bit width, thereby obtaining the activation vector of the next layer;

[0017] in, These are the zeros in the quantization function used to quantize the weight matrix. These are the zeros in the quantization function used to quantize the activation vector. The bias of the target layer after quantization. and These represent all-one matrices with the same shape as the weight matrix and activation vector, respectively.

[0018] In some alternative embodiments, the weighted hardware includes:

[0019] The second shifter, whose input is connected to the Tile-level cache, is used to process the target layer computation results after pooling and activation operations. Move to the right The bits are used to obtain the result of the operation. ; express The exponent, express The exponent;

[0020] a second adder, a first input terminal of which is connected to an output terminal of the second shifter, and a second input terminal of which is used for inputting a rounding bit, for performing a rounding operation on the operation result to obtain an operation result ;

[0021] a third adder, a first input terminal of which is connected to an output terminal of the second adder, and a second input terminal of which is used for inputting a zero point , for performing to obtain an operation result ;

[0022] and a limiting hardware, an input terminal of which is connected to an output terminal of the third adder, for limiting the operation result to a range of ;

[0023] wherein, is a zero point in a quantization function used for quantizing the activation vector, and are upper and lower bounds of the quantized value respectively.

[0024] In some optional embodiments, the re-quantization hardware comprises: the second shifter, the second adder and the limiting hardware;

[0025] a second shifter, an input terminal of which is connected to the Tile-level cache, for shifting the target layer calculation result after the pooling operation and the activation operation right by bits to obtain an operation result ; representing a mantissa of , and representing a mantissa of ;

[0026] the second adder comprises a calculation mode one and a calculation mode two executed in sequence; in the calculation mode one, a first input terminal of the second adder is connected to an output terminal of the second shifter, and a second input terminal is used for inputting a rounding bit, and the second adder is used for performing a rounding operation on the operation result to obtain an operation result ; in the calculation mode two, the first input terminal of the second adder is used for inputting the operation result , the second input terminal is used for inputting a zero point , and the second adder is used for performing to obtain an operation result ;

[0027] the limiting hardware, an input terminal of which is connected to an output terminal of the second adder, for limiting the operation result to a range of within a range;

[0028] wherein, is a zero point in a quantization function used for quantizing an activation vector, and are upper and lower bounds of quantized values, respectively.

[0029] Further, the amplitude limiting hardware comprises:

[0030] a first comparator, a positive input terminal of which is used for inputting , and a negative input terminal of which is used for inputting an operation result ;

[0031] a first multiplexer, a 0 input terminal of which is used for inputting the operation result , a 1 input terminal of which is used for inputting , and an address input terminal of which is connected with an output terminal of the first comparator;

[0032] a second comparator, a positive input terminal of which is used for inputting , and a negative input terminal of which is connected with an output terminal of the first multiplexer;

[0033] and a second multiplexer, a 0 input terminal of which is used for inputting , a 1 input terminal of which is connected with an output terminal of the first multiplexer, and an address input terminal of which is connected with an output terminal of the second comparator.

[0034] Further, in the same memristor array, the number of zero point columns is 1.

[0035] According to still another aspect of the present application, there is provided a method for deploying a quantized neural network based on a memristor accelerator, comprising:

[0036] deploying each layer in the quantized neural network into the memristor accelerator in sequence; the memristor accelerator is the above-mentioned memristor accelerator for deploying a quantized neural network provided by the present application;

[0037] for the first l layer in the quantized neural network, deploying it into the memristor accelerator comprises:

[0038] taking the first l layer as a target layer, mapping a weight matrix of the target layer after quantization to a regular column in the memristor array, and mapping a zero point corresponding to the weight matrix to a zero point column in the memristor array ;

[0039] inputting an activation vector of the target layer after quantization into the memristor array through a digital-to-analog converter;

[0040] ​Incorporating constants Quantization bias Cache to Tile level cache;

[0041] in, , L This indicates the number of layers in a quantized neural network.

[0042] Furthermore, in the quantization function, the formula for calculating the zero point is:

[0043]

[0044] Where z represents the zero point, and B represents the bit width of the quantization value. This indicates the quantization range, and s represents the scaling factor. This indicates the rounding operator.

[0045] In summary, the above-described technical solutions conceived in this invention can achieve the following beneficial effects:

[0046] (1) The memristor accelerator for deploying quantized neural networks provided by the present invention introduces an additional zero column in its memristor array to map the zeros of quantization parameters and integrates the calculation of zeros of quantization activation into the bias calculation. This can support asymmetric quantization, thereby matching the data range mapped by the memristor array with the actual data range, reducing the accuracy loss in the quantization data conversion process, and improving the inference accuracy of the quantized neural network deployed on the memristor accelerator. For low-precision quantized neural networks, the improvement in inference accuracy is more significant.

[0047] (2) In a preferred embodiment of the memristor accelerator for deploying quantized neural networks provided by the present invention, the weighting hardware includes shifters, adders and limiting hardware, which can perform a complete weighting operation. ,in, and These represent the data before and after weighting, respectively. express The exponent, express The exponent, This represents the zeros in the quantization function used to quantize the activation vector. and These are the upper and lower bounds of the quantization value, respectively. Because the weighting hardware in this invention can perform complete weighting operations, it can effectively reduce the accuracy loss during the weighting process, further improving the inference accuracy of quantized neural networks deployed on memristor accelerators. For low-precision quantized neural networks, the improvement in inference accuracy is even more significant. Attached Figure Description

[0048] Figure 1 An existing memristor accelerator architecture and a quantized neural network deployed on the memristor accelerator;

[0049] Figure 2 An example diagram of an existing quantized neural network bit width expansion problem;

[0050] Figure 3 An example diagram of data overflow in an existing re-quantization process;

[0051] Figure 4 A structure of a memristor in a memristor accelerator and a mapping diagram of a weight matrix provided by an embodiment of the present application;

[0052] Figure 5 A re-quantization hardware diagram provided by an embodiment of the present application;

[0053] Figure 6 An example diagram of re-quantization performed by re-quantization hardware in an embodiment of the present application. DETAILED DESCRIPTION

[0054] In order to make the objects, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as there is no conflict.

[0055] In the present application, the terms "first", "second", etc. (if any) in the present application and the accompanying drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.

[0056] Before the technical solutions of the present application are explained in detail, the neural network quantization related calculations are introduced.

[0057] Quantization can map high-precision floating-point numbers to low-precision integers. A commonly used quantization function is as follows:

[0058]

[0059] In the formula, is a real number to be quantized, is a scaling factor, is a zero point (an integer, the value is not necessarily 0, and in fact it is the quantization value corresponding to the real number 0), and are the upper and lower bounds of the quantization value, is a rounding operator, and the clipping function After quantization, the quantized representation of is an integer and quantization parameter and . The recovery of the quantized representation to a real value can be computed by a dequantization operation as follows:

[0060]

[0061] Obviously, since the clipping function and the rounding are not injective, the recovered real value is not exactly equal to .

[0062] Fixed-point numbers are widely used to represent data in existing memristor-based neural network accelerators. This means that the scaling factor for data quantization is an integer power of 2 (usually a negative integer). In this way, partial computations can be efficiently performed in a shift manner, which can provide further acceleration. Fixed-point quantization introduces a new quantization parameter - the mantissa , which can be computed by the following formula:

[0063]

[0064] where is the quantization range, is the bit-width of the quantized value, is the ceiling function. For fixed-point quantization, the scaling factor can be computed by equation (4) to ensure that it is an integer power of 2. Finally, the fixed-point representation of is composed of the integer and the quantization parameter and

[0065]

[0066] The computation of the forward propagation in the neural network inference process is shown in equation (5). Where , , and represent the activations of the th layer, the activations, weights and biases of the th layer, respectively. is an activation function, which is a non-linear function. The key operation (i.e. ) of this computation is the matrix vector multiplication (MVM).

[0067]

[0068] In quantized neural networks, since elements are quantized to integers, MVM is performed in the form of multiplication of integer matrix and integer vector. Take the integer matrix vector multiplication given in equation (6) as an example, the original matrix and vector only need 8 bits to store elements, but the result vector needs 16 bits to store. That is, the elements in the MVM result can have larger values, so more bits are needed for storage. Neural networks need to perform consecutive matrix vector multiplications, and modern neural networks usually contain tens of layers or even hundreds of layers, which exacerbates the problem of bit width expansion in quantized neural networks. As shown in the figure, the bit width of the activation can be as high as hundreds. It is impractical to support the computation and storage of data with such high and irregular bit widths. Therefore, the bit width of the activation must be reduced before it is passed to the next layer. This indispensable process of converting data in quantized representation from high bit width to low bit width is called re-quantization. Figure 2

[0069]

[0070] The deployment method used by existing memristor-based accelerators is not specifically designed for low-precision scenarios, which can cause a serious drop in final classification accuracy. By analyzing the deployment of quantized neural networks on existing accelerators, the accuracy drop comes not only from using low-precision representation, but also from limited data mapping capability and incomplete re-quantization implementation.

[0071] Limited data mapping capability: the existing data mapping method of the memristor-based accelerator only supports symmetric data range, which can not match the actual data range, resulting in additional precision loss. For weight mapping, the existing accelerator mainly uses two methods: one uses the difference between the positive array and the negative array, and the other represents the offset of 2 integer power (essentially the quantization parameter of zero point) through an additional unit column (i.e., a column storing the value "1"). For the former (referred to as the positive-negative array method), it uses the difference between two unsigned integers to represent a high-bit signed integer, and the represented data range is symmetric. For the latter (referred to as the unit column method), it uses the difference between an unsigned integer and a 2 integer power to represent a signed integer. The value of this 2 integer power is about half of the maximum representation value of the unsigned integer, so the unit column method also represents a symmetric data range. For activation mapping, the most commonly used is the two's complement representation method. Obviously, its represented range is also symmetric. In summary, whether it is activation or weight mapping method, only symmetric quantization is supported. Unfortunately, the data range of actual activations and weights can be extremely asymmetric. Since the data range represented by the mapping method does not match the actual data range, this will increase the precision loss in the conversion process of quantized data.

[0072] ​Incomplete weight quantization implementation: In existing memristor-based neural network accelerators, the weight quantization hardware is either not mentioned or too simple. Most of the previous accelerators do not mention the weight quantization hardware in their architecture. If the weight quantization operation is performed using the host computing device (such as CPU, GPU and FPGA). Offloading these operations to other computing devices will weaken the benefits of reducing data movement in the memristor-based neural network accelerator, greatly reducing its performance and energy efficiency. In addition, it is not suitable for scenarios such as edge computing that lack additional computing devices. One of the most advanced works uses a simplified weight quantization hardware, which is composed of only a shifter. Using this weight quantization hardware may face the problem of precision loss or overflow due to the lack of support for rounding and clipping in equation (1). Figure 3 An example of using such weight quantization hardware to cause overflow is shown, with 3 integers showing overflow. Overflow causes their sign bits to flip, and this serious overflow error introduces a great loss of precision.

[0073] In view of the problem of accuracy reduction in low-precision quantized neural network deployment on existing memristor-based accelerators, the present application provides a memristor accelerator for deploying a quantized neural network and a method for deploying a quantized neural network, the overall idea of which is to improve the memristor array in the memristor accelerator for deploying a quantized neural network to support asymmetric data to reduce the precision loss in the data conversion process and improve the inference accuracy of the deployed quantized neural network based on the observation that the actual data range of neural network activation and weight may be extremely asymmetric. On this basis, further improve the weight quantization hardware in the memristor accelerator for deploying a quantized neural network to realize complete weight quantization function to reduce the precision loss in the weight quantization process and further improve the inference accuracy of the quantized neural network.

[0074] The following is an example.

[0075] Example 1

[0076] A memristor accelerator for deploying a quantized neural network, as shown in Figure 1 , includes a plurality of processing units interconnected by a bus, a Tile-level cache connected to the bus, and a first adder, pooling hardware, activation hardware and weight quantization hardware connected to the Tile-level cache.

[0077] In this embodiment, the processing unit includes a plurality of memristor arrays, each row of each memristor array is connected to a digital-to-analog converter, and each column is connected to an analog-to-digital converter and a first shifter in turn; the columns in the memristor array include regular columns and zero columns, the regular columns are used to map the quantized weight matrix of the target layer to be deployed in the quantized neural network The zero-point column is used to map zero points. The digital-to-analog converter is used to input the quantized activation vector of the target layer into the memristor array. This enables the memristor array to perform matrix-vector multiplication between the weight matrix and activation vector of the target layer; the analog-to-digital converter and the first shifter are used sequentially to perform analog-to-digital conversion and shifting on the results of each column of operations;

[0078] Tile-level cache is used to cache the calculation results output by each processing unit and incorporates constants. Quantization bias ;

[0079] The first adder is used to combine the computation results of the target layer across each processing unit and the quantization bias. The target layer calculation results are obtained and cached in the Tile-level cache;

[0080] The pooling hardware and activation hardware are used to perform pooling and activation operations on the target layer computation results in turn, and the operation results are cached in the Tile-level cache.

[0081] The weighting hardware is used to weight the target layer computation results after pooling and activation operations, reducing its data bit width to a specified bit width, thereby obtaining the activation vector of the next layer;

[0082] in, These are the zeros in the quantization function used to quantize the weight matrix. These are the zeros in the quantization function used to quantize the activation vector. The bias of the target layer after quantization. and These represent all-one matrices with the same shape as the weight matrix and activation vector, respectively.

[0083] The memristor accelerator provided in this embodiment has the same overall architecture as a traditional memristor accelerator. In this embodiment, data is quantized into unsigned integers, and its zero point is set to... The relevant parameters are defined as follows:

[0084] B represents the bit width of the quantization value. Indicates the quantization range, and s represents the scaling factor. This indicates the rounding operator.

[0085] Considering that the weights and activation data ranges in neural networks are often asymmetric, it is desirable to support asymmetric data representation by improving the memristor array. Specifically, in this embodiment, an additional zero column is introduced in each memristor array to map the weight matrix. Corresponding zero point . Figure 4As shown, the embodiment realizes one example of asymmetric weight mapping, in which the quantized weights are stored in different cells of multiple arrays in the form of bit slices, Figure 4 The memristor cell shown in the middle is 2-bit, and the weights are quantized into 8 bits. It should be noted that the number of bits that the memristor cell can store and the number of bits that the quantized weights occupy here are only exemplary descriptions and should not be understood as the only limitation of the present application. In actual applications, the number of bits that the memristor cell can store and the number of bits that the quantized weights occupy can be other values. The elements or zeros in the quantized weight matrix are mapped to four memristor cells distributed in four arrays. Each array additionally adds a column (see the memristor cells in the dashed line box in the figure) for storing the zeros introduced by asymmetric weight quantization. The additional storage overhead introduced by adding a column is very small. For the most common 128x128 memristor array, the additional overhead is less than 1%. As for the calculation introduced by the zero term, the calculation ability of the array itself is used to complete it. The complete asymmetric quantized MVM calculation is given by formula (7).

[0086]

[0087] In formula (7), and respectively represent the quantized weight matrix W and the activation vector a , represents the quantized matrix vector multiplication operation result ; and are the zeros of the weight matrix and the activation vector , respectively. is an all-one matrix with all elements being 1, , are all-one matrices of the same shape as , . The additional calculation term introduced by the zero term is a vector with all elements being . Here, the function represents the sum of all elements of a matrix (vector is also a matrix). The calculations of are completed by the additional zero column and the regular column with the help of auxiliary peripheral circuits. By subtracting from the vector element by element, the calculation result of part 1 in formula (7) is completed.

[0088] In this embodiment, in order to realize arbitrary zero points to support asymmetric activation, the calculation of the zero points of the activation vector, i.e., part 2 in formula (7), is fused into the calculation of the bias only. , and are calculated in advance when offline quantization is performed, and they remain unchanged during inference. Therefore, the calculation result of part 2 is a constant vector, and this constant vector can be absorbed by the bias. In the calculation process of the bias involved in formula (5), the additional calculation introduced by the zero points of the activation is naturally completed. In this embodiment, the constant is fused into the original quantization bias to obtain a new quantization bias , and the new quantization bias is cached in the Tile-level cache. Through this pre-processing, the zero points introduced by the activation do not introduce any additional overhead at runtime.

[0089] Overall, this embodiment introduces an additional zero point column in the memristor array, introduces an additional quantization parameter zero point, supports asymmetric quantization, and realizes a new low-overhead mapping method. Specifically, by adding an additional zero point column and directly utilizing the calculation characteristics of the memristor array, the additional calculation caused by the zero points of the weights is completed in situ, and for the zero points of the activation vector, the calculation term introduced by the zero points is fused into the bias in advance. In this way, the asymmetric data range representation can be supported without introducing additional hardware overhead, the precision loss in the data mapping process can be effectively reduced, and the inference accuracy of the quantized neural network deployed on the memristor accelerator can be improved.

[0090] To further reduce the precision loss and improve the inference accuracy, this embodiment further proposes a re-quantization hardware that can support complete re-quantization. Before introducing the re-quantization hardware, the quantization representation of the input of the re-quantization hardware, i.e., the intermediate data, is described. The intermediate data needs to perform various signed arithmetic operations, and therefore the intermediate data is usually represented as a signed fixed-point number with a zero point of 0 to avoid repeatedly adding / subtracting a non-zero zero point when performing various operations. It is assumed that the fixed-point representation of the output of the intermediate data processing (i.e., the input of the re-quantization) is composed of an integer tensor with a zero point of 0 and a scale of , and the fixed-point representation of the activation vector to be obtained is composed of an integer tensor with a zero point of and a scale of (the zero points and the scale and other quantization parameters are obtained in advance offline). According to formulas (1), (2), and (4), the calculation involved in the re-quantization can be converted to:

[0091]

[0092] wherein, " indicates a shift operation, which is used to complete scaling, the rounding operator is used to reduce the precision loss of conversion into an integer, and the clipping function can reduce the impact of overflow. The quantization hardware circuit designed in this embodiment can complete the complete quantization process described above.

[0093] In order to simplify the hardware circuit as much as possible to reduce the hardware overhead while ensuring the realization of the corresponding function, the first step is completed using a shifter; when rounding the result after shifting, the simplest rounding scheme (i.e., 0 rounding 1) is adopted, and if the rounding bit is "1", 1 needs to be added; if the rounding bit is "0", no processing is performed. The rounding process is specifically implemented as the addition of an element of an integer tensor and its corresponding rounding bit, and therefore, in this embodiment, the rounding operation is completed using an adder; after rounding, a zero point needs to be added to the integer tensor, and the adder is reused in this embodiment to implement the addition operation; the last step is to use hardware to implement the clipping operation, as described above, the clipping function , and therefore, the maximum and minimum circuits are used to implement the clipping hardware to complete the calculation. Both of these circuits are composed of comparators and multiplexers.

[0094] It is easy to understand that the rounding bit is determined by the number of bits shifted to the right, which is the next bit of the remaining bit, that is, the highest bit in the discarded bit after shifting is the rounding bit.

[0095] Based on the above design idea, as shown in Figure 5 , in this embodiment, the quantization hardware includes:

[0096] a second shifter, a second adder, and clipping hardware;

[0097] The second shifter is connected to the Tile-level cache at the input end, and is used to move the target layer calculation result after the pooling operation and the activation operation to the right by bits to obtain an operation result ; , and , which represents the mantissa of ;

[0098] The second adder includes calculation mode one and calculation mode two executed in sequence; in calculation mode one, the first input end of the second adder is connected to the output end of the second shifter, the second input end is used to input the rounding bit, and the second adder is used to perform a rounding operation on the operation result to obtain an operation result ; in calculation mode two, the first input end of the second adder is used to input the operation result , the second input end is used for inputting zero point , the second adder is used for performing , to obtain operation result ;

[0099] the amplitude limiting hardware, the input end of which is connected to the output end of the second adder, is used for limiting the operation result in the range of ;

[0100] wherein, is zero point in the quantization function used for quantizing the activation vector, and are upper and lower bounds of the quantization value respectively.

[0101] As shown in Figure 5 , in the embodiment, the amplitude limiting hardware specifically comprises:

[0102] the first comparator, the positive input end of which is used for inputting , and the negative input end of which is used for inputting operation result ;

[0103] the first multiplexer, the 0 input end of which is used for inputting operation result , the 1 input end of which is used for inputting , and the address input end of which is connected to the output end of the first comparator;

[0104] the second comparator, the positive input end of which is used for inputting , and the negative input end of which is connected to the output end of the first multiplexer;

[0105] and the second multiplexer, the 0 input end of which is used for inputting , the 1 input end of which is connected to the output end of the first multiplexer, and the address input end of which is connected to the output end of the second comparator.

[0106] The components of the re-quantization hardware realized by the embodiment are only composed of shifters, adders and amplitude limiting hardware, and the hardware logic of each component is simple enough and will not introduce significant hardware overhead.

[0107] It should be noted that in some other embodiments of the present application, the operation for implementing the rounding operation and the operation for adding the rounding operation result and the activation zero point in the re-quantization hardware can also not reuse the same adder, but use different adders respectively, and correspondingly, the re-quantization hardware specifically comprises:

[0108] the second shifter, the input end of which is connected to the Tile level cache, is used for moving the target layer calculation result , which has undergone the pooling operation and the activation operation, rightward by the operation result ; the exponent of , the exponent of ; the exponent of

[0109] a second adder, a first input of which is connected to the output of the second shifter, and a second input of which is used for inputting the rounding bit, for performing a rounding operation on the operation result to obtain an operation result ;

[0110] a third adder, a first input of which is connected to the output of the second adder, and a second input of which is used for inputting the zero point , for performing to obtain an operation result ;

[0111] and a clipping hardware, an input of which is connected to the output of the third adder, for limiting the operation result to a range of ; the specific implementation of the clipping hardware is shown in Figure 5 ;

[0112] wherein is the zero point in the quantization function used for quantizing the activation vector, and are the upper and lower bounds of the quantized value, respectively.

[0113] Compared with the re-quantization hardware shown in Figure 5 , this implementation will use one more adder, but the overall hardware overhead is still very small.

[0114] Figure 6 The flow of the re-quantization hardware proposed in this embodiment for re-quantizing a specific integer tensor is shown, which is consistent with the integer tensor that needs to be re-quantized in Figure 3 , and is initially 8 bits, which needs to be reduced to 4 bits by the re-quantization hardware. The upper and lower bounds of the quantized value correspond to the maximum and minimum values of the 4-bit signed number, i.e. 7 (0111b) and -8 (1000b), where b is the suffix indicating the binary number.

[0115] By comparing the re-quantization processes of Figure 6 and Figure 3 , it can be seen that Figure 3The existing weight quantization hardware shown realizes weight quantization through shifting and direct truncation. The operation of right shifting without processing directly discards the low bits, which makes the numerical value smaller, and the expected deviation is negative, while the operation of direct truncation causes the problem of overflow. In contrast, the embodiment realizes complete weight quantization. The embodiment realizes the rounding operation in the weight quantization process, which can increase or decrease the numerical value, and the expected deviation is close to 0, thereby avoiding the single error direction of directly discarding the low bits after shifting, and reducing the precision loss as a whole. In addition, the embodiment realizes the limiting operation, which greatly alleviates the influence of overflow. Figure 6 In the weight quantization result, the four integers do not have the problem of sign flip caused by overflow, and compared with the weight quantization result of Figure 3 It can be known that the embodiment can effectively reduce the precision loss in the data weight quantization process.

[0116] Overall, the embodiment realizes the weight quantization hardware that can support complete weight quantization function by using simple devices such as shifters, adders, multiplexers, comparators, etc., can effectively reduce the precision loss in the data weight quantization process without introducing significant hardware overhead, and further improve the inference accuracy of the quantized neural network deployed on the memristor accelerator.

[0117] Embodiment 2:

[0118] A quantized neural network deployment method based on a memristor accelerator, comprising:

[0119] deploying each layer in the quantized neural network to the memristor accelerator in turn; the memristor accelerator is the memristor accelerator provided in the above embodiment 1 for deploying the quantized neural network;

[0120] for the first l layer in the quantized neural network, deploying it to the memristor accelerator, comprising:

[0121] taking the first l layer as a target layer, mapping the weight matrix of the target layer after quantization to a regular column in the memristor array, and mapping the corresponding zero point of the weight matrix to a zero point column in the memristor array;

[0122] inputting the activation vector of the target layer after quantization to the memristor array through a digital-to-analog converter;

[0123] caching the quantized bias fused with a constant to the Tile-level cache;

[0124] wherein, , LThe number of layers of the quantized neural network is represented.

[0125] In the quantization function, the calculation formula of the zero point is:

[0126]

[0127] It is easy to understand that the weight matrix is mapped to the conventional column in the memristor array, which is the same as the existing quantization neural network deployment method; the calculation of the zero point of the activation vector is a constant vector, which will be fused into the bias.

[0128] An accelerator usually contains multiple tiles. If the array in the existing tile can store the weights of all layers, the quantized neural network can be completely deployed on the accelerator. If it cannot completely store the weights of the entire network, the deployment of the entire quantized neural network can be realized through time division multiplexing.

[0129] It is easy for those skilled in the art to understand that the above description is only a preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A memristor accelerator for deploying quantized neural networks, characterized in that, include: Multiple processing units interconnected by a bus, a Tile-level cache connected to the bus, and a first adder, pooling hardware, activation hardware, and weighting hardware connected to the Tile-level cache. The processing unit includes multiple memristor arrays. Each row of each memristor array is connected to a digital-to-analog converter (DAC), and each column is sequentially connected to an DAC and a first shifter. The columns in the memristor arrays include regular columns and zero-point columns. The regular columns are used to map the quantized weight matrix of the target layer to be deployed in the quantized neural network. The zero-point column is used to map zero points. The digital-to-analog converter is used to input the quantized activation vector of the target layer into the memristor array. This enables the memristor array to perform matrix-vector multiplication between the weight matrix and activation vector of the target layer; the analog-to-digital converter and the first shifter are used sequentially to perform analog-to-digital conversion and shifting on the results of each column of operations; The Tile-level cache is used to cache the calculation results output by each processing unit and incorporates constants. Quantization bias ; The first adder is used to combine the operation results of the target layer on each processing unit and the quantization bias. The target layer calculation result is obtained and cached in the Tile-level cache; The pooling hardware and the activation hardware are used sequentially to perform pooling and activation operations on the target layer calculation results, and cache the operation results to the Tile-level cache. The weighting hardware is used to weight the target layer calculation results after pooling and activation operations, reducing its data bit width to a specified bit width, thereby obtaining the activation vector of the next layer. in, These are the zeros in the quantization function used to quantize the weight matrix. These are the zeros in the quantization function used to quantize the activation vector. The bias of the target layer after quantization. and These represent all-one matrices with the same shape as the weight matrix and activation vector, respectively; weighting is the process of converting the data in the quantized representation from high-bit width to low-bit width.

2. The memristor accelerator for deploying quantized neural networks as described in claim 1, characterized in that, The weighted hardware includes: The second shifter, whose input is connected to the Tile-level cache, is used to process the target layer calculation results after pooling and activation operations. Move to the right The bits are used to obtain the result of the operation. ; express The exponent, express The exponent; The second adder has its first input connected to the output of the second shifter, and its second input used to input the rounding bit, which is used to round the result. Perform a rounding operation to obtain the result. ; The third adder has its first input connected to the output of the second adder, and its second input used to input the zero point. It is used to execute The result of the calculation is obtained. ; And limiting hardware, whose input is connected to the output of the third adder, is used to convert the calculation result... Limited to Within the range; in, These are the zeros in the quantization function used to quantize the activation vector. and These are the upper and lower bounds of the quantized value, respectively.

3. The memristor accelerator for deploying quantized neural networks as described in claim 1, characterized in that, The weighted hardware includes: a second shifter, a second adder, and a limiting hardware; The second shifter, whose input is connected to the Tile-level cache, is used to process the target layer calculation results after pooling and activation operations. Move to the right The bits are used to obtain the result of the operation. ; express The exponent, express The exponent; The second adder includes a calculation mode one and a calculation mode two, executed sequentially. In the first calculation mode, the first input of the second adder is connected to the output of the second shifter, the second input is used to input the rounding bit, and the second adder is used to process the calculation result. Perform a rounding operation to obtain the result. In the second calculation mode, the first input terminal of the second adder is used to input the operation result. The second input terminal is used to input the zero point. The second adder is used to perform The result of the operation is obtained. ; The limiting hardware has its input connected to the output of the second adder, and is used to convert the calculation result... Limited to Within the range; in, These are the zeros in the quantization function used to quantize the activation vector. and These are the upper and lower bounds of the quantized value, respectively.

4. The memristor accelerator for deploying quantized neural networks as described in claim 2 or 3, characterized in that, The limiting hardware includes: The first comparator has its positive input terminal used for input. Its inverting input is used to input the calculation result. ; The first multiplexer uses its input terminal 0 to input the calculation result. Its input terminal 1 is used for input. Its address input terminal is connected to the output terminal of the first comparator; The second comparator has its positive input terminal used for input. Its inverting input terminal is connected to the output terminal of the first multiplexer; And a second multiplexer, whose input 0 is used for input. Its input terminal 1 is connected to the output terminal of the first multiplexer, and its address input terminal is connected to the output terminal of the second comparator.

5. The memristor accelerator for deploying quantized neural networks as described in claim 1, characterized in that, In the same memristor array, the number of zero columns is 1.

6. A method for deploying a quantized neural network based on a memristor accelerator, characterized in that, include: Each layer of the quantized neural network is sequentially deployed into a memristor accelerator; the memristor accelerator is the memristor accelerator for deploying a quantized neural network as described in any one of claims 1 to 5; For the first in the quantization neural network l Layers, which are deployed to the memristor accelerator, include: The first l The target layer is the quantized weight matrix of the target layer. Mapping to the regular columns of the memristor array, and mapping the corresponding zeros of the weight matrix to the zero columns of the memristor array. ; The activation vector of the target layer is quantized using a digital-to-analog converter. Input memristor array; Incorporating constants Quantization bias Cache to the Tile-level cache; in, , L This indicates the number of layers in a quantized neural network.

7. The method for deploying a quantized neural network based on a memristor accelerator as described in claim 6, characterized in that, In the quantization function, the formula for calculating the zero point is: Where z represents the zero point, and B represents the bit width of the quantization value. This indicates the quantization range, and s represents the scaling factor. This indicates the rounding operator.

Citation Information

Patent Citations

  • RRAM in-memory computing system array structure optimization-oriented method

    CN115879530A

  • U8 quantized data matrix multiplication acceleration convolution operation method and device

    CN118211013A