ReRAM convolution simplified implementation method
By performing high-low bit splitting and nested loop splitting on the input and weight data in ReRAM convolution operations, the problem of low accuracy in ReRAM convolution operations is solved, achieving higher computational accuracy and reducing circuit costs.
Patent Information
- Application Number
- CN202210334200.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2042-03-31
AI Technical Summary
Existing ReRAMs suffer from low precision during convolution operations, making it difficult to accurately adjust the weights of multiple bits. Furthermore, they require high precision from the ADC and DAC, leading to increased circuit implementation costs.
By splitting the input data and weight data into high and low bits, the multiplication operation is converted into a multiplication and addition operation. The nested loop is used to further split the data, reducing the accuracy requirements of the memristor and ADC/DAC. An adder is used to complete the final operation.
It significantly improves the accuracy of ReRAM convolution operations, reduces the requirements for the resistive state accuracy of memristor devices and the accuracy requirements of ADC/DAC circuits, and reduces the overall overhead cost.
Smart Images

Figure CN116127254B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of memristor data storage design, and more particularly to a ReRAM convolution simplification implementation method. BACKGROUND
[0002] Memristor (ReRAM) is a kind of nonlinear resistance with memory function, and its resistance value can be changed by controlling the change of current. If the high resistance value is defined as "1" and the low resistance value is defined as "0", the resistance can realize the function of storing data. Specifically, in actual use, the convolution operation of the memristor is usually used to realize the storage of different data. As shown in the accompanying drawings, the conductance G (the reciprocal of resistance R, R = 1 / G) of the memristor is taken as the convolution weight, and the input data is converted into voltage through the DAC circuit (digital-to-analog conversion circuit). Then the current I = V*G of each parallel ReRAM circuit. If a series of memristors are connected in parallel, the corresponding current output is the multiplication and addition result of all weights and data I = I1+I2+I3=V1*G1+V2*G2+V3*G3, and I is converted into data output through the ADC conversion circuit (analog-to-digital conversion circuit) again. This architecture of storage and calculation in one can greatly improve the performance and reduce the power loss. Figure 1 However, the corresponding problem is that the conductance of ReRAM needs to be accurately regulated in order to achieve the purpose of allowing users to set any weight value. Generally, the quantized weight value also needs to be 6~8
[0003] it, or even more. However, the process complexity of ReRAM itself cannot completely achieve the accurate regulation of multiple B it weights. This is a major technical problem of ReRAM used as convolution operation. At the same time, the accuracy required by ReRAM for ADC and DAC will also increase with the increase of calculation accuracy requirement. B SUMMARY
[0004] In view of the above problems, the purpose of the present application is to provide a ReRAM convolution simplification implementation method to solve the problem of low accuracy when implementing convolution operation through a memristor.
[0005] The ReRAM convolution simplification implementation method provided by the present application comprises:
[0006] sequentially obtaining each multiplication operation formula in the convolution to be operated;
[0007] respectively splitting the input data and the weight data in each multiplication operation formula to convert each multiplication operation formula into a corresponding multiplication and addition operation simplified formula;
[0008] The calculation results of each multiplication formula in the convolution to be operated are calculated in sequence by the multiplication-sum operation simplified formula.
[0009] The operation of the convolution to be operated is realized based on the calculation results of each multiplication formula.
[0010] In addition, preferably, the input data and the weight data in each multiplication formula are binary data.
[0011] In addition, preferably, the input data and the weight data in each multiplication formula are split into high and low bits to convert each multiplication formula into a corresponding multiplication-sum operation simplified formula.
[0012] For any multiplication formula, the input data A and the weight data B are obtained. A B The result of the multiplication formula is denoted as: ;
[0013] The input data A and the weight data B are split into high and low bits, so that the input data A is split into high-bit input data A and low-bit input data B. A B The weight data B is split into high-bit weight data B and low-bit weight data B. A A H A L B B H B L A L The bit width of A is L A , B L The bit width of B is L B ;
[0014] The input data A and the weight data B after splitting are substituted into the corresponding multiplication formula to obtain a corresponding multiplication-sum operation simplified formula.
[0015]
[0016] .
[0017] In addition, preferably, for any multiplication formula, the input data A of the matrix data type is converted into input data A of the vector data type, and then the input data A is split into high and low bits. A A A High-low bit splitting is performed.
[0018] In addition, preferably, the input data of the matrix data type is converted into input data of the vector data type to obtain n input vector data: A A 1~ A n ;
[0019] High-low bit splitting is performed on each of the input vector data.
[0020] The n input vector data and the weight data after splitting are substituted into the corresponding multiplication formula to obtain a corresponding multiplication and addition operation simplified formula. B
[0021] .
[0022] In addition, preferably, the calculation results of each multiplication formula in the convolution to be operated are calculated in sequence by the multiplication and addition operation simplified formula, including:
[0023] The multiplication and addition operation simplified formula is split into four multiplication operation simplified formulas.
[0024] =X1、 =X2、 =X3、 =X4
[0025] The calculation results of each multiplication formula in the convolution to be operated are calculated in sequence based on the four multiplication operation simplified formulas.
[0026] In addition, preferably, in the process of high-low bit splitting of the input data A and the weight data B :
[0027] Each multiplication formula in the convolution to be operated is cyclically split by nested loops to perform m high-low bit splitting on the input data A and the weight data B ; wherein,
[0028] m is an integer and ≥1.
[0029] In addition, preferably, in the process of cyclically splitting each multiplication formula in the convolution to be operated by nested loops:
[0030] Each multiplication operation simplified formula is further split to further split the value of A H into A H1 with A L1 , will be split into B H split into B H1 with B L1 ;
[0031] again split into A H1 , A L1 , B H1 , B L1 split, and so on until the operation precision of the convolution to be operated reaches a preset precision threshold.
[0032] In addition, preferably, the weight data is a memristor calculation matrix data.
[0033] In addition, preferably, in the process of implementing the operation of the convolution to be operated based on the calculation results of the respective multiplication operation formulas,
[0034] The operation result of the convolution to be operated is calculated by an adder based on the calculation results of the respective multiplication operation formulas.
[0035] Compared with the prior art, the ReRAM convolution simplified implementation method according to the application has the following beneficial effects:
[0036] The ReRAM convolution simplified implementation method provided by the application can significantly improve the precision of ReRAM convolution operation by splitting the input data and weight data into high and low bits, and the ReRAM convolution operation is no longer limited by the precision requirement of ADC / DAC; in addition, since the ReRAM convolution simplified implementation method provided by the application can cyclically split the input data and weight data into high and low bits, the precision of ReRAM convolution operation is always improved to the saturation value, so the requirement for the device resistance state precision of the memristor itself can be significantly reduced.
[0037] To achieve the above and related purposes, one or more aspects of the application include features that will be described in detail below and particularly pointed out in the claims. The following description and drawings detail certain illustrative aspects of the application. However, these aspects indicate only some of the various ways in which the principles of the application can be employed. In addition, the application is intended to include all such aspects and their equivalents. BRIEF DESCRIPTION OF DRAWINGS
[0038] Other objects and results of the present application will become more fully understood and appreciated with reference to the following description taken in conjunction with the accompanying drawings, in which:
[0039] Figure 1 A schematic diagram of multiplication and addition operation using an existing memristor according to an embodiment of the present application;
[0040] Figure 2 A schematic diagram of a ReRAM convolution simplified implementation method according to an embodiment of the present application;
[0041] Figure 3 A schematic diagram of a ReRAM convolution simplified implementation method according to an embodiment of the present application after loop splitting;
[0042] The same reference numbers in all the drawings indicate similar or corresponding features or functions. DETAILED DESCRIPTION
[0043] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of one or more embodiments. It can be evident, however, that embodiments can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form in order to facilitate describing one or more embodiments.
[0044] It needs to be explained in advance that before the specific structure of the ReRAM convolution simplified implementation method provided by the present application is introduced in detail, the principle of matrix convolution operation needs to be introduced simply, and the principle of matrix convolution operation is as follows (taking 3*3 matrix and 2*2 matrix convolution as an example):
[0045]
[0046] From the above formula, it can be seen that the convolution operation is mainly multiplication and addition operation. If the sizes of the multiplied matrices are inconsistent, there will be a shift operation, but the most basic operation is multiplication and addition, and each element of the matrix after convolution is the multiplication and addition result. Therefore, the multiplier is the core in the convolution operation, and therefore, the improvement and optimization of the multiplier directly determine the performance and overhead of the overall circuit architecture.
[0047] In addition, it also needs to be explained that ReRAM is a new type of memory device, which uses resistance to store 0 and 1 (high resistance state and low resistance state), and is a new type of storage technology similar to FLASH. Because ReRAM itself stores data by using resistance, according to the principle of Figure 1 , ReRAM is a good principle device for using resistance to do convolution operation. Now the operation architecture of convolution kernel using ReRAM is introduced Figure 1Current state will exist two problems.
[0048] 1. If we take the G value in the weight, the weight itself can reach 8Bit, 10Bit, or even directly use floating point 32Bit, which puts very high requirements on the multi-resistance state of the memistor itself, and the current parallel or memistor itself cannot meet such high precision requirements. Figure 1
[0049] 2. The input voltage represents the input data value, and the value itself may have the same precision as the weight, such as 10+Bit or 20+Bit, and under such high precision, the requirements for ADC or DAC conversion circuit are greatly increased, which further increases the circuit implementation cost of the memistor multiplier.
[0050] The ReRAM convolution simplification implementation method and the operation method provided by the application will be described in detail below with reference to the accompanying drawings.
[0051] Figure 2 The principle of the ReRAM convolution simplification implementation method according to the embodiment of the application is shown, Figure 3 The principle of the ReRAM convolution simplification implementation method according to the embodiment of the application after loop splitting is shown.
[0052] Combined with Figure 2 and Figure 3 It can be seen that the concept of the ReRAM convolution simplification implementation method provided by the application is to split the multiplication structure of the input in the convolution to be operated, thereby re-implementing the multiplier circuit in the convolution, reducing the A DC or weight precision requirements of the ReRAM multiplication operation.
[0053] The specific steps of the ReRAM convolution simplification implementation method according to the embodiment of the application will be described in detail below:
[0054] The ReRAM convolution simplification implementation method provided by the application comprises:
[0055] The multiplication operation formulas in the convolution to be operated are obtained in sequence;
[0056] The input data and weight data in each multiplication operation formula are respectively split into high and low bits to convert each multiplication operation formula into a corresponding multiplication and addition operation simplified formula;
[0057] The calculation results of each multiplication operation formula in the convolution to be operated are calculated in sequence through the multiplication and addition operation simplified formula;
[0058] The operation of the convolution to be operated is realized based on the calculation results of each multiplication operation formula.
[0059] It should be noted that, in general, in convolution operation, the input data and weight data in each multiplication formula are binary data.
[0060] Specifically, the input data and weight data in each multiplication formula are high-low split to convert each multiplication formula into a corresponding multiplication-addition operation simplified formula, including:
[0061] For any one multiplication formula, the input data A and the weight data B are obtained; then the result of the multiplication operation can be recorded as: ;
[0062] The input data A and the weight data B are high-low split, so that the input data A is split into high input data A H and low input data A L and the weight data B is split into high weight value B H and low weight value B L ; wherein, A L The bit width of L A , B L The bit width of L B ;
[0063] The split input data A and the weight data B are substituted into the corresponding multiplication formula to obtain the corresponding multiplication-addition operation simplified formula:
[0064]
[0065] .
[0066] In actual operation, for any one multiplication formula, we record the input data value as A , whose bit width is W A , and the weight value as B , whose bit width is W B . Then the multiplication result can be expressed as:
[0067] (1)
[0068] If the binary number A is split into high and low bits, the high bits are A H and the low bits are A L with a low bit width of L A . Similarly, the binary number B is split into B H , B L with a low bit width of L B . Then, equation (1) can be expressed as:
[0069]
[0070] .
[0071] In binary operations, the exponential power of 2 is equal to a shift operation, and the multiplication of A and B is changed into multiplication of four small bit widths and addition operation between them. In this way, the requirement for multiplication precision can be reduced. Moreover, the requirement for multiplication precision can be adjusted according to the actual situation and the adjustment of high and low bit widths of the multipliers A , B .
[0072] In addition, it should be noted that for any multiplication operation formula, the input data of the matrix data type A is converted into input data of the vector data type A , and then the input data A is split into high and low bits.
[0073] Specifically, the input data of the matrix data type A is converted into input data of the vector data type to obtain n input vector data: A 1~ A n ;
[0074] Each input vector data is split into high and low bits;
[0075] The split n input vector data and the weight data B are substituted into the corresponding multiplication operation formula to obtain a corresponding multiplication and addition operation simplified formula;
[0076] equation (2).
[0077] Through the above formula, a large bit width multiplication-addition circuit can be split into a small bit width multiplication-addition circuit in a split manner. This manner can reduce the requirement for the resistance state accuracy of the memristor, reduce the accuracy requirement for the input ADC or output DAC circuit, and reduce the overall cost and accuracy requirement.
[0078] The application will be described in detail below with reference to the accompanying drawings Figure 2 and 3 The working principle of the ReRAM convolution simplification implementation method provided by the application is described in detail.
[0079] By Figure 2 It can be seen that after the control logic converts the matrix data obtained from the data cache into vector data, there are n data A 1~ A n, each data is split into high and low bits, A H and A L , wherein the bit width of the high bit and the low bit can be set by itself, and the length of the low bit is L A The value of the memristor calculation array (i.e., the weight data) B H , B L , A 1~ A n The high and low bit data of the memristor calculation array will complete the multiplication-addition operation similar to the principle. Figure 1
[0080] Specifically, the multiplication-addition operation simplification formula is used to sequentially calculate the calculation results of each multiplication operation formula in the convolution to be operated, including:
[0081] The multiplication-addition operation simplification formula is split into four multiplication operation simplification formulas:
[0082] =X1、 =X2、 =X3、 =X4
[0083] Based on the four multiplication operation simplification formulas, the calculation results of each multiplication operation formula in the convolution to be operated are sequentially calculated.
[0084] Corresponding to the above Figure 2 , the four memristor calculation arrays correspond to the four multiplication-addition operations in formula 2, and finally the calculation result =X1、 =X2、 =X3、 =X4
[0085] These results are then shifted into the adder. After the addition is completed, the result is sent to the control logic, completing one multiply-accumulate operation in the convolution operation. The result is then sent to the control logic, which then controls the start of the next multiply-accumulate operation, until all the convolution operations are completed.
[0086] It should be noted that because the multiply-accumulate operation is split by the method in formula 2, the precision of the single memristor calculation array can be controlled according to the strength of the split (( L A , L B The length), and the best case can reduce the precision requirements of the memristor, ADC circuit and DAC circuit by nearly half.
[0087] In addition, in a preferred embodiment of the present application, the precision of the ReRAM convolution simplification implementation method provided by the present application can be further reduced by the method of nested loops, and the principle is as shown in Figure 3 Figure 3 It can be understood that Further reduce the device precision requirement architecture), in the process of high-low split of the input data A and the weight data B :
[0088] Each multiplication operation formula in the convolution to be operated can be split by nested loops to split the input data A and the weight data B m times; wherein,
[0089] m is an integer and ≥1.
[0090] In addition, preferably, in the process of splitting each multiplication operation formula in the convolution to be operated by nested loops:
[0091] Each multiplication operation formula is further split to further split the value of A H into A H1 and A L1 , B H is split into B H1 and B L1 ;
[0092] AgainA H1 、 A L1 、 B H1 、 B L1 Split, this cycle, until the operation precision of the convolution to be operated reaches the preset precision threshold.
[0093] The following is an example of the calculation part to illustrate the process of further improving the accuracy by using the nested loop method, as shown in The calculation part can be further reduced in accuracy requirement, and we will Figure 3 H The value of H1 and A L1 Split A H into A H1 and B H Split B H1 and B L1 This is equivalent to splitting the single memristor calculation array in the architecture in Figure 3 Again, this can further reduce the accuracy requirement of the memristor calculation array in Figure Three In theory, it can be reduced in a loop until it meets the accuracy requirement that the device itself can achieve.
[0094] It should be noted that the convolution simplification implementation method provided by the present application is based on a ReRAM (memristor) implementation, so in the actual calculation process, the weight data is the memristor calculation matrix data.
[0095] In addition, in the process of implementing the operation of the convolution to be operated based on the calculation results of each multiplication operation formula, the final operation result of the convolution to be operated needs to be completed by a preset adder based on the calculation results of each multiplication operation formula.
[0096] From the above specific embodiments, it can be seen that the ReRAM convolution simplification implementation method provided by the present application has the following advantages:
[0097] 1、The ReRAM convolution simplification implementation method provided by the present application can significantly improve the accuracy of ReRAM convolution operation by splitting the input data and weight data into high and low bits, and the ReRAM convolution operation is no longer limited by the accuracy requirement of ADC / DAC;
[0098] 2. The ReRAM convolution simplification implementation method provided by the application can perform cyclic high-low bit splitting on input data and weight data, so that the precision of ReRAM convolution operation is always improved to a saturation value, and therefore, the requirement for the resistance state precision of the memristor itself can be significantly reduced.
[0099] As described above with reference to Figure 2 and Figure 3 The ReRAM convolution simplification implementation method according to the application is described by way of example. However, those skilled in the art should understand that various improvements can be made to the ReRAM convolution simplification implementation method provided by the application without departing from the content of the application. Therefore, the protection scope of the application should be determined by the content of the appended claims.
Claims
1. A ReRAM convolution simplified implementation method, characterized in that, The application is applied to a compute-in-memory architecture, comprising: sequentially obtaining each multiplication operation formula in a convolution to be operated; splitting high and low bits of input data and weight data in each multiplication operation formula to convert each multiplication operation formula into a corresponding multiplication-addition operation simplified formula; sequentially calculating the calculation results of each multiplication operation formula in the convolution to be operated through the multiplication-addition operation simplified formula; implementing the operation of the convolution to be operated based on the calculation results of each multiplication operation formula; wherein the input data and the weight data in each multiplication operation formula are binary data; and splitting high and low bits of the input data and the weight data in each multiplication operation formula to convert each multiplication operation formula into a corresponding multiplication-addition operation simplified formula comprises: For any one multiplication formula, obtain its input data A and weight data B ; the result of the multiplication can be recorded as: ; performing high-low split on the input data A and the weight data B to split the input data A into high-bit input data A H and low-bit input data A L and split the weight data B into high-bit weight value B H and low-bit weight value B L ; wherein, A L the bit width of the high-bit input data L A , B L the bit width of the low-bit input data L B ; The input data after splitting A and the weight data B Substitute into the corresponding multiplication formula to get the corresponding multiplication and addition operation formula: 。 2. The ReRAM convolution simplification implementation method of claim 1, wherein For any one multiplication formula, the input data of matrix data type is first converted into input data of vector data type A , and then the input data A is split into high and low bits. A 3. The ReRAM convolution simplification implementation method of claim 2, wherein Input data of matrix data type is converted to input data of vector data type to get n input vector data: A Input data of matrix data type is converted to input data of vector data type to get n input vector data: A 1~ A n ; splitting high and low bits of each input vector data; The split n input vector data and the weight data are multiplied B The corresponding multiplication and addition operation simplification formula is obtained by substituting the corresponding multiplication operation formula. 。 4. The ReRAM convolution simplification implementation method of claim 3, wherein sequentially calculating the calculation results of each multiplication operation formula in the convolution to be operated through the multiplication-addition operation simplified formula comprises: splitting the multiplication-addition operation simplified formula into four multiplication operation simplified formulas; = X1, = X2, = X3, = X4 sequentially calculating the calculation results of each multiplication operation formula in the convolution to be operated based on the four multiplication operation simplified formulas.
5. The ReRAM convolution simplification implementation method of claim 4, wherein in response to the input data A and the weight data B being split into high and low bits: The multiplication operation formula in the convolution operation to be operated is split by means of nested loops, so as to split the input data A and the weight data B m times of high-low splitting; wherein, m is an integer and ≥ 1.
6. The ReRAM convolution simplification implementation method of claim 5, wherein in the process of circularly splitting each multiplication operation formula in the convolution to be operated through a nested loop, For each multiplication operation simplification formula is further split to A H value is further split into A H1 and A L1 B H is split into B H1 and B L1 ; Again A H1 , A L1 , B H1 , B L1 Split, and so on until the operation precision of the convolution to be operated reaches a preset precision threshold.
7. The ReRAM convolution simplification implementation method of any one of claims 1 to 6, wherein the weight data is a memristor calculation matrix data.
8. The ReRAM convolution simplification implementation method of claim 1, wherein in the process of implementing the operation of the convolution to be operated based on the calculation results of each multiplication operation formula, calculating the operation result of the convolution to be operated through an adder based on the calculation results of each multiplication operation formula.
Citation Information
Patent Citations
Convolution operation circuit and convolution operation method
WO2021081854A1