Storage and calculation integrated block multiplication and addition array

By employing a block multiply-accumulate array and a shift-accumulate circuit in the in-memory computing chip, the problem of long computation time in in-memory computing chips has been solved, achieving high-speed and efficiency improvement in vector dot product operations.

CN120909985AActive Publication Date: 2025-11-0758TH RES INST OF CETC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511026604.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-07
Estimated Expiration
2045-07-24

AI Technical Summary

Technical Problem

In existing in-memory computing chips, non-blocked in-memory multiply-accumulate circuits have long operation times and low efficiency, which cannot effectively improve the computing speed of AI algorithms and reduce power consumption.

Method used

By employing a storage-integrated block multiply-accumulate array, weight values ​​are programmed into storage units in blocks, and dot product operations between vectors and weights are performed using analog-to-digital conversion circuits and shift-accumulate circuits, reducing computation steps and time.

Benefits of technology

The vector dot product operation time of the in-memory computing chip is reduced by half, and the computing speed is doubled, realizing high-speed in-memory computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909985A_ABST
    Figure CN120909985A_ABST
Patent Text Reader

Abstract

The invention discloses a storage and calculation integrated block multiplication and addition array, and relates to the technical field of storage and calculation integrated circuit structure design, and the storage and calculation integrated block multiplication and addition array comprises the following steps: carrying out one-time vector point integration to form a four-row storage and calculation array to complete partial multiplication and addition, and obtaining four partial multiplication and addition results PP0, PP1, PP3 and PP3; the one-bit calculation result of each partial multiplication and addition result PPI (i = 0, 1, 2, 3) is realized by adopting an accumulation mode; the one-bit calculation result of each partial multiply-add result PPi is calculated in a configurable clock beat; the partial multiply-accumulate results PPI (i = 0, 1, 2, 3) are multiplied by 20, 24, 24 and 28 respectively, namely, the partial multiply-accumulate results PPI (i = 0, 1, 2, 3) are shifted leftwards by 0, 4, 4 and 8 bits respectively, and four shifted partial multiply-accumulate results PPSi (i = 0, 1, 2, 3) are obtained; the four partial multiply-accumulate results PPSi (i = 0, 1, 2, 3) are subjected to additive operation in a clock beat through a full adder. According to the block multiplication and addition circuit, the vector dot product operation time of storage and calculation integrated implementation can be reduced by half, the effective rate is doubled, and storage and calculation integrated high-speed calculation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of memory-computing integrated circuits, and particularly relates to a memory-computing integrated block multiplication-addition array. BACKGROUND

[0002] Artificial intelligence (AI) is mainly supported by neural network algorithms and hardware circuits, such as general-purpose computing chips CPU, etc. to complete large-scale multiplication-addition operations in neural network computation, but there are disadvantages of slow speed and large circuit power consumption. Therefore, in order to improve efficiency, AI algorithms need to run on dedicated AI chips.

[0003] The AI chip is mainly composed of neural network hardware circuits or accelerators inside. The main function of the neural network accelerator or circuit is to accelerate the convolution operation or matrix multiplication in the AI algorithm. When implementing matrix multiplication and convolution operation, the AI chip implemented by pure digital circuit uses adder and multiplier in the digital circuit; the larger the neural network scale, the more digital adders and multipliers needed inside the AI accelerator, resulting in large circuit area and large power consumption.

[0004] The memory-computing integrated chip uses analog signals such as current and resistance to complete multiplication and addition in the AI algorithm. The weight factor in multiplication is programmed and burned in the memory-computing unit in the form of charge or conductance. When performing convolution operation or matrix multiplication, there is no need to read the weight factor from the off-chip memory (such as DDR memory), but only one programming is needed to store the weight factor in the non-volatile memory device, which speeds up the operation and solves the problem of "memory wall".

[0005] The memory-computing integrated chip uses non-volatile memory devices inside, which has small device cell area and low power consumption, and can be deployed on a large scale. The power consumption of the memory-computing integrated chip is 1 to 2 orders of magnitude less than that of the AI chip implemented by digital circuit in the same area. At present, the memory-computing unit inside the memory-computing integrated chip is mainly implemented by SRAM, RRAM, MRAM and Flash or CAS unit. The non-blocked memory-computing integrated multiplication-addition circuit has long operation time and low efficiency. SUMMARY

[0006] The present application aims to provide a memory-computing integrated block multiplication-addition array to solve the problems in the background art.

[0007] To solve the above technical problems, the present application provides a memory-computing integrated block multiplication-addition array, comprising a plurality of rows of memory-computing arrays and an analog-to-digital conversion unit AD corresponding to each row of memory-computing arrays; wherein one intersection point of a row line and a column line represents a memory-computing unit;

[0008] The input data of the memory-compute array is a vector V=(E0, E1, E2,...), each element Ej of the vector is 8 bits wide, the high 4 bits are represented as Ej_H[7:4], and the value of each bit is d7, d6, d5, d4, respectively; the low 4 bits are represented as Ej_L[3:0], and the value of each bit is d3, d2, d1, d0, respectively; the weight value is Wj[7:0], the high 4 bits are represented as Wj[7:4], and the low 4 bits are represented as Wj[3:0]; j is a positive integer, representing the jth element or the jth weight; the dot product of the vector and the weight is represented as: E0×W0+E1×W1+E2×W2+....;

[0009] The external circuit programs the low 4 bits Wj[3:0] of the weight into the memory-compute unit at position (0, j), programs the high 4 bits Wj[7:4] of the weight into the memory-compute unit at position (1, j), programs the low 4 bits Wj[3:0] of the weight into the memory-compute unit at position (2, j), programs the high 4 bits Wj[7:4] of the weight into the memory-compute unit at position (3, j), and j represents the jth column of the memory-compute array; the weight W is represented as an analog signal in the form of charge quantity or conductance value;

[0010] According to the high 4 bits and the low 4 bits of the vector element Ej, the high 4 bits and the low 4 bits of the weight Wj, the dot product of the vector and the weight is further represented as:

[0011] Ej_L×Wj[3:0]+Ej_L×Wj[7:4]×2 4 +Ej_H×Wj[3:0]×2 4 +Ej_H×Wj[7:4]×2 8 .

[0012] In an embodiment, after the memory-compute unit is programmed with the weight, a unit vector Vu is input, the dot product of the unit vector and the weight is calculated, and the weight programming result of the memory-compute array is read back, including:

[0013] The first element E0 of the unit vector Vu is 1, that is, d7=0, d6=0, d5=0, d4=0, d3=0, d2=0, d1=0, d0=1; the second element E1 is 0, that is, d7=0, d6=0, d5=0, d4=0, d3=0, d2=0, d1=0, d0=0; the third element E2 is 0, that is, d7=0, d6=0, d5=0, d4=0, d3=0, d2=0, d1=0, d0=0; the multiplication and addition result of the unit vector and the weight is E0×W0+E1×W1+E2×W2=(d0×W0[3:0], d0×W0[7:4])=(W0[3:0], W0[7:4]), and the multiplication and addition values are read by the analog-to-digital conversion circuits AD0 and AD1 as res0 and res1, respectively.

[0014] Analog-to-digital conversion circuit ADi integrates each row current of the multiply-add array and converts it into a digital signal, and the multiply-add process propagates from the input vector element Ej to the output of ADi, which needs to be measured; i = 0, 1, 2, …;

[0015] The measurement is started with a 2-clock period increment method, and the measurement value is compared with the theoretical value. If the measurement result at the kth clock beat is close to or equal to the theoretical value, it is considered that the multiply-add time of the 1-bit data of the calculation vector and the weight is k clock beats.

[0016] The control circuit uses a first counter to count the clock beats, and generates an enable signal every k beats, and the registers are respectively sampled and saved.

[0017] In an embodiment, each bit of each element Ej of the vector Vu is multiplied and added, and the corresponding bit multiply-add results of multiple elements are accumulated, including:

[0018] In the control circuit, when the first counter counts to k, the second counter starts counting, and counts from 1 to 4; the registers R0, R1, R2, and R3 are cleared at this moment; the outputs res_bit0, res_bit1, res_bit2, and res_bit3 of the registers T0, T1, T2, and T3 are respectively left shifted by the counting value of the second counter, and four adders ADDER0, ADDER1, ADDER2, and ADDER3 are used to respectively complete four accumulation operations to obtain four partial multiply-add results in the first result: PP0(0), PP1(0), PP2(2), and PP3(3).

[0019] In an embodiment, the high and low parts of each element of the vector Vu are right shifted by 1 bit, the multiply-add of each bit is completed, and the shift accumulation circuit is used to implement, including:

[0020] Ej_H and Ej_L are right shifted by 1 bit, at this moment, d1 of each element is located at d0 bit, and d5 is located at d4 bit, the simulation multiply-add result reaches resk(k = 0, 1, 2, 3) after k periods, the registers Tk(k = 0, 1, 2, 3) sample the multiply-add results of each row respectively, the registers Rk(k = 0, 1, 2, 3) complete the multiply-add of each row, and four partial multiply-add results in the second result: PP0(1), PP1(1), PP3(1), and PP3(1) are obtained; the above process is repeated twice to complete the multiply-add of each bit of each element and the weight.

[0021] In an embodiment, the fourth partial multiply-accumulate result PP0(3) is not shifted, PP1(3) is left shifted by 4 bits; PP2(3) is left shifted by 4 bits, and PP3(3) is left shifted by 8 bits; the results obtained after left shifting of each partial multiply-accumulate are PPS0, PPS1, PPS2 and PPS3 respectively; the multiply-accumulate results PPS0, PPS1, PPS2 and PPS3 after shifting are added together by a full adder FA and stored in a register M.

[0022] The application provides a memory-compute integrated block multiply-add array, which can reduce the vector dot product operation time of the memory-compute integration by half, effectively improve the speed by one time, and realize high-speed computation of the memory-compute integration. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 is a structural schematic diagram of the memory-compute integrated block array provided by the application.

[0024] Figure 2 is a structural schematic diagram of a shift-add circuit for multiply-accumulate results of the block array provided by the application.

[0025] Figure 3 is a data timing schematic diagram provided by the application.

[0026] Figure 4 is a weight programming adjustment method flowchart. DETAILED DESCRIPTION

[0027] The application provides a memory-compute integrated block multiply-add array, which can reduce the vector dot product operation time of the memory-compute integration by half, effectively improve the speed by one time, and realize high-speed computation of the memory-compute integration.

[0028] Figure 1 The memory-compute integrated block array in the figure is composed of 4 rows of memory-compute arrays and 4 analog-to-digital conversion units AD, and can complete one multiply-add of 8-bit wide data and 8-bit wide weights. Figure 1There are four row lines and three column lines, respectively, the 0th row line, the 1st row line, the 2nd row line, the 3rd row line, the 0th column line, the 1st column line, and the 2nd column line, so there are 12 storage and calculation units in total. The storage and calculation unit is programmed to write weights, and the weight value is in the form of an analog quantity, such as a charge quantity or a conductance value. When a signal is applied to the input end of the column: if the signal is 1, the storage and calculation unit of the column is opened to contribute a current of the weight value W to the row line; if the input of the corresponding column is 0, the storage and calculation unit of the column is closed, and the current contributed to the row line is 0. Therefore, the current value on the 0th row line is W0[3:0] x di(j) + W1[3:0] x di(j) + W2[3:0] x di(j), where i represents the i-th bit of the vector element, j represents the j-th column of the array, and so on.

[0029] The input data of the storage and calculation array is a vector V=(E0, E1, E2, E3), each element Ej[7:0] of the vector is 8 bits wide, and the weight value is Wj[7:0], j(=0, 1, 2, 3) represents the j-th element or the j-th weight. The dot product of the vector and the weight is represented as:

[0030] E0 x W0 + E1 x W1 + E2 x W2 (1)

[0031] An element E of the vector can be represented as (d7, d6, d5, d4, d3, d2, d1, d0), so the multiplication and accumulation process in equation (1) can be decomposed as follows:

[0032] (d7, d6, d5, d4, d3, d2, d1, d0)0 x W0 + (d7, d6, d5, d4, d3, d2, d1, d0)1 x W1 + (d7, d6, d5, d4,

[0033] d3, d2, d1, d0)2 x W2 (2)

[0034] Continue to decompose equation (2) to obtain the following formula:

[0035] (d70 x W0 + d71 x W1 + d72 x W2, d60 x W0 + d61 x W1 + d62 x W2,

[0036] d50 x W0 + d51 x W1 + d52 x W2,

[0037] d40 x W0 + d41 x W1 + d42 x W2, d30 x W0 + d31 x W1 + d32 x W2,

[0038] d20 x W0 + d21 x W1 + d22 x W2,

[0039] d10 x W0 + d11 x W1 + d12 x W2, d00 x W0 + d01 x W1 + d02 x W2) (3)

[0040] Continuing to decompose equation (3), the following equation is obtained:

[0041] (d70 x W0 + d71 x W1 + d72 x W2) x 2 7 + (d60 x W0 + d61 x W1 + d62 x W2) x 2 6 +

[0042] (d50 x W0 + d51 x W1 + d52 x W2) x 2 5 + (d40 x W0 + d41 x W1 + d42 x W2) x 2 4 +

[0043] (d30 x W0 + d31 x W1 + d32 x W2) x 2 3 + (d20 x W0 + d21 x W1 + d22 x W2) x 2 2 + (d10 x W0 + d11 x W1 + d12 x W2) x 2 1 + (d00 x W0 + d01 x W1 + d02 x W2) x 2 0 (4)

[0044] The calculation process in equation (4) is actually a process of shift multiplication and accumulation. The weight W is multiplied and accumulated with each bit of the vector element. The vector element is 8 bits, so 8 times are needed to complete the multiplication and accumulation of the weight W and the vector element E. In order to reduce the operation time, the calculation process in equation (4) is divided into 4 blocks, each block calculating the multiplication and accumulation of a part of the vector element and a part of the weight.

[0045] The high 4-bit part of the vector element is represented as Ej_H[7:4], and the values on each bit are d7, d6, d5, and d4, respectively. The low 4-bit part is represented as Ej_L[3:0], and the values on each bit are d3, d2, d1, and d0, respectively. The weight value is Wj[7:0], the high 4-bit part is represented as Wj[7:4], and the low 4-bit part is represented as Wj[3:0]. Equation (4) is further decomposed to obtain the following calculation formula

[0046] (d70 x W0[7:4] + d71 x W1[7:4] + d72 x W2[7:4]) x

[0047] 2 7 + (d60 x W0[7:4] + d61 x W1[7:4] + d62 x W2[7:4]) x 2 6 +

[0048] (d50xW0[7:4]+d51xW1[7:4]+d52xW2[7:4])x

[0049] 2 5 +(d40xW0[7:4]+d41xW1[7:4]+d42xW2[7:4])x2 4 (5)

[0050] (d70xW0[3:0]+d71xW1[3:0]+d72xW2[3:0])x

[0051] 2 7 +(d60xW0[3:0]+d61xW1[3:0]+d62xW2[3:0])x2 6 +

[0052] (d50xW0[3:0]+d51xW1[3:0]+d52xW2[3:0])x

[0053] 2 5 +(d40xW0[3:0]+d41xW1[3:0]+d42xW2[3:0])x2 4 (6)

[0054] (d30xW0[7:4]+d31xW1[7:4]+d32xW2[7:4])x2 3

[0055] (d20xW0[7:4]+d21xW1[7:4]+d22xW2[7:4])x2 2 +

[0056] (d10xW0[7:4]+d11xW1[7:4]+d12xW2[7:4])x

[0057] 2 1 +(d00xW0[7:4]+d01xW1[7:4]+d02xW2[7:4])x2 0 (7)

[0058] (d30xW0[3:0]+d31xW1[3:0]+d32xW2[3:0])x2 3

[0059] (d20xW0[3:0]+d21xW1[3:0]+d22xW2[3:0])x2 2 +

[0060] (d10xW0[3:0]+d11xW1[3:0]+d12xW2[3:0])x

[0061] 2 1 +(d00xW0[3:0]+d01xW1[3:0]+d02xW2[3:0])x2 0 (8)

[0062] From formula (5), (6), (7) and (8), it can be seen that each item calculation process only needs 4 times of accumulation, compared with 8 times of accumulation in formula (4), the time needed is reduced by half.

[0063] The process of calculating multiplication and accumulation in formula (4) needs 1 row of storage and calculation array, and the process of calculating multiplication and accumulation in formula (5), (6), (7) and (8) needs 4 rows of storage and calculation array. Therefore, the storage and calculation integrated block multiplication and addition array of the application needs 4 rows of storage and calculation array to complete the dot product of an 8-bit wide vector and an 8-bit wide weight, as shown in Figure 1 .

[0064] The current of each row in the storage and calculation integrated block array is converted into a digital signal by an analog-to-digital conversion unit AD, and the bit width is 8 bits. The multiplication and addition calculation of each bit of the vector unit and the weight is called one multiplication and addition. The result resk (k represents the row number, taking 0, 1, 2, 3) of each multiplication and addition is accumulated after being shifted, and exists in the temporary register Rk (k represents the row number, taking 0, 1, 2, 3) of the circuit. The results PPi (i = 0, 1, 2, 3) of the fourth multiplication and addition are respectively left shifted by 0 bit, 4 bits, 4 bits and 8 bits, and then full addition is performed to obtain the final dot product result mac of the vector and the weight, as shown in Figure 2 .

[0065] Please refer to Figure 4 , which shows a weight programming adjustment method provided by an embodiment of the application, in combination with Figure 1 , Figure 2 , Figure 3 , the method comprises:

[0066] (1) Step 101, the values on each bit in the high 4-bit part Ej_H[7:4] of the input data Ej of the storage and calculation array are respectively d7, d6, d5 and d4, and the values on each bit in the low 4-bit part Ej_L[3:0] are respectively d3, d2, d1 and d0. The weight value is Wj[7:0], the high 4-bit part is represented as Wj[7:4], and the low 4-bit part is represented as Wj[3:0]. The external circuit programs (writes) the low 4 bits Wj[3:0] of the weight to the storage and calculation unit at position (0, j), programs the high 4 bits Wj[7:4] of the weight to the storage and calculation unit at position (1, j), programs the low 4 bits Wj[3:0] of the weight to the storage and calculation unit at position (2, j), and programs the high 4 bits Wj[7:4] of the weight to the storage and calculation unit at position (3, j). Wherein j represents the jth column of the storage and calculation array, and the weight W is represented as an analog signal in the form of charge quantity or conductance value.

[0067] (2) Step 102, the memory-computing array generates changes from the input data E di (i = 0, 1, 2, 3, 4, 5, 6, 7) to the ADi (i = 0, 1, 2, 3) output stable digital signal results, which requires a certain delay, expressed in the number of clock beats. Therefore, in order to obtain the correct multiplication and addition result, the delay, i.e. the accurate number of clock beats, needs to be measured. Set the input element Ej to 1, i.e. d7 = 0, d6 = 0, d5 = 0, d4 = 0, d3 = 0, d2 = 0, d1 = 0, d0 = 1; the second element E1 = 0, i.e. d7 = 0, d6 = 0, d5 = 0, d4 = 0, d3 = 0, d2 = 0, d1 = 0, d0 = 0; the third element E2 = 0, i.e. d7 = 0, d6 = 0, d5 = 0, d4 = 0, d3 = 0, d2 = 0, d1 = 0, d0 = 0.

[0068] The analog-digital conversion circuit AD0 integrates the current of the 0th row of the multiplication and addition array and converts it into a digital signal, and the time for the input vector element E0 to propagate to the output of AD0 is measured: the measurement starts with 2 clock periods and increases by 1, and the measured value is compared with the theoretical value. If the result measured at the kth (for example, k = 3) clock beat is close to or equal to the theoretical value, it is considered that the delay of the memory-computing multiplication and addition array and the analog-digital conversion module AD is k clock beats.

[0069] (3) Step 103, the control circuit uses the counter timer1 to count the clock beats, and generates an enable signal every k beats, and the registers T0, T1, T2, T3 sample and save the output results res0, res1, res2 and res3 of the corresponding analog-digital conversion circuits AD0, AD1, AD2 and AD3, respectively.

[0070] When timer1 counts to k, the registers R0, R1, R2, R3 are cleared at this moment, and the counter timer2 starts counting from 1 to 4. The outputs res_bit0, res_bit1, res_bit2, res_bit3 of T0, T1, T2, T3 are left shifted by timer2 bits, i.e. left shifted by the count value of timer2, and four accumulations are completed through the four adders ADDER0, ADDER1, ADDER2, ADDER3, respectively, to obtain four partial multiplication and addition results in the first result: PP0(0), PP1(0), PP2(2) and PP3(3), where i = 0, 1, 2, 3, as shown in Figure 2 and Figure 3 .

[0071] (4) Step 104, Ej_H and Ej_L (j=0, 1, 2) are right shifted by one bit, at this time, d1 of Ej_L is in d0 bit, d5 of Ej_H is in d4 bit, after k cycles, the multiplication and addition result reaches res (i=0, 1, 2, 3), registers Ti (i=0, 1, 2, 3) sample the multiplication and addition result of each row respectively, registers Ri (i=0, 1, 2, 3) complete multiplication and accumulation of each row, and four partial multiplication and addition results are obtained in the second time: PP0(1), PP1(1), PP3(1) and PP3(1). The above process is repeated twice, and each bit of each element is multiplied and accumulated with the weight.

[0072] (5) Step 105, the fourth time partial multiplication and accumulation result PP0(3) is not shifted, PP1(3) is left shifted by 4 bits to complete the process of ×2 4 , PP2(3) is left shifted by 4 bits to complete the process of ×2 4 , and PP3(3) is left shifted by 8 bits to complete the process of ×2 8 . The results obtained after left shifting of each partial multiplication and accumulation are PPS0, PPS1, PPS2 and PPS3 respectively.

[0073] (6) Step 106, the multiplication and accumulation results PPS0, PPS1, PPS2 and PPS3 after shifting are added together by using full adder FA, and are stored in register M, and the vector dot product result mac is obtained, that is, Ej_L×Wj[3:0]+Ej_L×Wj[7:4]×2 4 +Ej_H×Wj[3:0]×2 4 +Ej_H×Wj[7:4]×2 8 .

[0074] In order to make the description simple, the decomposition of each weight value in the above embodiment and all possible combinations of different weight programming of the calculation unit are not described, however, as long as there is no contradiction in these combinations, it should be considered that it is within the scope of the present application.

[0075] The above description is only a description of the preferred embodiment of the present application, and is not any limitation on the scope of the present application, any modification and modification of the above disclosure made by a person skilled in the art of the present application is within the protection scope of the claims.

Claims

1. A memory compute integrated block multiply-add array, comprising: The memory-computing array comprises four rows of memory-computing arrays and an analog-to-digital conversion unit AD corresponding to each row of memory-computing arrays; one intersection of a row line and a column line represents one memory-computing unit; The input data of the memory-computing array is a vector V=(E0, E1, E2,...), each element Ej of the vector has a bit width of 8 bits, the high 4 bits are represented as Ej_H[7:4], and the value of each bit is d7, d6, d5, d4, respectively; the low 4 bits are represented as Ej_L[3:0], and the value of each bit is d3, d2, d1, d0, respectively; The weight value is Wj[7:0], the high 4 bits are represented as Wj[7:4], and the low 4 bits are represented as Wj[3:0]; j is a positive integer, representing the jth element or the jth weight; the dot product of the vector and the weight is represented as: E0×W0+E1×W1+E2×W2+....; An external circuit programs the low 4 bits Wj[3:0] of the weight into the memory-computing unit at position (0, j), programs the high 4 bits Wj[7:4] of the weight into the memory-computing unit at position (1, j), programs the low 4 bits Wj[3:0] of the weight into the memory-computing unit at position (2, j), programs the high 4 bits Wj[7:4] of the weight into the memory-computing unit at position (3, j), and j represents the jth column of the memory-computing array; the weight W is represented as an analog signal in the form of a charge quantity or a conductance value; According to the high 4 bits and the low 4 bits of the vector element Ej, the high 4 bits and the low 4 bits of the weight Wj, the dot product of the vector and the weight is further represented as: Ej_L x Wj[3:0] + Ej_L x Wj[7:4] x 2 4 +Ej_H x Wj[3:0] x 2 4 +Ej_H x Wj[7:4] x 2 8 .

2. The in-memory computing block-multiply-add array of claim 1, wherein, After the memory-computing unit is programmed with the weight, a unit vector Vu is input, the dot product of the unit vector and the weight is calculated, and the weight programming result of the memory-computing array is read back, including: The first element E0 of the unit vector Vu is 1, that is, d7=0, d6=0, d5=0, d4=0, d3=0, d2=0, d1=0, and d0=1; the second element E1 is 0, that is, d7=0, d6=0, d5=0, d4=0, d3=0, d2=0, d1=0, and d0=0; the third element E2 is 0, that is, d7=0, d6=0, d5=0, d4=0, d3=0, d2=0, d1=0, and d0=0; the multiplication and addition result of the unit vector and the weight is E0×W0+E1×W1+E2×W2=(d0×W0[3:0], d0×W0[7:4])=(W0[3:0], W0[7:4]), and the multiplication and addition values are read by the analog-to-digital conversion circuits AD0 and AD1 as res0 and res1, respectively; The analog-to-digital conversion circuit ADi integrates and converts the current of each row of the multiplication and addition array into a digital signal, and the multiplication and addition process propagates from the input vector element Ej to the output of the ADi, and the time period needs to be measured; i=0, 1, 2,...; The measurement is started from the 2nd clock cycle and is incremented, and the measurement value is compared with the theoretical value; if the measurement result at the kth clock beat is close to or equal to the theoretical value, the multiplication and addition time of the 1-bit data of the calculation vector and the weight is considered to be k clock beats; The control circuit counts the clock beats using a first counter, and generates an enable signal every k beats, and the registers sample and save respectively.

3. The in-memory computing block-multiply-add array of claim 2, wherein, Each bit of each element Ej of the vector Vu is multiplied and added, and the corresponding bit multiplication and addition results of multiple elements are accumulated, including: In the control circuit, when the first counter counts to k, the second counter starts counting, and counts from 1 to 4; the registers R0, R1, R2 and R3 are cleared at this moment; the outputs res_bit0, res_bit1, res_bit2 and res_bit3 of the registers T0, T1, T2 and T3 are respectively left shifted by the counting value of the second counter, and four accumulations are respectively completed through four adders ADDER0, ADDER1, ADDER2 and ADDER3, to obtain four partial multiplication and addition results in the first result: PP0(0), PP1(0), PP2(2) and PP3(3).

4. The in-memory computing block-multiply-add array of claim 3, wherein, The high and low parts of each element of the vector Vu are right shifted by 1 bit, and the multiplication and addition of each bit are completed, which is realized through a shift accumulation circuit, including: Ej_H and Ej_L are both right shifted by 1 bit, at this moment, d1 of each element is located at d0 bit, and d5 is located at d4 bit, and after k cycles of analog multiplication and addition, the simulation multiplication and addition result reaches resk(k=0,1,2,3), the registers Tk(k=0,1,2,3) sample the multiplication and addition results of each row respectively, and the registers Rk(k=0,1,2,3) complete the multiplication and accumulation of each row, to obtain four partial multiplication and addition results in the second result: PP0(1), PP1(1), PP3(1) and PP3(1); the above process is repeated twice, and each bit of each element is multiplied and accumulated with the weight.

5. The in-memory computing block-multiply-add array of claim 4, wherein, The fourth result PP0(3) of the partial multiplication and accumulation does not perform shift, PP1(3) is left shifted by 4 bits; PP2(3) is left shifted by 4 bits, and PP3(3) is left shifted by 8 bits; the results obtained after left shifting of each partial multiplication and accumulation are PPS0, PPS1, PPS2 and PPS3 respectively; The shifted multiplication and accumulation results PPS0, PPS1, PPS2 and PPS3 are added together through a full adder FA and stored in a register M.

Citation Information

Patent Citations

  • Parallel multiply-add device based on SRAM

    CN110750232A

  • Storage and calculation integrated circuit, and data operation method based on storage and calculation integrated circuit

    CN112558917A

  • Convolution operation circuit and operation method thereof

    CN113869498A

  • Neural network acceleration method, neural network accelerator, chip and electronic equipment

    CN116402106A

  • Multiplier and chip

    CN116522967A