Multi-bit data in-memory computing array structure, SRAM and electronic devices

By designing a multi-bit data in-memory computing array structure, adjusting the pulse width using the voltage-controlled delay unit, realizing multiplication and multiplication accumulation calculation of five-bit input and weights, the problem of limited single-bit calculation accuracy in the existing technology is solved, and the inference accuracy and efficiency of the AI ​​system are improved.

CN119669147BActive Publication Date: 2025-05-13ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510201815.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-13
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing nonvolatile in-memory computing circuits only support multiplication in-memory computing of single-bit input and weights, resulting in limited system-level inference accuracy, limiting the development of AI technology.

Method used

A multi-bit data in-memory computing array structure is designed, including a multi-column computing array and a voltage-controlled delay unit. By adjusting the pulse width of the reference signal, multiplication and multiplication accumulation calculation of five-bit input and weight is realized.

Benefits of technology

Multi-bit calculation with five-bit input and weight is realized, which improves system-level inference accuracy and efficiency, and solves the finite problem of single-bit calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669147B_ABST
    Figure CN119669147B_ABST
Patent Text Reader

Abstract

The present application relates to a multi-bit data in-memory computing array structure, an SRAM and an electronic device, wherein the multi-bit data in-memory computing array structure is used to determine the multiplication and accumulation results of a five-bit input and a five-bit weight, and includes multiple columns of multi-bit data in-memory computing arrays. The core of the multi-bit data in-memory computing array is to characterize the calculation result by the pulse width adjustment amount of the reference signal. Since the pulse width adjustment amount can be accumulated, when it is necessary to implement the multiplication and accumulation calculation of the five-bit input and the five-bit weight, it is only necessary to combine the multiple columns of the multi-bit data in-memory computing array in the form of rows, and the reference signal output by each voltage-controlled delay circuit in the previous column is the reference signal received by the corresponding voltage-controlled delay circuit in the next column, which solves the problem that the current non-volatile in-memory computing circuit usually only supports the multiplication and accumulation in-memory calculation of single-bit input and weight, and can only provide limited system-level reasoning accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of integrated circuits, and in particular to a multi-bit data storage computing array structure, SRAM and electronic equipment. Background Art

[0002] Computing In Memory (CIM) breaks the von Neuma architecture of traditional computers and embeds computing circuits into memory, integrating storage and computing, thereby greatly reducing data migration and memory access consumption. Nonvolatile Computing In Memory is extremely advantageous in micro AI devices that require nonvolatile data storage and low power consumption and battery power. Current nonvolatile in-memory computing technology solutions support binary neural networks (BNNs) or binary weight networks (BWNs), which to a certain extent reduces storage requirements and improves energy efficiency. However, BNNs and BWNs are only applicable to simple networks, which usually only support multiplication of single-bit input and weights and multiply and accumulate (MAC) in-memory computing. Therefore, when applied to complex applications, they can only provide limited system-level reasoning accuracy, which limits the further development of AI technology.

[0003] Currently, no effective solution has been proposed to the problem that current non-volatile in-memory computing circuits usually only support in-memory multiplication of single-bit input and weight, and can only provide limited system-level reasoning accuracy. Summary of the invention

[0004] The present invention provides a multi-bit data in-memory computing array structure, SRAM and electronic device to solve the problem that current non-volatile in-memory computing circuits generally only support in-memory multiplication calculations of single-bit inputs and weights, and can only provide limited system-level reasoning accuracy.

[0005] In a first aspect, the present invention provides a multi-bit data in-memory computing array for determining a five-bit input IN MSB IN3IN2IN1IN0 and five-bit weight W MSB The multiplication calculation result of W3W2W1W0 includes a first calculation circuit, four second calculation circuits and four groups of voltage-controlled delay units corresponding to different bits;

[0006] The output of the first calculation circuit is , the outputs of the four second calculation circuits are A2, B2, A1 and B1 respectively;

[0007] when When, A2 = VSS, B2 = IN3IN2, A1 = VSS, B1 = IN1IN0;

[0008] when When, A2 = IN3IN2, B2 = VDD, A2 = IN1IN0, B2 = VDD;

[0009] The control input C of each group of voltage-controlled delay units is the corresponding bit weight and includes a high-order voltage-controlled delay circuit and a low-order voltage-controlled delay circuit. The two voltage-controlled inputs A and B of the high-order voltage-controlled delay circuit are A2 and B2 respectively, and the two voltage-controlled inputs A and B in the low-order voltage-controlled delay circuit are A1 and B1 respectively. The number of unit adjustment amounts contained in the sum of four times the pulse width adjustment amount of the high-order voltage-controlled delay circuit for the reference signal and the pulse width adjustment amount of the low-order voltage-controlled delay circuit for the reference signal is the calculation result of the corresponding bit;

[0010] When the control input C is 0, the voltage-controlled delay circuit maintains the pulse width of the reference signal;

[0011] When the control input C is 1, the voltage-controlled delay circuit adjusts the pulse width of the reference signal and the adjustment strategy is: the pulse width of the reference signal is linearly positively correlated with the voltage-controlled input A and linearly negatively correlated with the voltage-controlled input B.

[0012] In a second aspect, the present invention provides a multi-bit data in-memory calculation array structure for determining a five-bit input IN MSB IN3IN2IN1IN0 and five-bit weight W MSB The multiplication and accumulation result of W3W2W1W0 includes a plurality of columns of the multi-bit data in-memory calculation array as described in the first aspect;

[0013] The reference signal output by each voltage-controlled delay circuit in the previous column is the reference signal received by the corresponding voltage-controlled delay circuit in the next column.

[0014] In a third aspect, the present invention provides an SRAM that uses the multi-bit data in-memory computing array structure described in the second aspect to implement multiplication and accumulation calculations of five-bit inputs and five-bit weights.

[0015] In a fourth aspect, the present invention provides an electronic device, comprising a memory and a processor, wherein the memory is the SRAM described in the third aspect.

[0016] Compared with the related art, the multi-bit data in-memory computing array provided by the present invention can realize the multiplication calculation of five-bit input and weight, and supports the formation of a multi-bit data in-memory computing array structure in a combined form to realize the multiplication and accumulation calculation of five-bit input and five-bit weight, which can provide greater system-level reasoning accuracy and efficiency, and solves the problem that the current non-volatile in-memory computing circuit usually only supports the in-memory multiplication and accumulation calculation of single-bit input and weight, and can only provide limited system-level reasoning accuracy.

[0017] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a schematic diagram of a multi-bit data in-memory computing array structure provided in this embodiment;

[0019] Figure 2 is a schematic diagram of a multi-bit data in-memory computing array provided in this embodiment;

[0020] Figure 3 is a schematic diagram of a first calculation circuit provided in this embodiment;

[0021] Figure 4 is a schematic diagram of a local read-write module provided in this embodiment;

[0022] Figure 5 is a schematic diagram of a second calculation circuit provided in this embodiment;

[0023] Figure 6 is a schematic diagram of a voltage-controlled delay circuit provided in this embodiment. DETAILED DESCRIPTION

[0024] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.

[0025] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by people with general skills in the technical field to which this application belongs. The words "one", "a", "a", "the", "these" and the like in this application do not indicate a quantitative limitation, and they may be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" may mean: A exists alone, A and B exist at the same time, and B exists alone. Generally, the character " / " indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.

[0026] In this embodiment, a multi-bit data in-memory calculation array structure is provided to determine the five-bit input IN MSB IN3IN2IN1IN0 and five-bit weight W MSB The multiplication and accumulation result of W3W2W1W0.

[0027] Reference Figure 1 The multi-bit data in-memory computing array structure includes multiple columns of multi-bit data in-memory computing arrays. The reference signal output by any voltage-controlled delay circuit in the previous column is the reference signal received by the corresponding voltage-controlled delay circuit in the next column.

[0028] The multi-bit data in-memory calculation array provided in this embodiment is used to determine the five-bit input IN MSB IN3IN2IN1IN0 and five-bit weight W MSB The multiplication calculation result of W3W2W1W0 passes the multiplication calculation result of the previous column of multi-bit data in-memory calculation array to the next column of multi-bit data in-memory calculation array, thereby realizing the accumulation calculation between the multiplication calculation results.

[0029] Reference Figure 2Specifically, the multi-bit data memory calculation array provided in this embodiment includes a first calculation circuit, four second calculation circuits, and four groups of voltage-controlled delay units corresponding to different bits, and for storing W respectively. MSB In the multi-bit input and weight, IN3IN2IN1IN0 is a four-bit data input, W3W2W1W0 is a four-bit data weight, IN MSB is the symbol input, indicating the symbol of IN3IN2IN1IN0, W MSB is the symbol weight, indicating the sign of W3W2W1W0.

[0030] The output OUT of the first calculation circuit MSB for , the outputs of the four second calculation circuits are A2, B2, A1 and B1 respectively. When A2 = VSS, B2 = IN3IN2, A1 = VSS, B1 = IN1IN0; when When, A2 = IN3IN2, B2 = VDD, A1 = IN1IN0, B1 = VDD.

[0031] The first calculation circuit is used to realize the calculation between the symbol input and the symbol weight, and the first circuit constitutes the symbol operation module, and the second calculation circuit is used to realize the calculation between the data input and the symbol weight. Calculation between.

[0032] Reference Figure 3 Specifically, in this embodiment, each group of storage arrays includes multiple 6T storage cells, and the first storage node and the second storage node of the 6T storage cell are connected to the local bit line LBL and the local bit line LBLB respectively. When the first storage node and the second storage node are 1 and 0 respectively, it means that the stored weight is 1; when the first storage node and the second storage node are 0 and 1 respectively, it means that the stored weight is 0. The first calculation circuit includes a PMOS tube P1 and a PMOS tube P2, the sources of the PMOS tube P1 and the PMOS tube P2 are connected to the local bit line LBL and the local bit line LBLB respectively, and the gates of the PMOS tube P1 and the PMOS tube P2 are connected to the symbol input IN respectively. MSB And symbol input IN MSB The drains of the PMOS tube P1 and the PMOS tube P2 are connected and form the output node of the first calculation circuit, which is used to output The first computing circuit is used to store W MSB The memory array shares the same pair of local bit lines LBL and local bit lines LBLB.

[0033] For example, the word line WL of row 0 is turned on and W MSB=1 as an example: when IN MSB = 0, P1 is turned on, P2 is turned off, OUT MSB =1, the product of the two is negative; when IN MSB = 1, P2 is turned on, P1 is turned off, OUT MSB =0, the product of the two is positive; when the word line WL of row 0 is turned on and W MSB =0 as an example: when IN MSB = 0, P1 is turned on, P2 is turned off, OUT MSB =0, the product of the two is a positive value; when IN MSB = 1, P2 is turned on, P1 is turned off, OUT MSB =1, then the product of the two is a negative value.

[0034] It should be noted that the first calculation circuit is essentially a symbol input IN MSB and symbol weight W MSB Therefore, in other embodiments, other structures of XOR circuits can also be used to implement the symbol input IN MSB and symbol weight W MSB XOR calculation.

[0035] It should be noted that the storage array is also equipped with basic local read and write modules. Figure 4 In this embodiment, the local read / write module includes NMOS tubes N1 and N2. The gate of NMOS tube N1 is connected to the horizontal word line HWL, the drain is connected to the local bit line LBL, and the source is connected to the global bit line GBL; the gate of NMOS tube N2 is connected to the horizontal word line HWL, the drain is connected to the local bit line LBLB, and the source is connected to the global bit line GBLB. When the horizontal word line HWL and the word line WL of the n-th row 6T storage unit are turned on, the write signal is transmitted to the local bit line LBL / LBLB through the global bit line GBL / GBLB, and then the data is written into the n-th row 6T storage unit.

[0036] The four second calculation circuits constitute an input signal column channel module for processing data input. Specifically, the four-bit data input IN3IN2IN1IN0 is divided into two groups of 2-bit signals, IN3IN2 and IN1IN0, and the two groups of 2-bit signals are respectively used as inputs of different second calculation circuits. Among them, the second calculation circuit receiving IN3IN2 is defined as a high-weight calculation circuit, and the second calculation circuit receiving IN1IN0 is defined as a low-weight calculation circuit.

[0037] Specifically, in this embodiment, the second calculation circuit includes a first signal channel and a second signal channel, the outputs of the first signal channel and the second signal channel are connected to form the output of the second calculation circuit, and the first signal channel is The second signal channel is turned on when in the second calculation circuit with an output of A2, the inputs of the first signal channel and the second signal channel are IN3IN2 and VSS respectively; in the second calculation circuit with an output of B2, the inputs of the first signal channel and the second signal channel are VDD and IN3IN2 respectively; in the second calculation circuit with an output of A1, the inputs of the first signal channel and the second signal channel are IN1IN0 and VSS respectively; in the second calculation circuit with an output of B1, the inputs of the first signal channel and the second signal channel are VDD and IN1IN0 respectively.

[0038] Each second calculation circuit includes two signal channels, and the two signal channels are respectively connected to different signals. The output of the first calculation circuit The conduction states of the two signal channels can be controlled so that the second calculation circuit can output different signals, that is, when When A2 = VSS, B2 = IN3IN2, A1 = VSS, B1 = IN1IN0; when When, A2 = IN3IN2, B2 = VDD, A1 = IN1IN0, B1 = VDD.

[0039] For the specific structure of the signal channel, any transmission gate structure can be used.

[0040] Reference Figure 5 For example, in this embodiment: the second calculation circuit with output A2 includes: NMOS tube N3, PMOS tube P3, NMOS tube N4 and PMOS tube P4, the gate of NMOS tube N3 is connected to , the drain is connected to the data input IN3IN2 and the source is used to output A2, the gate of the PMOS tube P3 is connected The reverse signal, drain connected to data input IN3IN2 and source used to output A2, the gate of NMOS tube N4 connected The reverse signal, drain connected to VSS and source output A2, the gate of PMOS tube P4 connected , the drain is connected to the ground VSS and the source is output A2; the second calculation circuit with the output B2 includes: NMOS tube N5, PMOS tube P5, NMOS tube N6 and PMOS tube P6, NMOS tube N5 , the drain is connected to the power supply VDD and the source is used to output B2, the gate of the PMOS tube P5 is connected The reverse signal of the NMOS tube N6 is connected to the drain of the power supply VDD and the source of the NMOS tube N6. The reverse signal, the drain is connected to the data input IN3IN2 and the source is used to output B2, the gate of the PMOS tube P6 is connected , the drain is connected to the data input IN3IN2 and the source is used to output B2; the second calculation circuit with the output A1 includes NMOS tube N7, PMOS tube P7, NMOS tube N8 and PMOS tube P8, the gate of NMOS tube N7 is connected , the drain is connected to the data input IN1IN0 and the source is used to output A1, the gate of the PMOS tube P7 is connected The reverse signal, drain connected to data input IN1IN0 and source used to output A1, the gate of NMOS tube N8 connected The reverse signal, drain connected to VSS and source used to output A1, the gate of PMOS tube P8 connected , the drain is connected to VSS and the source is used to output A1; the second calculation circuit with output B1 includes NMOS tube N9, PMOS tube P9, NMOS tube N10 and PMOS tube P10, the gate of NMOS tube N9 is connected to , the drain is connected to the power supply VDD and the source is used to output B1, the gate of the PMOS tube P9 is connected The reverse signal, drain connected to power supply VDD and source used to output B1, the gate of NMOS tube N10 connected The reverse signal and drain of the PMOS tube P10 are connected to the data input IN1IN0 and the source is used to output B1. , the drain is connected to the data input IN1IN0 and the source is used for output B1.

[0041] For the above-mentioned second calculation circuit, MOS transistors such as N3~N6, P3~P6 corresponding to data input IN3IN2 constitute a high-weight calculation circuit, and the output signals are A2 and B2; MOS transistors such as N7~N10, P7~P10 corresponding to data input IN1IN0 constitute a low-weight calculation circuit, and the output signals are A1 and B1.

[0042] When data input IN3IN2 = 00, the data input voltage value is V 00 ; When data input IN3IN2 = 01, the data input voltage value is V 01 ; When data input IN3IN2 = 10, the data input voltage value is V 10 ; When data input IN3IN2 = 11, the data input voltage value is V 11 ; The same applies when data is input into IN1IN0.

[0043] when When: in the high-weight calculation circuit, the transmission gate composed of P4, N4 and P6, N6 is turned on, the transmission gate composed of P3, N3 and P5, N5 is cut off, A2 = VSS, B2 = the voltage value of IN3IN2; in the low-weight calculation circuit, the transmission gate composed of P8, N8 and P10, N10 is turned on, the transmission gate composed of P7, N7 and P9, N9 is cut off, A1 = VSS, B1 = the voltage value of IN1IN0.

[0044] when When: in the high-weight calculation circuit, the transmission gate composed of P3, N3 and P5, N5 is turned on, the transmission gate composed of P6, N6 and P8, N8 is cut off, A2 = the voltage value of IN3IN2, B2 = VDD; in the low-weight calculation circuit, the transmission gate composed of P3, N3 and P5, N5 is turned on, the transmission gate composed of P6, N6 and P8, N8 is cut off, A2 = the voltage value of IN1IN0, B2 = VDD.

[0045] Reference Figure 6 , the control input C of each group of voltage-controlled delay units is the corresponding bit weight and includes a high-order voltage-controlled delay circuit and a low-order voltage-controlled delay circuit. The two voltage-controlled inputs A and B of the high-order voltage-controlled delay circuit are A2 and B2 respectively, and the two voltage-controlled inputs A and B in the low-order voltage-controlled delay circuit are A1 and B1 respectively. The number of unit adjustments contained in the sum of four times the pulse width adjustment amount of the high-order voltage-controlled delay circuit for the reference signal and the pulse width adjustment amount of the low-order voltage-controlled delay circuit for the reference signal is the calculation result of the corresponding bit. The voltage-controlled inputs of each group of voltage-controlled delay units are the same, which are A2 and B2, A1 and B1 respectively, but the control inputs of each group of voltage-controlled delay units are not the same, which are the data weights of the corresponding bits, namely W3, W2, W1, and W0 respectively. This means that the final outputs of different groups of voltage-controlled delay units are used to characterize the data weights of the corresponding bits and the multiplication calculation results of their corresponding two-bit data inputs. When the control input C is 0, the voltage-controlled delay circuit maintains the pulse width of the reference signal; when the control input C is 1, the voltage-controlled delay circuit adjusts the pulse width of the reference signal and the adjustment strategy is: the pulse width of the reference signal is linearly positively correlated with the voltage-controlled input A and linearly negatively correlated with the voltage-controlled input B. The voltage-controlled delay circuit also has an input port IN and an output port OUT, which are respectively used to receive a reference signal and output a reference signal, and the reference signal is a rectangular wave signal.

[0046] The functions of the voltage-controlled delay circuit are as follows:

[0047] When C = 0 and CB (the inverse signal of C) = 1, the delay amount of the voltage-controlled delay circuit is not affected by the voltage value of the voltage-controlled port. After the rectangular wave is delayed by the voltage-controlled delay circuit, the delay amount of the rising edge is t0, and the delay amount of the falling edge is t0, that is, the pulse width of the rectangular wave remains unchanged.

[0048] When C = 1 and CB = 0, there are the following situations:

[0049] When A = 0, B = VDD, after the rectangular wave is delayed by the voltage-controlled delay unit, the delay amount of the rising edge is t0, and the delay amount of the falling edge is t0; when A = 0, B = V ij When the rectangular wave is delayed by the voltage-controlled delay unit, the delay of the falling edge is t0, and the delay of the rising edge is t0 + x , where: V ij = V 00 When the rising edge delay is t0 +0 , V ij = V 01 When the rising edge delay is t0 +1 , V ij = V 10 When the rising edge delay is t0 +2 , V ij = V 11 When the rising edge delay is t0 +3 ; A = V ij When B = VDD, after the rectangular wave is delayed by the voltage-controlled delay unit, the delay of the rising edge is t0, and the delay of the falling edge is t0 + x , where: V ij = V 00 When the falling edge delay is t0 +0 , V ij = V 01 When the falling edge delay is t0 +1 , V ij =V 10 When the falling edge delay is t0 +2 , V ij = V 11 When the falling edge delay is t0 +3 .

[0050] In this embodiment, the voltage-controlled delay circuit is composed of a current-starved inverter and an inverter connected in series. The output of the current-starved inverter is connected to the input of the inverter. The current-starved inverter has corresponding ports, namely two voltage-controlled input ports (receiving voltage-controlled inputs A and B), two control input ports (receiving control inputs C and CB) and the input port IN of the voltage-controlled delay circuit. The output port of the inverter is the output port OUT of the voltage-controlled delay circuit.

[0051] From the above description, it can be seen that when the control input C is 0, the voltage-controlled delay circuit maintains the pulse width of the reference signal; when the control input C is 1, the voltage-controlled delay circuit adjusts the pulse width of the reference signal, and the larger the voltage-controlled input A, the larger the falling edge delay and the larger the pulse width, and the larger the voltage-controlled input B, the larger the rising edge delay and the smaller the pulse width, and the change in the delay is linearly related to the corresponding voltage-controlled input, that is, the pulse width of the reference signal is linearly positively correlated with the voltage-controlled input A and linearly negatively correlated with the voltage-controlled input B.

[0052] In this embodiment, for a group of voltage-controlled delay units, the number of unit adjustment amounts contained in the sum of four times the pulse width adjustment amount of the high-order voltage-controlled delay circuit for the reference signal and the pulse width adjustment amount of the low-order voltage-controlled delay circuit for the reference signal is the multiplication result of the corresponding bit input and the weight. It is the unit adjustment amount, that is, the minimum adjustment amount of the voltage-controlled delay circuit for the reference signal.

[0053] The four groups of voltage-controlled delay units correspond to four bits respectively. Therefore, the multiplication result of the four-bit data input and the data weight can be obtained through the four groups of voltage-controlled delay units. Then, the multiplication result from the high bit to the low bit is added according to the weight ratio of 8 / 4 / 2 / 1, and the five-bit input IN can be obtained. MSB IN3IN2IN1IN0 and five-bit weight W MSB The multiplication calculation result of W3W2W1W0, in which the product of symbol input and symbol weight has been integrated into the multiplication calculation of quad-bit data input and data weight.

[0054] In summary, the multi-bit data in-memory computing array provided in this embodiment can realize five-bit input IN MSB IN3IN2IN1IN0 and five-bit weight W MSB The multiplication calculation of W3W2W1W0 includes one sign bit and four data bits. And the core is to characterize the calculation result by the pulse width adjustment amount of the reference signal. Since the pulse width adjustment amount can be accumulated, when it is necessary to realize the five-bit input IN MSB IN3IN2IN1IN0 and five-bit weight W MSBWhen performing the multiplication and accumulation calculation of W3W2W1W0, it is only necessary to combine the multi-column multi-bit data in-memory calculation array in the form of rows, and use the reference signal output by each voltage-controlled delay circuit in the previous column as the reference signal received by the corresponding voltage-controlled delay circuit in the next column. The correspondence here means that they all correspond to the same bit and are all high-bit delay circuits or low-bit delay circuits. For example, there is a correspondence between the high-bit delay circuit corresponding to the second bit in the previous column and the high-bit delay circuit corresponding to the second bit in the next column. The single-bit calculation result in the previous column can then be passed to the calculation result of the corresponding bit in the next column. The multi-bit data in-memory calculation array provided in this embodiment supports the formation of a multi-bit data in-memory calculation array structure in a combined form to realize a five-bit input IN MSB IN3IN2IN1IN0 and five-bit weight W MSB The multiplication and accumulation calculation of W3W2W1W0 can provide greater system-level reasoning accuracy and efficiency, solving the problem that the current non-volatile in-memory computing circuits usually only support in-memory multiplication and accumulation calculations of single-bit input and weight, and can only provide limited system-level reasoning accuracy.

[0055] Furthermore, the multi-bit data in-memory calculation array structure provided in this embodiment also includes a TDC quantization module and a digital shift adder module, which are respectively used to implement the quantization of the pulse width adjustment amount and the addition of the quantization results in the above process.

[0056] Specifically, the TDC quantization module is used to quantize the number of unit adjustment amounts contained in the pulse width adjustment amount between the reference signal output by any voltage-controlled delay circuit in the first column and the reference signal output by the corresponding voltage-controlled delay circuit in the last column.

[0057] Exemplarily, the control inputs of the four voltage-controlled delay units corresponding to the same bit in the four columns are 1, 1, 1, and 0 respectively, the operation results of the sign bits of the four columns are positive, negative, positive, and positive respectively, and the voltage-controlled inputs of the four voltage-controlled delay units are as follows: A = VSS, B = V 01 , A = V 10 , B = VDD, A = VSS, B = V 11 , A = VSS, B = V 01 The theoretical quantization result of TDC is (1×1)-(2×1)+(3×1)+(1×0)=2. Among them, after the first voltage-controlled delay circuit, the rising edge of the rectangular wave is delayed by t0+1. , falling edge delay t0; after the second voltage-controlled delay circuit, the rectangular wave rising edge delay t0, falling edge delay t0 + 2 ; After the third voltage-controlled delay circuit, the rising edge of the rectangular wave is delayed by t0 + 3 , falling edge delay t0; after the fourth column voltage-controlled delay circuit, the rectangular wave rising edge delay t0, and the falling edge delay t0.

[0058] It should be noted that the delay of the rising edge greater than t0 will narrow the pulse width of the rectangular wave. For example, after the rectangular wave is delayed by the first row of voltage-controlled delay units, the rising edge delay of the rectangular wave is t0 + 1 , the falling edge delay t0, the rectangular wave pulse width is narrowed by 1 relative to the initial rectangular wave The delay of the falling edge is greater than t0, which will make the pulse width of the rectangular wave wider. For example, after the rectangular wave is delayed by the second voltage-controlled delay unit, the rising edge delay of the rectangular wave is t0, and the falling edge delay is t0+ 2 The pulse width of this rectangular wave is 2 times wider than that of the first rectangular wave. .

[0059] Based on this, the initial pulse width of the rectangular wave can be set to 10 , after the first row of voltage-controlled delay units, the rising edge of the rectangular wave is delayed by t0 + 1 , the falling edge delay is t0, and the rectangular wave pulse width becomes narrower; after the second row of voltage-controlled delay units, the rectangular wave rising edge delay is t0, and the falling edge delay is t0+ 2 , the rectangular wave pulse width becomes wider; after the third voltage-controlled delay unit, the rectangular wave rising edge delay t0 + 3 , the falling edge is delayed by t0, and the rectangular wave pulse width becomes narrower; after the fourth voltage-controlled delay unit, the rectangular wave rising edge is delayed by t0, the falling edge is delayed by t0, and the rectangular wave pulse width remains unchanged; after the four-stage voltage-controlled delay unit delay, the rectangular wave pulse width is 8 , the pulse width of the initial rectangular wave becomes narrower by 10 - 8 = 2.

[0060] Therefore, in this embodiment, it is only necessary to convert the last column of rectangular wave pulse width into a binary number (which contains the unit pulse width adjustment amount). The number of pulse widths), and then the binary number set according to the initial rectangular wave pulse width (the unit pulse width adjustment amount contained in the initial rectangular wave pulse width) By subtracting the binary number from the number of columns (the number of columns that contain 2 bits by 1 bit), we can get the result of the four columns of 2-bit × 1-bit multiplication and accumulation in the above example.

[0061] As described above, the multi-bit multiplication and accumulation process of the multi-bit data in-memory calculation array structure in this embodiment is explained, which illustrates that the multiplication calculation result in the previous column can be accumulated to the multiplication calculation result in the next column. Further, in the multi-bit data in-memory calculation array structure, the number of columns of the multi-bit data in-memory calculation array determines the maximum number of accumulation items that can be performed by the structure, and the number of rows of the multi-bit data in-memory calculation array determines the number of groups of five-bit multiplication and accumulation that can be performed simultaneously by the structure. For example, in this embodiment, there are 4 multi-bit data in-memory calculation arrays in each column, indicating that 4 groups of multi-bit ratio multiplication and accumulation calculations can be performed in parallel, and the multi-bit data in-memory calculation array has a total of 32 columns, indicating that a maximum of 32 multi-bit ratio multiplication calculation results can be accumulated.

[0062] In this embodiment, an SRAM (static random access memory) is also provided, and the multi-bit data in-memory calculation array structure provided in this embodiment is used to implement the multiplication and accumulation calculation of five-bit input and five-bit weight.

[0063] This embodiment also provides an electronic device, including a memory and a processor, wherein the memory is the SRAM provided in this embodiment.

[0064] It should be understood that the specific embodiments described herein are only used to explain the application, rather than to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of this application.

[0065] Obviously, the drawings are only some examples or embodiments of the present application. For ordinary technicians in the field, the present application can also be applied to other similar situations based on these drawings without creative work. In addition, it is understandable that although the work done in this development process may be complicated and lengthy, for ordinary technicians in the field, certain changes in design, manufacturing or production based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient content disclosed in this application.

Claims

1. A multi-bit data in-memory computing array for determining a five-bit input IN MSB IN3IN2IN1IN0 and five-bit weight W MSB The multiplication result of W3W2 W1W0 is characterized by: It includes a first calculation circuit, four second calculation circuits and four groups of voltage-controlled delay units corresponding to different bits respectively; The output of the first calculation circuit is , the outputs of the four second calculation circuits are A2, B2, A1 and B1 respectively; when When, A2 = VSS, B2 = IN3IN2, A1 = VSS, B1 = IN1IN0; when When, A2 = IN3IN2, B2 = VDD, A1 = IN1IN0, B1 = VDD; The control input C of each group of voltage-controlled delay units is the corresponding bit weight and includes a high-order voltage-controlled delay circuit and a low-order voltage-controlled delay circuit. The two voltage-controlled inputs A and B of the high-order voltage-controlled delay circuit are A2 and B2 respectively, and the two voltage-controlled inputs A and B in the low-order voltage-controlled delay circuit are A1 and B1 respectively. The number of unit adjustment amounts contained in the sum of four times the pulse width adjustment amount of the high-order voltage-controlled delay circuit for the reference signal and the pulse width adjustment amount of the low-order voltage-controlled delay circuit for the reference signal is the calculation result of the corresponding bit; When the control input C is 0, the voltage-controlled delay circuit maintains the pulse width of the reference signal; When the control input C is 1, the voltage-controlled delay circuit adjusts the pulse width of the reference signal and the adjustment strategy is: the pulse width of the reference signal is linearly positively correlated with the voltage-controlled input A and linearly negatively correlated with the voltage-controlled input B.

2. The multi-bit data in-memory computing array according to claim 1, characterized in that: Also included are the functions for storing W MSB , W3, W2, W1, and W0.

3. The multi-bit data in-memory computing array according to claim 1, characterized in that: Each group of memory arrays includes a plurality of 6T memory cells, wherein the first storage node and the second storage node of the 6T memory cells are connected to the local bit lines LBL and LBLB respectively; The first calculation circuit includes PMOS tubes P1 and P2, the sources of P1 and P2 are connected to local bit lines LBL and LBLB respectively, and the gates of P1 and P2 are connected to IN MSB and IN MSB The drains of P1 and P2 are connected and constitute an output node of the first calculation circuit; The first computing circuit is used to store W MSB The memory array shares the same pair of local bit lines LBL and LBLB.

4. The multi-bit data in-memory computing array according to claim 1, characterized in that: The second calculation circuit includes a first signal channel and a second signal channel. The outputs of the first signal channel and the second signal channel are connected to form the output of the second calculation circuit. The second signal channel is turned on when When conducting; In the second calculation circuit with output A2, the inputs of the first signal channel and the second signal channel are IN3IN2 and VSS respectively; In the second calculation circuit with output B2, the inputs of the first signal channel and the second signal channel are VDD and IN3IN2 respectively; In the second calculation circuit whose output is A1, the inputs of the first signal channel and the second signal channel are IN1IN0 and VSS respectively; In the second calculation circuit whose output is B1, the inputs of the first signal channel and the second signal channel are VDD and IN1IN0 respectively.

5. The multi-bit data in-memory computing array according to claim 1, characterized in that: The voltage-controlled delay circuit is composed of a current-starved inverter and an inverter connected in series.

6. A multi-bit data in-memory computing array structure for determining a five-bit input IN MSB IN3IN2IN1IN0 and five-bit weight W MSB The multiplication and accumulation result of W3W2 W1W0 is characterized by: comprising a plurality of columns of a multi-bit data in-memory computing array as claimed in any one of claims 1 to 5; The reference signal output by each voltage-controlled delay circuit in the previous column is the reference signal received by the corresponding voltage-controlled delay circuit in the next column.

7. The multi-bit data in-memory computing array structure according to claim 6, characterized in that: It also includes a TDC quantization module and a digital shift adder module; The TDC quantization module is used to quantize the number of unit adjustment amounts contained in the pulse width adjustment amount between the reference signal output by any voltage-controlled delay circuit in the first column and the reference signal output by the corresponding voltage-controlled delay circuit in the last column; For any group of voltage-controlled delay units, the digital shift adder module is used to amplify the quantization result of the high-bit voltage-controlled delay circuit by four times and then add it to the quantization result of the low-bit voltage-controlled delay circuit to obtain the multiplication calculation result of the corresponding bit input and weight; And the digital shift adder module is used to perform weighted summation on the multiplication calculation results of different bit inputs and weights according to the bit relationship to obtain the multiplication and accumulation result of five-bit input and five-bit weight.

8. The multi-bit data in-memory computing array structure according to claim 7, characterized in that: There are 4 multi-bit data in-memory calculation arrays in each column, and there are 32 columns of multi-bit data in-memory calculation arrays in total.

9. A SRAM, characterized in that: The multi-bit data in-memory computing array structure described in any one of claims 6 to 8 is used to implement multiplication and accumulation calculations of five-bit inputs and five-bit weights.

10. An electronic device comprising a memory and a processor, characterized in that: The memory is the SRAM described in claim 9.

Citation Information

Patent Citations

  • Multi-bit in-memory computing device with symbols

    CN114895869A

  • Charge domain storage calculation circuit and storage calculation circuit with positive and negative number operation function

    CN115910152A