In-Memory Computing Device
By dividing the memory array into p×q tiles with adjustable bit line switches and using a staircase adder, the memory-in-computation device reduces data storage and power consumption, improving computational efficiency for DNNs in AIoT applications.
Patent Information
- Application Number
- CN202010939405.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-28
- Filing Date
- 2020-09-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-09-09
AI Technical Summary
In multiplication and addition operations of multiple input signals and multiple weights, the data storage requirements and power consumption of the computing device in the memory are high, which has become an urgent problem.
The memory array is divided into p×q memory partitions, and the number of bits of weight is adjusted by adjusting the number of conduction of the bit line selection switch. At the same time, the step adder is used to perform addition operations to reduce data storage requirements.
It effectively reduces the data storage requirements and hardware costs of computing devices in memory, and improves computing efficiency.
Smart Images

Figure CN114115797B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an in-memory computing device, and more particularly to an in-memory computing device that can reduce the demand for data storage. Background Art
[0002] In recent years, artificial intelligence (AI) accelerators for deep neural networks (DNNs) used in edge computing have become increasingly important for the integration and implementation of artificial intelligence-based Internet of Things (AIoT) applications. In addition to the traditional von Neumann computing architecture, an in-memory computing (CIM) architecture that can further improve computing efficiency has been proposed.
[0003] However, in the multiply-accumulate operations of multiple input signals and multiple weights, a large range and a large amount of data will inevitably be generated. Therefore, how to reduce the data storage requirements and power consumption of in-memory computing devices has become an important issue for those skilled in the art. Summary of the Invention
[0004] The present invention provides an in-memory computing device that can reduce the demand for data storage.
[0005] The in-memory computing device of the present invention includes a memory array, p×q analog-to-digital converters, and a stepped adder. The memory array is divided into p×q storage partitions (tiles), where p and q are both positive integers greater than 1. Each storage partition has multiple partition bit lines, which are respectively coupled to corresponding global bit lines through multiple bit line selection switches. The bit line selection switches are respectively turned on or off according to multiple control signals. The memory array receives multiple input signals. The analog-to-digital converters are respectively coupled to multiple global bit lines of the storage partitions. The analog-to-digital converters respectively convert the electrical signals on the global bit lines to generate p×q digital sub-output values. The stepped adder is coupled to the analog-to-digital converters and performs an addition operation on the sub-output values to generate an operation result.
[0006] Based on the above, the present invention divides the memory array into multiple storage partitions, and each storage partition can adjust the number of bits of the weight by adjusting the number of turned-on bit line selection switches. Multiple storage partitions respectively generate multiple sub-output values according to the received input signals. Then, the stepped adder performs an addition operation on the sub-output values to generate an operation result. Through the above architecture, in the in-memory computing device, the data storage requirements for performing multiply-accumulate operations can be reduced, the hardware cost and power consumption can be effectively reduced, and the computing speed can be improved. Description of the Drawings
[0007] Figure 1 Schematic diagram showing an in - memory computing device according to an embodiment of the present invention.
[0008] Figure 2A Schematic diagram showing a partitioning method of a memory array in an in - memory computing device according to an embodiment of the present invention.
[0009] Figure 2B And Figure 2C Schematic diagram showing a method for adjusting the number of weight bits according to an embodiment of the present invention.
[0010] Figure 3A And Figure 3B Schematic diagrams respectively showing different implementation manners of a storage partition according to an embodiment of the present invention.
[0011] Figure 4 Schematic diagram showing an implementation manner of a stepped adder according to an embodiment of the present invention.
[0012] Figure 5 Schematic diagram showing an implementation manner of a first sub - stepped adder according to an embodiment of the present invention.
[0013] Figure 6 Schematic diagram showing an implementation manner of a second sub - stepped adder according to an embodiment of the present invention.
[0014] Figure 7 Schematic diagram showing a partial circuit of an in - memory computing device according to another embodiment of the present invention.
[0015]
Symbol Description
[0016] 100, 700: In - memory computing device
[0017] 110, 200: Memory array
[0018] 120, 400: Stepped adder
[0019] 411~41p, 500: First sub - stepped adder
[0020] 420, 600: Second sub - stepped adder
[0021] 720: Normalization circuit
[0022] 721: Multiplier
[0023] 722: Full adder
[0024] 730: Quantizer
[0025] 731: Divider
[0026] AD11~ADqp: Analog - to - digital converter
[0027] BF: Offset parameter
[0028] BL1 - BLL: Partition bit lines
[0029] BLT1 - BLT6: Bit line selection switches
[0030] CDR1 - CDRp: First direction operation results
[0031] CR: Operation result
[0032] CSL: Common source line
[0033] CT1 - CT6: Control signals
[0034] DEN: Reference value
[0035] FAD11 - FADN1, FAD11a - FADM1a: Full adders
[0036] GBL, GBL_11 - GBL_qp: Global bit lines
[0037] IN11 - INp4: Input signals
[0038] LA1 - LAN, LB1 - LBM: Layers
[0039] MB11 - MBqp: Storage partitions
[0040] MC1 - MCK + 1: Memory cells
[0041] NCR: Adjusted operation result
[0042] OCR: Output operation result
[0043] SF: Scaling factor
[0044] SF11 - SFN1, SF11a - SFM1a: Shifters
[0045] ST1 - STK + 1: Selection switches
[0046] SV11 - SVqp: Sub - output values Detailed implementation manners
[0047] Please refer to Figure 1 , Figure 1Schematic diagram of an in-memory computing device according to an embodiment of the present invention. The in-memory computing device 100 can be applied to the calculation of deep neural networks (DNNs). The in-memory computing device 100 includes a memory array 110, p×q analog-to-digital converters AD11 to ADqp, and a stepped adder 120. In this embodiment, the memory array 110 can be divided into p×q memory partitions MB11 to MBqp, where both p and q are positive integers greater than 1. Each memory partition MB11 to MBqp can receive a plurality of input signals. A plurality of memory cells in each memory partition MB11 to MBqp provide a plurality of weight values, and perform a multiply-accumulate operation based on the input signals and the weight values.
[0048] The analog-to-digital converters AD11 to ADqp are respectively coupled to the memory partitions MB11 to MBqp. In this embodiment, each memory partition MB11 to MBqp has a global bit line. The analog-to-digital converters AD11 to ADqp are respectively coupled to the global bit lines of the memory partitions MB11 to MBqp. The analog-to-digital converters AD11 to ADqp respectively perform analog-to-digital conversion operations on the electrical signals on the global bit lines of the memory partitions MB11 to MBqp, and thereby generate p×q sub-output values SV11 to SVqp. Among them, the above electrical signals can be voltage signals or current signals.
[0049] In this embodiment, each memory partition MB11 to MBqp has a plurality of partition bit lines. All the partition bit lines corresponding to the same memory partition are coupled to the corresponding global bit line.
[0050] The stepped adder 120 is coupled to the analog-to-digital converters AD11 to ADqp. The stepped adder 120 receives the sub-output values SV11 to SVqp, and performs an addition operation on the sub-output values AD11 to ADqp, and thereby generates an operation result CR.
[0051] Please refer to the following Figure 2A , Figure 2AA schematic diagram showing the partitioning method of the memory array in the in-memory computing device according to an embodiment of the present invention. In this embodiment, the memory array 200 includes a plurality of NOR flash memory cells. The memory array 200 can be divided into p×q memory partitions MB11~MBqp, including p memory partition columns and q memory partition rows. Both p and q are positive integers greater than 1. The memory partitions MB11~MBqp respectively have a plurality of global bit lines GBL_11~GBL_qp. The memory partitions MB11~MBqp arranged in the same column can receive the same input signal. For example, the memory partitions MB11, MBq1 arranged in the same column receive the input signals IN11~IN14, and the memory partitions MB1p, MBqp arranged in the same column receive the input signals INp1~INp4.
[0052] Corresponding to the memory partitions MB11~MBqp, the in-memory computing device according to an embodiment of the present invention is provided with p×q analog-to-digital converters AD11~ADqp. The analog-to-digital converters AD11~ADqp are respectively coupled to the global bit lines GBL_11~GBL_qp, and perform analog-to-digital conversion operations on the electrical signals on the global bit lines GBL_11~GBL_qp to respectively generate p×q sub-output values.
[0053] Here, the memory cells in the memory array 200 can be pre-programmed with values of 0 or 1, and by making the word line voltage received by the memory cells selected or unselected, the memory cells can provide the required weight values.
[0054] Incidentally, in Figure 2A the embodiment, if the memory array 200 has j rows of memory cells (which can provide j-bit weight values), the maximum weight value that the memory array 200 can provide can be up to 2 j -1, resulting in a large data storage requirement. Through the embodiment of the present invention, the memory array 200 is divided into p×q memory partitions MB11~MBqp, and by respectively arranging the global bit lines GBL_11~GBL_qp in the memory partitions MB11~MBqp, on the premise of providing j-bit weight values, the maximum weight value that each memory partition MB11~MBqp can provide can be up to q×(2 j / q -1), effectively reducing the data storage requirement.
[0055] On the other hand, when the memory array 200 has i columns of memory cells (which can provide i-bit input signals), by dividing the memory array 200 into p×q memory partitions MB11~MBqp, the number of input signals corresponding to a single memory partition can be reduced to i / p, and therefore, the maximum value of the input signals corresponding to the p memory partitions in this embodiment can be p×(2 i / p-1).
[0056] Based on the above, through the differentiation method of the p×q storage partitions MB11 to MBqp in the embodiments of the present invention, the data storage requirements of the memory array 200 having i columns of memory cells and j rows of memory cells can be reduced from (2 i -1)×(2 j -1) to p×q×(2 i / p -1)×(2 j / q -1). Taking i = j = 8, p = q = 4 as an example, the data storage requirements can be reduced from 65025 to 144.
[0057] In addition, please refer to Figure 2B and Figure 2C , Figure 2B and Figure 2C which are schematic diagrams showing the weight bit adjustment method in the embodiments of the present invention. In Figure 2B , taking the storage partition MB11 of Figure 2A as an example, the storage partition MB11 has a plurality of bit line selection switches BLT1 to BLT6. The partition bit lines (local bit lines) BL1 to BL6 are respectively coupled to the global bit line GBL_11 through the bit line selection switches BLT1 to BLT6. The bit line selection switches BLT1 to BLT6 are respectively controlled by the control signals CT1 to CT6 to be turned on or off respectively. Among them, the number of the turned-on bit line selection switches BLT1 to BLT6 can represent the weight bits of the storage partition MB11.
[0058] Taking Figure 2C as an example, among them, the bit line selection switches BLT1 to BLT3 are respectively turned on according to the control signals CT1 to CT3, and the bit line selection switches BLT4 to BLT6 are respectively turned off according to the control signals CT4 to CT6. Under this condition, the three partition bit lines BL1 to BL3 effectively connected to the global bit line GBL_11 can present weights of two bits that can be encoded as 000, 001, 011, and 111. Of course, when the weight bits need to be adjusted, it can be achieved by adjusting the number of the turned-on bit line selection switches BLT1 to BLT6.
[0059] Regarding the implementation manners of each storage partition, please refer to Figure 3A and Figure 3B which are respectively schematic diagrams showing different implementation manners of the storage partitions in the embodiments of the present invention. In Figure 3A , the storage partition MB11 includes a plurality of memory cells MC1 to MCK+1, where the memory cells MC1 to MCK+1 are two-transistor (2T) NOR flash memory cells.
[0060] The memory cells MC1 to MCK+1 are arranged in a planar manner. The memory partition MB11 further includes a plurality of selection switches ST1 to STK+1 corresponding to the memory cells MC1 to MCK+1 respectively. The selection switches ST1 to STK+1 are composed of transistors. The memory cells MC1 to MCK+1 and the corresponding selection switches ST1 to STK+1 are sequentially connected in series between the corresponding partition bit lines and the common source line CSL. Taking the memory cells MC1 and MC2 as examples, the memory cell MC1 and the selection switch ST1 are sequentially connected in series between the partition bit line BL1 and the common source line CSL; the memory cell MC2 and the selection switch ST2 are sequentially connected in series between the partition bit line BL1 and the common source line CSL. In addition, all the partition bit lines BL1 to BLL in the memory partition MB11 are coupled to the global bit line GBL.
[0061] In the memory partition MB11, the control terminals of the selection switches ST1 and ST2 receive input signals IN11 and IN12 respectively. The gates of the memory cells MC1 and MC2 receive signals MG1 and MG2 respectively. The memory cells MC1 and MC2 form an anti-OR type flash memory element with a 2T architecture. When performing in-memory computing operations, according to the input signals IN11 and IN12, the selection switches ST1 and ST2 respectively provide currents, and then according to the transduction values provided by the memory cells MC1 and MC2 as weight values, and a multiply-accumulate result is generated. The voltage generated according to the multiply-accumulate operation can be transmitted on the partition bit line BL1 to the global bit line GBL.
[0062] In the present embodiment, the global bit line GBL is coupled to the analog-to-digital converter AD11. The analog-to-digital converter AD11 can convert the voltage on the global bit line GBL and obtain a sub-output value in digital format.
[0063] It is worth mentioning that the hardware architecture of the analog-to-digital converter AD11 can be implemented by using an analog-to-digital conversion circuit well-known to those skilled in the art, without specific limitations.
[0064] In addition, in Figure 3B , the memory partition MB11 can be implemented by using a three-dimensional architecture. And each partition bit line BL1 to BLL can be coupled to two or more memory cells. Figure 3B In Figure 3A the same as the operation mode of the memory partition MB11 in the embodiment, it will not be elaborated here.
[0065] Please refer to the following Figure 4 , Figure 4A schematic diagram showing an implementation of a stepped adder according to an embodiment of the present invention. The stepped adder 400 includes a plurality of first sub-stepped adders 411 to 41p and a second sub-stepped adder 420. Corresponding to a memory array divided into p×q storage partitions, the number of the first sub-stepped adders 411 to 41p may be p. According to the example of FIG. 2, the first sub-stepped adders 411 to 41p may respectively correspond to p columns of storage partitions, and each of the first sub-stepped adders 411 to 41p is coupled to q analog-to-digital converters of the corresponding column of storage partitions. In FIG. 2, taking the column of storage partitions where the storage partitions MB11 to MBq1 are located as an example, the first sub-stepped adder 411 may be coupled to the analog-to-digital converters AD11 to ADq1 corresponding to the storage partitions MB11 to MBq1 respectively.
[0066] The first sub-stepped adder 411 performs an addition operation on the sub-output values generated by the analog-to-digital converters AD11 to ADq1 to generate a first-direction operation result CDR1. Similarly, the first sub-stepped adders 412 to 41p may respectively generate a plurality of first-direction operation results CDR2 to CDRp through the performed addition operations.
[0067] The second sub-stepped adder 420 is coupled to the first sub-stepped adders 411 to 41p. The second sub-stepped adder 420 is used to perform an addition operation on the first-direction operation results CDR2 to CDRp respectively generated by the first sub-stepped adders 411 to 41p, and thereby generate an operation result CR.
[0068] Regarding the implementation details of each of the first sub-stepped adders 411 to 41p and the second sub-stepped adder 420, reference may be made to the following Figure 5 and Figure 6 implementation manners.
[0069] Figure 5A schematic diagram showing an implementation of a first sub - stepped adder according to an embodiment of the present invention. The first sub - stepped adder 500 is coupled to q analog - to - digital converters AD11~AD1q corresponding to the same storage partition column. The first sub - stepped adder 500 has N layers LA1~LAN, where N = log2 q. In this embodiment, in each of the layers LA1~LAN, there is one or more full - adders and shifters. Among them, the first layer LA1 includes full - adders FAD11~FAD1A and shifters SF11~SF1A. The full - adders FAD11~FAD1A respectively correspond to the shifters SF11~SF1A, and are staggeredly and sequentially coupled to the analog - to - digital converters AD11~AD1q. Among them, the number of full - adders FAD11~FAD1A is q / 2, and the number of shifters SF11~SF1A is also q / 2. The full - adders FAD11~FAD1A respectively receive the sub - output values generated by the odd - numbered analog - to - digital converters AD11, AD13, …, and the shifters SF11~SF1A respectively receive the sub - output values generated by the even - numbered analog - to - digital converters AD12, …, AD1q, and perform bit - shifting operations.
[0070] Among them, in this embodiment, the shifters SF11~SF1A are used to shift the received sub - output values in the high - order direction. In this embodiment, the bit - shifting amount of the shifters SF11~SF1A in the first layer LA1 is equal to j / q, where j is the total number of rows of storage units in the memory array. The full - adders FAD11~FAD1A respectively receive the outputs of the shifters SF11~SF1A and perform full - add operations.
[0071] Incidentally, in this embodiment, the bit - shifting amount of the shifters in the second layer is equal to 2×j / q, and so on. In addition, in the r - th layer of the first sub - stepped adder 500, there are q / 2 r full - adders and q / 2 r shifters. The full - adders and shifters in the same layer are staggeredly arranged in sequence and are respectively coupled to the output terminals of the full - adders in the previous layer, where 1 < r ≤ N.
[0072] The N - th layer LAN includes a single full - adder FADN1 and a single shifter SFN1. The bit - shifting amount of the shifter SFN1 is equal to 2×(log2 q - 1)×j / q. The full - adder FADN1 then generates a first - direction operation result CDR1.
[0073] The hardware architectures of the full - adders FAD11~FADN1 and shifters SF11~SFN1 in this embodiment can be implemented using full - add circuits and digital - shift circuits well - known to those skilled in the art, without specific limitations.
[0074] Figure 6 A schematic diagram showing an implementation of a second sub - stepped adder according to an embodiment of the present invention. The second sub - stepped adder 600 includes a plurality of layers LB1 - LBM, where each layer LB1 - LBM includes at least one full adder and at least one shifter. Among them, the second sub - stepped adder 600 has M layers LB1 - LBM, and M = log2 p, where p is the number of the first - direction operation results CDR1 - CDRp. In this embodiment, the first layer LB1 has full adders FAD11a - FAD1Ba and shifters SF11a - SF1Ba. The full adders FAD11a - FAD1Ba respectively correspond to the shifters SF11a - SF1Ba and are arranged alternately in pairs. A plurality of first input terminals of the full adders FAD11a - FAD1Ba respectively receive the odd - numbered first - direction operation results CDR1, CDR3,..., CDRp - 1, and a plurality of second input terminals of the full adders FAD11a - FAD1Ba are respectively coupled to the output terminals of the shifters SF11a - SF1Ba. The input terminals of the shifters SF11a - SF1Ba respectively receive the even - numbered first - direction operation results CDR2, CDR4,..., CDRp. The number of the full adders FAD11a - FAD1Ba and the number of the shifters SF11a - SF1Ba are both equal to p / 2.
[0075] In addition, in the s - th layer of the second sub - stepped adder 600, there are p / 2 s full adders and p / 2 s shifters. The full adders and shifters in the same layer are arranged alternately in sequence and are respectively coupled to the output terminals of the full adders in the previous layer, where 1 < s ≤ M. In the last layer LBM, there is a single full adder FADM1a and a single shifter SFM1a. The single full adder FADM1a is used to generate the operation result CR.
[0076] In this embodiment, the shifters SF11a - SF1Ba of the second sub - stepped adder 600 are used to shift the received sub - output values in the high - order direction. The shift amounts of the shifters SF11a - SF1Ba in the first layer LB1 are the same and are both equal to i / p, where i is the number of bits of the input signal. In addition, the shift amounts of the shifters in the second layer of the second sub - stepped adder 600 can all be 2×i / p, and so on. The shift amount of the shifter SFM1a in the last layer LBM can be 2×(log2 p - 1)×i / p.
[0077] The hardware architectures of the full adders FAD11a to FADM1a and the shifters SF11a to SFM1a in this embodiment can be implemented using full add circuits and digital shift circuits well-known to those skilled in the art, without specific limitations. Additionally, the hardware architectures of the full adders FAD11a to FADM1a in this embodiment can be the same as or different from Figure 5 the hardware architectures of the full adders FAD11 to FADN1 in the Figure 5 embodiment. The hardware architectures of the shifters SF11a to SFM1a in this embodiment can be the same as or different from
[0078] Please refer to the following Figure 7 , Figure 7 which is a schematic diagram showing a partial circuit of the in-memory computing device according to another embodiment of the present invention. The in-memory computing device 700 further includes a normalization circuit 720 and a quantizer 730. The normalization circuit 720 is coupled to the output of the stepped adder 710 to receive the operation result CR generated by the stepped adder 710. The normalization circuit 720 includes a multiplier 721 and a full adder 722. The multiplier 721 receives the operation result CR and the scaling factor SF, and multiplies the operation result CR and the scaling factor SF. The full adder 722 then receives the output of the multiplier 721 and further receives the offset parameter BF. The full adder 722 is used to add the output of the multiplier 721 and the offset parameter BF, and thereby generate an adjusted operation result NCR.
[0079] The above-mentioned scaling factor SF and offset parameter BF can be set by the designer himself / herself to normalize the operation result CR to a reasonable numerical range for facilitating subsequent operations.
[0080] The quantizer 730 is coupled to the normalization circuit 720, receives the adjusted operation result NCR, and divides the adjusted operation result NCR by a reference value DEN to generate an output operation result OCR. In this embodiment, the quantizer 730 can be a divider 731. Among them, the reference value DEN can be a non-zero preset value preset by the designer, without specific limitations.
[0081] The hardware architectures of the above-mentioned full adder 722, multiplier 721, and divider 731 can be implemented using full adder circuits, multiplier circuits, and divider circuits well-known in the art respectively, without specific limitations.
[0082] Incidentally, the in-memory computing device 700 of this embodiment can be applied to a Convolutional Neural Network (CNN).
[0083] In summary, the present invention divides the memory array into p×q storage partitions, and then cooperates with a stepped adder to complete the required multiplication and addition operations. Under the architecture of the present invention, the number of weighted bits can be adjusted according to the number of conductive bit-line selection switches. Moreover, the magnitude of the values generated during the operation process can be reduced, the demand for data storage can be effectively reduced, the burden on the hardware can be reduced, and the computing efficiency can be increased.
[0084] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An in-memory computing device, wherein, Comprising: A memory array, which is divided into p×q memory partitions, where p and q are both positive integers greater than 1. The memory array receives a plurality of input signals, and these memory partitions are respectively coupled to a plurality of global bit lines. Each of these memory partitions has a plurality of partition bit lines, and these partition bit lines are respectively coupled to corresponding global bit lines through a plurality of bit line selection switches. These bit line selection switches are respectively turned on or off according to a plurality of control signals, and the number of turned-on bit line selection switches in each of these memory partitions represents the weighted number of bits of each of these memory partitions; p×q analog-to-digital converters, which are respectively coupled to these global bit lines, and these analog-to-digital converters respectively convert the electrical signals on these global bit lines to generate p×q digital sub-output values; And A stepped adder, which is coupled to these analog-to-digital converters and performs an addition operation on these sub-output values to generate an operation result.
2. The in-memory computing device according to claim 1, wherein, The stepped adder includes: p first sub-stepped adders, each of these first sub-stepped adders is respectively coupled to q of these analog-to-digital converters, and these first sub-stepped adders respectively generate p first-direction operation results; and A second sub-stepped adder, which is coupled to these first sub-stepped adders and generates the operation result according to these first-direction operation results.
3. The in-memory computing device according to claim 2, wherein, Each of the first sub-stepped adders has N layers, and each layer includes at least one full adder and at least one shifter, where N = .
4. The in-memory computing device according to claim 3, wherein, The first layer of each of these first sub-stepped adders has q / 2 of the at least one full adder and q / 2 of the at least one shifter, and the at least one full adder and the at least one shifter are arranged alternately in sequence to respectively receive the corresponding these sub-output values.
5. The in-memory computing device according to claim 4, wherein, In the r-th layer of each of the first sub-stepped adders, there are q / 2 r of the at least one full adder and q / 2 r of the at least one shifter. The at least one full adder and the at least one shifter in the same layer are arranged alternately in sequence and are respectively coupled to the output ends of the at least one full adder in the previous layer, where 1 < r ≤ N.
6. The in-memory computing device according to claim 2, wherein, The second sub-step adder has M layers, each layer including at least one full adder and at least one shifter, where M = , and the at least one full adder and the at least one shifter are arranged alternately to respectively receive these first-direction operation results.
7. The in-memory computing device according to claim 6, wherein, The first layer of the second sub-stepped adder has p / 2 of the at least one full adder and p / 2 of the at least one shifter, and the at least one full adder and the at least one shifter are arranged alternately in sequence to respectively receive these first-direction operation results.
8. The in-memory computing device according to claim 6, wherein, In the s-th layer of the second sub-step adder, there are p / 2 s of the at least one full adder and p / 2 s of the at least one shifter. The at least one full adder and the at least one shifter in the same layer are arranged alternately in sequence and are respectively coupled to the output ends of the at least one full adder in the previous layer, where 1 < s ≤ M.
9. The in-memory computing device according to claim 1, wherein, Further comprising: A normalization circuit, which is coupled to the stepped adder, normalizes the operation result according to a scaling factor and an offset parameter, and generates an adjusted operation result.
10. The in-memory computing device according to claim 9, wherein, Further comprising: A quantizer, which is coupled to the normalization circuit, performs a quantization operation on the adjusted operation result according to a reference value, and generates an output operation result.
11. The in-memory computing device according to claim 10, wherein, The quantizer is a divider, which divides the adjusted operation result by the reference value to generate the output operation result.
12. The in-memory computing device according to claim 9, wherein, The normalization circuit includes: A multiplier, which multiplies the operation result by the scaling factor; and A full adder, which adds the output of the multiplier and the offset parameter to generate the adjusted operation result.
13. The in-memory computing device according to claim 1, wherein, The memory array includes a plurality of two-transistor NOR flash memory cells.
14. The in-memory computing device according to claim 13, wherein, The control terminal of a selection transistor in each of these two-transistor NOR flash memory cells receives one of these input signals.
Citation Information
Patent Citations
In-memory computing chip based on NAND Flash, storage device and terminal
CN211016545U
Hierarchical dynamic memory array architecture using read amplifiers separate from bit line sense amplifiers
US6198682B1