Computing memory cell and processing array device using the same

By designing a dual-port SRAM cell, using a cross-coupled inverter and isolation circuit, the bandwidth bottlenecks and stability problems existing in the read and write operations of the existing 6T SRAM cell are solved, achieving more efficient computing and lower power consumption.

CN120029967APending Publication Date: 2025-05-23GSI TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510062300.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2017-09-19
Filing Date
2017-11-06
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing 6T SRAM cells have bandwidth bottlenecks and stability problems in read and write operations, especially when multiple units are turned on, which can easily lead to data jumps and insufficient write driver strength.

Method used

A dual-port SRAM cell is designed, using a cross-coupled inverter and an isolation circuit. Through the independent operation of the read and write ports, the read and write performance of the unit is improved, and the influence of bit line level is reduced through the isolation circuit.

Benefits of technology

Faster sensing and higher number of computing objects are achieved, power consumption and bit line discharge time are reduced, and data jumps and insufficient write driver strength is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029967A_ABST
    Figure CN120029967A_ABST
Patent Text Reader

Abstract

The invention relates to a computing memory cell and a processing array device using the same. A memory cell that may be used for computation and a processing array using the memory cell are capable of performing logical operations including Boolean AND, Boolean OR, Boolean NAND, or Boolean NOR. The memory cell may have a read port with an isolation circuit that isolates data stored in a storage cell of the memory cell from a read bit line.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent application is a divisional application of a Chinese invention patent application with application number 2017800857350 and invention name “Computational storage unit and processing array device using storage unit”, which entered the national phase of the People’s Republic of China based on the PCT international application with application number PCT / US2017 / 060227 and application date of November 6, 2017. Claiming priority / related applications

[0002] This application claims priority under 35 USC §§ 119, 364, and 365 to U.S. Nonprovisional Patent Application Serial No. 15 / 709,379, filed on September 19, 2017, and entitled “Computational Memory Cell and Processing Array Device Using Memory Cells,” U.S. Nonprovisional Patent Application Serial No. 15 / 709,382, filed on September 19, 2017, and entitled “Computational Memory Cell and Processing Array Device Using Memory Cells,” and U.S. Nonprovisional Patent Application Serial No. 15 / 709,385, filed on September 19, 2017, and entitled “Computational Memory Cell and Processing Array Device Using Memory Cells,” all of which are hereby granted under 35 USC §§ 119, 364, and 365. 119(e) and 120, and in turn claims the benefit of and priority to U.S. Provisional Patent Application Serial No. 62 / 430,762, filed on December 6, 2016, and entitled “Computational Dual Port Sram Cell and Processing Array Device Using the Dual Port Sram Cells,” the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure generally relates to memory cells that can be used for computing. Background Art

[0004] Arrays of memory cells such as dynamic random access memory (DRAM) cells, nonvolatile memory cells, nonvolatile storage devices, or static random access memory (SRAM) cells, or content addressable memory (CAM) cells are well-known devices used in various computer or processor-based devices to store digital bits of data. Various computer and processor-based devices may include computer systems, smart phone devices, consumer electronics, televisions, network switches and routers, etc. Arrays of memory cells are typically packaged in an integrated circuit, or may be packaged within an integrated circuit that also has a processing device within the integrated circuit. Different types of typical memory cells have different performance and characteristics that distinguish each type of memory cell. For example, DRAM cells take longer to access; lose their data content unless they are periodically refreshed; but are relatively inexpensive to manufacture due to the simple structure of each DRAM cell. On the other hand, SRAM cells have faster access times; do not lose their data content unless power is removed from the SRAM cell; and are relatively more expensive due to the greater complexity of each SRAM cell compared to DRAM. CAM cells have the unique capability of being able to easily address the contents within the cell, but are more expensive to manufacture because each CAM cell requires more circuitry to implement the content addressing capability.

[0005] Various computing devices that can be used to perform calculations on digital binary data are also known. The computing devices may include microprocessors, CPUs, microcontrollers, etc. These computing devices are usually manufactured on integrated circuits, but may also be manufactured on integrated circuits that also have a certain amount of memory, which is integrated on the integrated circuit. In these known integrated circuits with computing devices and memory, the computing device performs calculations on digital binary data bits, and the memory is used to store various digital binary data including, for example, instructions executed by the computing device and data on which the computing device operates.

[0006] Recently, devices have been introduced that use memory arrays or storage units to perform computing operations. In some of these devices, a processor array can be formed from the storage units to perform the calculations. These devices can be referred to as in-memory computing devices.

[0007] Big data operations are data processing operations that must process large amounts of data. Machine learning uses artificial intelligence algorithms to analyze data and typically requires large amounts of data to perform. Big data operations and machine learning are also typically very computationally intensive applications that often encounter input / output issues due to bandwidth bottlenecks between the computing device and the memory where the data is stored. The in-memory computing devices described above can be used, for example, for these big data operations and machine learning applications because the in-memory computing devices perform the computations in memory, thereby eliminating the bandwidth bottleneck.

[0008] In-memory computing devices typically use well-known standard SRAM or DRAM or CAM memory cells that can perform calculations. For example, Figure 1 A standard 6T SRAM cell that can be used for computing is shown. The standard 6T SRAM cell may have a bit line (BL) and a complementary bit line (BLb) and a word line (WL) connected to the cell. The cell may include two access transistors (M13, M14), and each access transistor has a source coupled to the bit line (BL and BLb), respectively. Each access transistor also has a gate, and the gates of the two access transistors are connected to the word line (WL), as shown in FIG. Figure 1 As shown. The drain of each access transistor can be connected to a pair of inverters (I11, I12) that are cross-coupled to each other. The side of the cross-coupled inverters closest to the bit line BL can be labeled as D, and the other side of the cross-coupled inverters closest to the complementary bit line (BLb) can be labeled as Db. As is known in the art, the cross-coupled inverters act as storage elements of the SRAM cell, and reading data from / writing data to the SRAM cell is known in the art and will now be described in more detail.

[0009] When two cells connected to the same bit line are turned on, the bit line (BL) can perform an AND function of two bits of data stored in the cells. During the read cycle, both BL and BLb have static pull-up transistors, and if the data in both cells are logic high "1", BL remains at 1. If any of the data in the cells is logic low "0" or both are logic low "0", BL is pulled to a lower level and will be logic 0. By sensing the BL level, 2 cells are used to perform an AND function. Similarly, if 3 cells are turned on, the BL value is the result of the AND function of the data stored in the 3 cells. During a write operation, multiple word lines (WL) can be turned on, so multiple cells can be written at the same time. In addition, writing (or selective writing) can be performed selectively, which means that during the write cycle, if both BL and BLb remain high, writing will not be performed.

[0010] Figure 1 The cell shown has its disadvantages. During the read cycle, when multiple cells are turned on, if all cells except one cell store a low logic value "0", the BL voltage level is the ratio of the pull-down transistor of the "0" cell to the BL pull-up transistor. If the BL voltage level is too low, this will cause the cell storing a logic "1" to jump to a logic "0". As a result, it seems desirable to have a strong BL pull-up transistor to allow more cells to be turned on. However, if only 1 cell contains "0" data during the read, the strong BL pull-up transistor will make the "0" signal small, making it difficult to sense the data.

[0011] During the write cycle, Figure 1 The cells in the memory cell also have disadvantages. If multiple cells to which data is to be written are valid, the BL driver for writing needs to be strong enough to make the latch device ( Figure 1 In addition, the more WLs that are turned on during a write cycle, the stronger the write driver needs to be, which is not desirable.

[0012] During the selective write cycle, Figure 1 The cells in the BL also have disadvantages. Specifically, the BL pull-up transistor needs to be strong to combat the "0" stored in multiple valid cells. Similar to the read cycle described above, when all cells except one are valid and contain a "0", the lone cell containing a "1" is susceptible to instability caused by the lower BL level.

[0013] Therefore, it is desirable to be able to compute but not have Figure 1 The disadvantages of a typical 6T SRAM cell are shown. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 A typical six-transistor static random access memory cell is shown; Figure 2 A first embodiment of a dual-port SRAM cell that can be used for computing is shown; Figure 3 Show that it can be incorporated Figure 2 , Figure 4 , Figure 6 ,or Figure 7 A processing array device of a dual-port SRAM cell; Figure 4 A second embodiment of a dual-port SRAM cell that can be used for computing is shown; Figure 5 It is for Figure 4 The write port truth table of the dual-port SRAM cell; Figure 6 A third embodiment of a dual-port SRAM cell that can be used for computing is shown; Figure 7 A fourth embodiment of a dual-port SRAM cell that can be used for computing is shown; Figure 8 and Fig. 9 Shown can be used in Figure 2 , Figure 4 , Figure 6 ,or Figure 7 Two examples of latched inverters in a dual-port SRAM cell are shown; Fig.10 shows an embodiment of a dual-port SRAM cell that can be used for computing; and Fig.11 Another embodiment of a dual-port SRAM cell that can be used for computing is shown. DETAILED DESCRIPTION

[0015] The present disclosure is particularly applicable to static random access memory (SRAM) cells or arrays of cells, or processing arrays having different layouts as described below, and the present disclosure will be described in the context. However, it will be understood that SRAM devices and processing arrays using SRAM cells have greater practicality because each SRAM cell can be configured / laid out differently than the embodiments described below, and changes to the configuration / layout of the dual-port SRAM cells that can be used for calculations fall within the scope of the present disclosure. For illustrative purposes, dual-port SRAM cells are disclosed below and in the accompanying drawings. However, it is to be understood that SRAM computing cells and processing arrays can also be implemented using SRAM cells having three or more ports, and the present disclosure is not limited to the dual-port SRAM cells disclosed below. It is also to be understood that SRAM cells having three or more ports can be constructed slightly differently than the dual-port SRAM shown in the accompanying drawings, and those skilled in the art will understand how to construct these three-port or more port SRAMs for the following disclosure.

[0016] In addition, although SRAM cells are used in the following examples, it is to be understood that the disclosed memory cells for computing and processing arrays using the memory cells can be implemented using various different types of memory cells (including DRAM, CAM, non-volatile memory cells, and non-volatile memory devices), and these implementations using the various types of memory cells fall within the scope of the present disclosure.

[0017] Figure 2 A first embodiment of a dual-port SRAM cell 20 that can be used for computing is shown, which overcomes the Figure 1The disadvantages of the typical SRAM cell shown in FIG. 1 are as follows. The dual-port SRAM cell may include Figure 2 The two access transistors M23 and M24 coupled together and the two cross-coupled inverters I21 and I22 are shown to form an SRAM cell. The SRAM cell can be operated as a storage latch and can have a read port and a write port so that the SRAM cell is a dual-port SRAM cell. The two inverters are cross-coupled because the input of the first inverter is connected to the output of the second inverter, and the output of the first inverter is coupled to the input of the second inverter, as shown in FIG. Figure 2 The write word line carries the signal and is called WE (see Figure 2 ), while the write bit line and its complement are referred to as WBL and WBLb, respectively. The write word line (WE) is coupled to the gate of each of the two access transistors M23, M24 that are part of the SRAM cell. Figure 2 As shown, each of the write bit line and its complement (WBL and WBLb) is coupled to the source of the corresponding access transistor M23, M24, and the drain of each of these access transistors M23, M24 is coupled to each side of the cross-coupled inverter (in Figure 2 are coupled as marked D and Db).

[0018] Figure 2 The circuit in may also have a read word line RE, a read bit line RBL, and a read port formed by transistors M21, M22 coupled together to form an isolation circuit. The read word line RE may be coupled to the gate of transistor M21 forming part of the read port, and the read bit line is coupled to the drain terminal of transistor M21. The gate of transistor M22 may be coupled to the Db output terminal from the cross-coupled inverters I21, I22, and the source of transistor M22 may be coupled to ground.

[0019] In operation, the dual-port SRAM cell can read data stored in the latch using a read word line (RE) for addressing / activating the dual-port SRAM cell and a signal on a read bit line (RBL) for reading data stored in the dual-port SRAM cell. The dual-port SRAM cell can write data into the dual-port SRAM cell by addressing / activating the dual-port SRAM cell using a signal on a write word line (WE) and then writing the data into the dual-port SRAM cell using a write bit line (WBL, WBLb).

[0020] During read, multiple cells can be switched on (where Figure 2 Only a single unit is shown, but Figure 3Multiple cells are shown in FIG. 1 ) to perform an AND function between the data stored in the cells that are turned on. For example, Figure 3 Some cells in the columns of the processing array 30 in (such as cell 00, ..., cell m0) can be activated by the RE signal for each of these cells. Therefore, at the beginning of the read cycle, RBL is precharged to high, and if the Db signals of all cells turned on by RE are "0", RBL remains high, because although the gate of transistor M21 is turned on by the RE signal, the gate of M22 is not turned on because the Db signal is low. As a result, the RBL line is not connected to the ground to which the source of transistor M22 is connected, and the RBL line is not discharged. Cell 20 can be operated as a dual-port SRAM cell. The write operation is activated by WE, and the data is written by toggling of WBL and WBLb. The read operation is activated by RE, and the read data is accessed on RBL. Cell 20 can also be used for calculations, where RBL is also used for logical operations. If the Db signal of any or all of the cells is "1", RBL is discharged to 0 because the gate of M22 is turned on and the RBL line is connected to ground. As a result, RBL=NOR(Db0, Db1, etc.), where Db0, Db1, etc. are the complementary data of the SRAM cells that have been turned on by the RE signal. Alternatively, RBL=NOR(Db0, Db1, etc.)=AND(D0, D1, etc.), where D0, D1, etc. are the real data of the cells that have been turned on by the RE signal.

[0021] like Figure 2 As shown, the Db signal of cell 20 can be coupled to the gate of transistor M22 to drive the RBL line. However, unlike the typical 6T cell, the Db signal is isolated from the RBL line and its signal / voltage level by means of transistors M21 and M22 (together forming an isolation circuit). Since the Db signal / value is isolated from the RBL line and signal / voltage level, compared to Figure 1 In a typical SRAM cell, the Db signal is not susceptible to the lower bit line level caused by multiple "0" data stored in multiple cells. Figure 2 As a result, the cell (and a device composed of a plurality of cells) provides more operands for Boolean functions (such as the AND function described above, and the NOR function / OR function / NAND function described below) and search operations because there is no limit to how many cells can be turned on to drive the RBL. In addition, in Figure 2In the disclosed cell, the RBL line is precharged (without using a static pull-up transistor like a typical 6T cell), so the cell can provide faster sensing because all the current generated by the cell is used to discharge the bit line capacitance, and no current is consumed by the static pull-up transistor, so that the bit line discharge rate can be more than twice faster than a typical SRAM cell. Without the additional current consumed by the static pull-up transistor, sensing for the disclosed cell also requires less power, and the discharge current is reduced by more than half.

[0022] Figure 2 The write port of the cell in is operated in the same manner as the 6T typical SRAM cell described above. Figure 2 The write cycle and selective write cycle of the cell in has the same restrictions as the 6T cell discussed above. In addition to the above AND functions, Figure 2 The SRAM cell 20 in the embodiment can also perform the NOR function by storing inverted data. Specifically, if D is stored at the gate of M22 instead of Db, then RBL=NOR(D0, D1, etc.). It is understood by those skilled in the art that Figure 2 The unit configuration shown may be altered slightly to accomplish this, but such modifications fall within the scope of the present disclosure.

[0023] Figure 3 Show that it can be incorporated Figure 2 , Figure 4 , Figure 6 ,or Figure 7 A dual-port SRAM cell processing array device 30, wherein each cell such as cell 00, ..., cell 0n and cell m0, ..., cell mn is Figure 2 , Figure 4 , Figure 6 ,or Figure 7 The unit is formed as shown. Figure 3 Array 30 is an array of cells arranged as shown. Processing array 30 can perform calculations using the computing power of dual-port SRAM cells as described above. Array device 30 can be formed by M word lines (such as RE0, WE0, ..., REm, WEm) and N bit lines (such as WBL0, WBLb0, RBL0, ..., WBLn, WBLbn, RBLn). Array device 30 may also include a word line generator (WL generator) that generates word line signals, and a plurality of bit line read / write logics (such as BL read / write logic 0, ..., BL read / write logic n) that perform read operations and write operations using bit lines. Depending on the use of processing array device 30, array device 30 may be manufactured on an integrated circuit, or may be integrated into another integrated circuit.

[0024] In a read cycle, the word line generator may generate one or more RE signals in one cycle to turn on / activate one or more cells, and the RBL lines of the cells activated by the RE signals constitute an AND function or an NOR function, the output of which is sent to the corresponding BL read / write logic. The BL read / write logic processes the RBL result (the result of the AND operation or the NOR operation) and sends the result back to its WBL / WBLb for use in / write back to the same cell, or sends the result back to the adjacent BL read / write logic for use in / write back to the adjacent cell, or sends it out of the processing array. Alternatively, the BL read / write logic can store the RBL result from its own bit line or from the adjacent bit line in a latch within the BL read / write logic so that during the following cycle or later cycles, the read / write logic can perform logic using the latched data as the RBL result.

[0025] In the write cycle, the word line generator generates one or more WE signals for the cell to which data is to be written. The BL read / write logic processes the write data from its own RBL, or the write data from the adjacent RBL, or the write data from outside the processing array. The ability of the BL read / write logic to process data from adjacent bit lines means that data can be shifted from one bit line to an adjacent bit line, and one or more bit lines or all bit lines in the processing array can be shifted at the same time. The BL read / write logic can also decide not to write for a selective write operation based on the RBL result. For example, if RBL=1, the data on the WBL line can be written to the cell. If RBL=0, no write operation is performed.

[0026] Figure 4 A second embodiment of a dual-port SRAM cell 40 is shown that can be used for computing. The read port operation of the cell is similar to Figure 2 The unit in is the same, but with improved write port operation as described above. Figure 4 In the unit, a pair of cross-coupled inverters I41 and I42 form a latch as a storage element. Figure 4 The cells in have the same isolation circuits (M41, M42) for the read bit lines as described above.

[0027] Transistors M43, M44 and M45 form a write port. The unit can be Figure 3 The array device 30 is shown arranged with WE running horizontally and WBL and WBLb running vertically. Figure 51 shows the truth table of the write port. If WE is 0, no write is performed. If WE is 1, the storage node D and its complement Db are written through WBL and WBLb. Specifically, if WBL=1 and WBLb=0, then D=1 and Db=0; and if WBL=0 and WBLb=1, then D=0 and Db=1. If both WBL and WBLb are 0, no write is performed, and the data stored is the data stored in the storage element before the current write cycle (such as Figure 5 D(n-1) shown). Therefore, in the case of WBL=WBLb=0, the cell can perform a selective write function. In the cell, M45 is activated by a write word line (WE) signal coupled to the gate of M45, and M45 pulls the sources of transistors M43 and M44 to ground.

[0028] See also Figure 4 ,and Figure 2 Unlike the dual-port cell in , the WBL and WBLb lines of this cell are driving the gates of transistors M44 and M43, not the sources. Therefore, the drive strength of WBL and WBLb is not limited by the number of cells that are turned on. In a selective write operation, WBL and WBLb do not require strong devices to maintain the WBL and WBLb signal levels, and there is no limit to how many cells can be turned on. As Figure 2 The unit in Figure 4 The units in can also be used in Figure 3 in the processing array.

[0029] During a write cycle, the WE signal of each unselected cell is 0, but one of the signals on WBL and WBLb is 1. For example, in Figure 3 In , for the cell m0 to be written, WEm is 1, and for the cell 00 not to be written, WE0 is 0. Figure 4 In the example, D and Db of the unselected cells should maintain their original values. However, if D of the unselected cells stores "1", and the drain of M45 is 0 and WBLb is 1, the gate of the access transistor M43 is turned on, and the capacitance charge of the node D is the charge shared with the capacitance of the drain of the node N from M45 and the sources of M43 and M44. The high level of D is reduced by this charge sharing, and if the node N capacitance is high enough, the level will be reduced so that the I41 and I42 latches jump to the opposite data.

[0030] Figure 6 A third embodiment of a dual port SRAM cell 60 is shown which may be used for computing. As with the other embodiments above, this cell may be used in the processing array 30 described above. Figure 6The cells in have the same isolation circuits (M61, M62) for the read bit lines as described above. Cell 60 also has the same isolation circuits (M61, M62) for the read bit lines as described above. Figure 4 The same cross-coupled inverters I61, I62 as in the case of the cells in FIG. 1 and two access transistors M63, M64 having their respective gates coupled to the write bit line and the complementary write bit line. Figure 6 In the unit, Figure 4 The M45 transistor in the embodiment can be divided into a first write port transistor M65 and a second write port transistor M66, so that M63, M64, M65 and M66 form a write port circuit. Therefore, node D can only share charge with the drain of M65 and the source of M63, while the source of M64 no longer affects node D, and the high voltage level of node D can be kept high to avoid data jumping to the opposite state. This improves the disadvantage of charge sharing of unselected cells. Used to modify Figure 4 Another way to increase the capacitance of node D by having I41 and I42 with larger gate sizes is to increase the capacitance of node D. Note that node Db is not susceptible to the additional capacitance of M42.

[0031] Figure 7 A fourth embodiment of a dual port SRAM cell 70 that can be used for computing is shown. As with the other embodiments, this cell can be used in the processing array 30 described above. Figure 7 The cell in has the same isolation circuit (M71, M72) for the read bit line as described above. Cell 70 also has the same cross-coupled inverters I71, I72 as above, and two access transistors M75, M76, the two access transistors M75, M76 having their respective gates coupled to the write word line WE. The SRAM cell may also include transistors M73, M74, the gates of which are coupled to the write bit line and the complementary write bit line. Transistors M73, M74, M75 and M76 form a write port circuit. Cell 70 and Figure 6 Unit 60 in operates similarly.

[0032] return Figure 4 , Figure 6 and Figure 7 , latching devices (e.g., Figure 4 I41 and I42 in can be simple inverters. To complete a successful write, Figure 4The driving strength of the series transistors M43 and M45 in the I42 needs to be stronger than the pull-up PMOS transistor of I42, and the ratio needs to be about 2 to 3 times, so that the driving strength of the transistors M43 and M45 can be optimally 2-3 times as strong as the pull-up PMOS transistor of I42. In advanced technologies such as 28nm or better, the layout of the PMOS transistors and the NMOS transistors preferably have equal lengths. Therefore, when using 28nm or better feature sizes to manufacture Figure 4 , Figure 6 and Figure 7 When the PMOS transistors of I41 and I42 are connected in series, they may be two or more PMOS transistors, such as Figure 8 For ease of layout, one or more of the series PMOS transistors may be connected to ground, such as Fig. 9 shown. Figure 8 and Fig. 9 The latched inverter in can be used in all embodiments of the above-mentioned SRAM cell.

[0033] return Figure 2 , the read port transistors M21 and M22 (isolation circuit) can be PMOS instead of Figure 2 NMOS shown. If the M21 and M22 transistors are PMOS (where the source of M22 is coupled to VDD), RBL is precharged to 0; and if Db of one or more cells that are turned on is 0, RBL is 1; and if Db of all cells is 1, RBL is 0. In other words, RBL=NAND(Db0, Db1, etc.)=OR(D0, D1, etc.), where D0, D1, etc. are the real data of the cells that are turned on, and Db0, Db1, etc. are complementary data. The NAND function can also be performed by storing inverted data, so that if D is stored at the gate of M22 instead of Db, then RBL=NAND(D0, D1, etc.). The read port formed by the PMOS can be used in Figure 2 , Figure 4 , Figure 6 or Figure 7 All dual-port cells can be used for OR and NAND functions.

[0034] Figure 3 The processing array 30 in Figure 3 The array shown has dual-port SRAM cells in different configurations. For example, Figure 3 The processing array 30 in FIG. 3 may have some dual-port SRAM cells with NMOS read port transistors and some dual-port SRAM cells with PMOS read port transistors. The processing array 30 may also have other combinations of dual-port SRAM cells.

[0035] Depend on Figure 2 , Figure 4 , Figure 6 and Figure 7 The processing array composed of dual-port SRAM cells shown in FIG. Figure 3 An example of an application of (shown in ) is a search operation. For a 1-bit search operation, 2 cells store real (D) data and complementary (Db) data along the same bit line. The search is performed by inputting the search keyword S as the RE for the real data, and inputting the complement of S, Sb, as the RE for the complementary data. If S=1, Sb=0, then RBL=D=AND(S,D). If S=0, Sb=1, then RBL=Db=AND(Sb,Db). Therefore RBL=OR(AND(S,D),AND(Sb,Db))=XNOR(S,D). In other words, if S=D, then RBL=1, and if S≠D, then RBL=0.

[0036] As another example, for an 8-bit word search, the data of the 8-bit word is stored in 8 cells D[0:7] along the same bit line, and the complementary data of the 8-bit word is also stored in another 8 cells Db[0:7] along the same bit line as the real data. The search keyword can be input as the 8-bit S[0:7] of the RE applied to the real data cell D[0:7], and the 8-bit Sb[0:7] (complement of S) of the RE applied to the complementary data cell Db[0:7]. The bit lines can be written as RBL = AND(XNOR(S[0], D[0]), XNOR(S[1], D[1]), ..., XNOR(S[7], D[7]). If all 8 bits match, then RBL is 1. If any one or more bits do not match, then RBL = 0. By placing multiple data words along the same word line, and placing each word on a bit line in parallel, a parallel search can be performed in one operation. In this way, the search results for each bit line in the array are generated in one operation.

[0037] Depend on Figure 2 , Figure 4 , Figure 6 and Figure 7 The processing array composed of dual-port SRAM cells shown in FIG. Figure 3 In other words, multiple RE and WE signals on the same bit line can be turned on at the same time to simultaneously execute read logic on the read bit line and write logic on the write bit line. Figure 1 The performance of the cells and processing arrays is improved over the typical single-port SRAM shown.

[0038] Therefore, a dual-port static random access memory computing unit is disclosed, which has an SRAM cell with a latch, a read port for reading data from the SRAM cell and a write port for writing data to the SRAM cell, and an isolation circuit, wherein the isolation circuit isolates a data signal representing a piece of data stored in the latch of the SRAM cell from a read bit line. The read port may have a read word line coupled to the isolation circuit and activating the isolation circuit and a read bit line coupled to the isolation circuit, and the write port has a write word line, a write bit line, and a complementary write bit line coupled to the SRAM cell. In the unit, the isolation circuit may also include: a first transistor, whose gate is coupled to the read word line; and a second transistor, whose gate is coupled to the data signal, and the first transistor and the second transistor of the isolation circuit are both NMOS transistors or both PMOS transistors. The data signal of the unit may be a data signal or a complementary data signal. The SRAM cell may also have: a first inverter and a second inverter, the first inverter having an input and an output, the second inverter having an input coupled to the output of the first inverter and an output coupled to the input of the first inverter; a first access transistor coupled to the input of the first inverter and the output of the second inverter, and coupled to a write bit line; and a second access transistor coupled to the output of the first inverter and the input of the second inverter, and coupled to a complementary write bit line. The write port may also include a write word line coupled to the gate of the first access transistor and the gate of the second access transistor, and a write bit line and a complementary write bit line respectively coupled to the source of each access transistor.

[0039] In another embodiment, the SRAM cell further includes: a first inverter and a second inverter, the first inverter having an input and an output, the second inverter having an input coupled to the output of the first inverter and an output coupled to the input of the first inverter; a first access transistor coupled to the input of the first inverter and the output of the second inverter, and a gate of the first access transistor coupled to a write bit line; and a second access transistor coupled to the output of the first inverter and the input of the second inverter, and a gate of the second access transistor coupled to a complementary write bit line. In other embodiments, the write port further includes a write word line coupled to the gate of the write port transistor, and a drain of the write port transistor is coupled to a source of the first access transistor and a source of the second access transistor.

[0040] In yet another embodiment, the SRAM cell further comprises: a first inverter having an input and an output, and a second inverter having an input coupled to the output of the first inverter and an output coupled to the input of the first inverter; a first access transistor coupled to the input of the first inverter and the output of the second inverter, and having a gate coupled to a write bit line; and a second access transistor coupled to the output of the first inverter and the input of the second inverter, and having a gate coupled to a write complementary bit line. In this embodiment, the write port further comprises a write word line coupled to the gate of each of the first write port transistor and the second write port transistor, the drain of the first write port transistor is coupled to the source of the first access transistor, and the drain of the second write port transistor is coupled to the source of the second access transistor.

[0041] In another embodiment, the SRAM cell further includes: a first inverter having an input and an output, and a second inverter having an input coupled to the output of the first inverter and an output coupled to the input of the first inverter; a first access transistor coupled to the input of the first inverter and the output of the second inverter, and a gate of the first access transistor coupled to a write word line; and a second access transistor coupled to the output of the first inverter and the input of the second inverter, and a gate of the second access transistor coupled to the write word line. In this embodiment, the write port further includes: a first write port transistor having a gate coupled to a complementary write bit line and a drain coupled to a source of the first access transistor; and a second write port transistor having a gate coupled to the write bit line and a drain coupled to a source of the second access transistor.

[0042] Each of the different embodiments of the dual-port static random access memory computing unit can perform selective write operations and can perform Boolean AND operations, Boolean NOR operations, Boolean AND NOR operations, or Boolean OR operations. Each of the different embodiments of the dual-port static random access memory computing unit can also perform search operations.

[0043] A processing array is also disclosed, which has: a plurality of dual-port SRAM cells arranged in an array; a word line generator, which is coupled to a write word line signal and a read word line signal of each dual-port SRAM cell in the array; and a plurality of bit line read and write logic circuits, which are coupled to a read bit line, a write bit line, and a complementary write bit line of each dual-port SRAM cell. In the processing array, each dual-port SRAM cell is coupled to a write word line and a read word line, the signals of which are generated by the word line generator, and each dual-port SRAM cell is also coupled to a read bit line, a write bit line, and a complementary write bit line sensed by one of the plurality of bit line read and write logic circuits, and each dual-port SRAM cell has an isolation circuit, which isolates a data signal representing a piece of data stored in a latch of the SRAM cell from the read bit line. In the processing array, one or more of the dual-port SRAM cells are coupled to a read bit line and perform a computing operation. The processing array can utilize the dual-port SRAM cells disclosed above. The processing array can perform selective write operations and can perform Boolean AND operations, Boolean OR operations, Boolean AND NOT operations, or Boolean OR operations. The processing array can also perform search operations. The processing array can also perform parallel shift operations on one or more bit lines or all bit lines simultaneously to shift data from one bit line to an adjacent bit line.

[0044] As described above, the disclosed computing SRAM cell and processing array can be implemented using SRAM cells with more than 2 ports (such as 3-port SRAM, 4-port SRAM, etc.). For example, the SRAM computing cell can be a 3-port cell with 2 read ports and 1 write port. In this non-limiting example, the 3-port SRAM cell can be used to more efficiently perform operations such as Y=OR(AND(A,B), AND(A,C)). In the case of using a 3-port SRAM, the value of the variable A is used twice because 2 read ports are used. In this example operation, Y can be calculated in one cycle where the result of AND(A,B) is on RBL1 and the result of AND(A,C) is on RBL2; and in the same cycle, RBL2 data can be sent to RBL1 to complete the OR operation to produce the final result. Therefore, compared to 2 cycles of a dual-port cell, the logic equation / operation can be completed in 1 cycle where the word line is triggered once to produce a result. Similarly, a 4-port SRAM cell can also be used, and the present disclosure is not limited to any specific number of ports of the SRAM cell.

[0045] Fig.10 An embodiment of a dual-port SRAM cell 100 is shown that can be used for computing. Fig.10 The units in Figure 2The same isolation circuits (M101, M102) for reading bit lines, the same storage latches (I101, I102), the same access transistors (M103, M104), the same write bit line and complementary write bit line, and the same read word line. However, in Fig.10 the selective write implementation is different. The active low write word line WEb is connected to one input terminal of the NOR gate (I103), and the other input terminal is connected to the active low selective write control signal SWb to control the gates of the access transistors M103 and M104. SWb runs in the same direction as the bit line. In this implementation, writing to the cell occurs only when both the write word line and the selective write control signal are active.

[0046] Fig.11 Another embodiment of the dual-port SRAM cell 110 that can be used for computing is shown. Fig.11 Similar to Fig.10 where the selective write control signal SW is combined with the write word line WE to control the selective write operation. Two access transistors M113 and M115 are connected in series to couple the storage latch to the write bit line WBL, and similarly, two access transistors M114 and M116 are connected in series to couple the storage latch to the complementary write bit line WBLb. The gates of M113 and M114 are coupled to WE, and the gates of M115 and M116 are coupled to SW. SW runs in the same direction as the bit line. Writing to the cell occurs only when both the write word line and the selective write signal are active.

[0047] For purposes of explanation, the foregoing description has been presented with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of the disclosure and its practical application, thereby enabling others skilled in the art to best utilize the disclosure and various embodiments with various modifications suited to the particular use contemplated.

[0048] The systems and methods disclosed herein can be implemented by one or more components, systems, servers, household appliances, other sub-components, or can be distributed among these elements. When implemented as a system, such a system can include and / or involve components such as software modules, general-purpose CPUs, RAM, etc., especially those found in a general-purpose computer. In embodiments where the innovation lies in the server, such a server can include or involve components such as CPUs, RAM, etc., those found in a general-purpose computer.

[0049] In addition, the systems and methods herein can be implemented by implementations with completely different or completely different software, hardware and / or firmware components other than those set forth above. With respect to these other components (e.g., software, processing components, etc.) and / or computer-readable media related to or embodying the present invention, for example, the innovative aspects herein can be implemented consistently with many general-purpose or special-purpose computing systems or configurations. Various exemplary computing systems, environments, and / or configurations applicable to the innovations herein may include, but are not limited to: software or other components implemented within or on a personal computer, a server or server computing device (such as a routing / connectivity component), a handheld or laptop device, a multiprocessor system, a microprocessor-based system, a set-top box, a consumer electronic device, a network PC, other existing computer platforms, a distributed computing environment including one or more of the above-mentioned systems or devices, etc.

[0050] In some instances, various aspects of the system and method can be implemented or performed by logic and / or logic instructions including program modules, which are, for example, performed in association with these components or circuits. Typically, program modules can include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific instructions herein. The present invention can also be implemented in the environment of distributed software, computer or circuit settings, wherein the circuits are connected by communication buses, circuits or links. In a distributed setting, control / instructions can appear in both local and remote computer storage media including memory storage devices.

[0051] The software, circuits and components herein may also include and / or utilize one or more types of computer-readable media. Computer-readable media may be any available media resident on, associated with, or accessible to these circuits and / or computing components. As an example and not limitation, computer-readable media may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, digital versatile disk (DVD) or other optical storage, tape, disk storage or other magnetic storage device, or any other medium that can be used to store desired information and can be accessed by computing components. Communication media may include computer-readable instructions, data structures, program modules and / or other components. In addition, communication media may include wired media, such as a wired network or a direct wired connection, but any such type of media herein does not include transient media. Any combination of the above is also included within the scope of computer-readable media.

[0052] In this specification, the term assembly, module, equipment, etc. may refer to any type of logic or functional software element, circuit, block and / or process that can be implemented in various ways. For example, the functions of various circuits and / or blocks can be combined with each other into any other number of modules. Each module can even be implemented as a software program stored on a tangible memory (e.g., random access memory, read-only memory, CD-ROM memory, hard disk drive, etc.) to be read by a central processing unit to realize the innovative functions herein. Alternatively, a module may include programming instructions transmitted to a general-purpose computer or processing / graphic hardware by a transmission carrier. Moreover, a module may be implemented as a hardware logic circuit that realizes the functions covered by the innovation herein. Finally, a module may be implemented using a dedicated instruction (SIMD instruction), a field programmable logic array, or any mixture thereof that provides desired level performance and cost.

[0053] As disclosed herein, features consistent with the present disclosure may be implemented via computer hardware, software, and / or firmware. For example, the systems and methods disclosed herein may be embodied in various forms, including, for example, data processors (such as computers that also include a database), digital electronic circuits, firmware, software, or combinations thereof. In addition, although some of the disclosed embodiments describe specific hardware components, the systems and methods consistent with the innovation herein may be implemented using any combination of hardware, software, and / or firmware. In addition, the above-mentioned features and other aspects and principles of the innovation herein may be implemented in various environments. Such environments and related applications may be specifically constructed to perform various routines, processes, and / or operations according to the present invention, or they may include general-purpose computers or computing platforms that are selectively activated or reconfigured by code to provide the necessary functions. The processes disclosed herein are not inherently related to any particular computer, network, architecture, environment, or other device, but may be implemented by an appropriate combination of hardware, software, and / or firmware. For example, various general-purpose machines may be used with programs written in accordance with the teachings of the present invention, or it may be more convenient to construct a dedicated device or system to perform the desired methods and techniques.

[0054] Various aspects of the methods and systems described herein, such as logic, may also be implemented as functions programmed into any of a variety of circuits, including programmable logic devices ("PLDs") such as field programmable gate arrays ("FPGAs"), programmable array logic ("PAL") devices, electrically programmable logic and memory devices, standard cell-based devices, and application specific integrated circuits. Some other possibilities for implementing various aspects include: memory devices, microcontrollers with memory such as EEPROM, embedded microprocessors, firmware, software, and the like. In addition, various aspects may be embodied in microprocessors with software-based circuit simulation, discrete logic (sequential and combinatorial), custom devices, fuzzy (neural) logic, quantum devices, and a mix of any of the above device types. The underlying device technology may be provided in a variety of component types, such as metal oxide semiconductor field effect transistor ("MOSFET") technology such as complementary metal oxide semiconductor ("CMOS"), bipolar technology such as emitter coupled logic ("ECL"), polymer technology (e.g., silicon conjugated polymers and metal conjugated polymer-metal structures), mixed analog and digital, and the like.

[0055] It should also be noted that any number of combinations of data and / or instructions implemented in hardware, firmware, and / or various machine-readable or computer-readable media can be used to enable various logics and / or functions disclosed herein, depending on their behavior, register transfer, logic components, and / or other characteristics. Computer-readable media that can embody such formatted data and / or instructions include, but are not limited to, various forms of non-volatile storage media (e.g., optical, magnetic, or semiconductor storage media), but again do not include transient media. Unless the context clearly requires, throughout the specification, the words "include", "comprise", etc. should be interpreted in an inclusive sense relative to the exclusive or exhaustive sense; that is, in the sense of "including but not limited to". Words using the singular or plural also include the plural or singular, respectively. In addition, the words "herein", "hereafter", "above", "below", and words of similar meaning are used as a whole in this application rather than referring to any particular part of this application. When the word "or" is used in relation to a list of two or more items, the word covers all of its following interpretations: any item in the list, all items in the list, and any combination of items in the list.

[0056] Although certain currently preferred embodiments of the present invention have been specifically described herein, it is obvious to those skilled in the art that changes and modifications may be made to the various embodiments shown and described herein without departing from the spirit and scope of the present invention. Therefore, the present invention is intended to be limited only to the extent required by applicable legal provisions.

[0057] While the foregoing has been made with reference to specific embodiments of the present disclosure, those skilled in the art will appreciate that changes may be made to the embodiments without departing from the principles and spirit of the present disclosure, the scope of which is defined by the appended claims.

Claims

1. A storage computing unit, include: A storage unit including a first node (D) and a second node (Db); Read ports, which include: a first transistor including a first terminal coupled to a read bit line (RBL), a gate coupled to a read word line (RE), and a second terminal; and a second transistor including a first terminal coupled to the second terminal of the first transistor, a gate coupled to the second node of the storage unit, and a second terminal coupled to ground; wherein, during a read operation, the read word line is asserted and read data appears on the read bit line; and Write ports, which include: a third transistor including a first terminal coupled to the first node of the storage cell, a gate coupled to a complement of a write bit line (WBLb), and a second terminal; a fourth transistor including a first terminal coupled to the second terminal of the third transistor, a gate coupled to a write word line (WE), and a second terminal coupled to ground; and a fifth transistor including a first terminal coupled to the second node of the storage cell, a gate coupled to a write bit line (WBL), and a second terminal; During a write operation, the write word line is enabled and data is written into the storage unit using the write bit line and the complement of the write bit line.

2. The storage computing unit according to claim 1, in, The second terminal of the fifth transistor is coupled to the second terminal of the third transistor and the first terminal of the fourth transistor.

3. The storage computing unit according to claim 2, in, When the read word line is valid and the read word lines coupled to other storage computing units coupled to the read bit line are valid, the read bit line contains the result of an AND operation performed on the data stored in the storage computing unit and the data stored in the other storage computing units.

4. The storage computing unit according to claim 2, in, When the read word line is valid and the read word lines coupled to other storage computing units coupled to the read bit line are valid, the read bit line contains the result of an OR-invalid operation performed on the complement of the data stored in the storage computing unit and the complement of the data stored in the other storage computing units.

5. The storage computing unit according to claim 1, wherein the write port further include: A sixth transistor includes a first terminal coupled to the second terminal of the fifth transistor, a gate coupled to the write word line, and a second terminal coupled to ground.

6. The storage computing unit according to claim 5, in, When the read word line is valid and the read word lines coupled to other storage computing units coupled to the read bit line are valid, the read bit line contains the result of an AND operation performed on the data stored in the storage computing unit and the data stored in the other storage computing units.

7. The storage computing unit according to claim 5, in, When the read word line is valid and the read word lines coupled to other storage computing units coupled to the read bit line are valid, the read bit line contains the result of an OR-invalid operation performed on the complement of the data stored in the storage computing unit and the complement of the data stored in the other storage computing units.

8. The storage computing unit according to claim 1, in, When the read word line is valid and the read word lines coupled to other storage computing units coupled to the read bit line are valid, the read bit line contains the result of an AND operation performed on the data stored in the storage computing unit and the data stored in the other storage computing units.

9. The storage computing unit according to claim 1, in, When the read word line is valid and the read word lines coupled to other storage computing units coupled to the read bit line are valid, the read bit line contains the result of an OR-invalid operation performed on the complement of the data stored in the storage computing unit and the complement of the data stored in the other storage computing units.

10. A system, include: A plurality of storage computing units arranged in a column and coupled to a read bit line (RBL), a write bit line (WBL) and a complement (WBLb) of the write bit line, wherein each of the plurality of storage computing units comprises: A storage unit including a first node (D) and a second node (Db); Read ports, which include: a first transistor including a first terminal coupled to the read bit line (RBL), a gate coupled to a read word line (RE), and a second terminal; and a second transistor including a first terminal coupled to the second terminal of the first transistor, a gate coupled to the second node of the storage unit, and a second terminal coupled to ground; wherein, during a read operation, the read word line is asserted and read data appears on the read bit line; and Write ports, which include: a third transistor including a first terminal coupled to the first node of the storage cell, a gate coupled to the complement of the write bit line (WBLb), and a second terminal; a fourth transistor including a first terminal coupled to the second terminal of the third transistor, a gate coupled to a write word line (WE), and a second terminal coupled to ground; and a fifth transistor including a first terminal coupled to the second node of the storage cell, a gate coupled to the write bit line (WBL), and a second terminal; During a write operation, the write word line is enabled and data is written into the storage unit using the write bit line and the complement of the write bit line. 11 . The system of claim 10 , wherein in each of the plurality of storage computing units, the second terminal of the fifth transistor is coupled to the second terminal of the third transistor and the first terminal of the fourth transistor.

12. The system according to claim 10, in, In each of the plurality of storage computing units, the write port further comprises: A sixth transistor includes a first terminal coupled to the second terminal of the fifth transistor, a gate coupled to the write word line, and a second terminal coupled to ground.

13. The system according to claim 10, in, The read bit lines contain results of an AND operation performed on data stored in the plurality of storage computing units when corresponding read word lines for the plurality of storage computing units are asserted.

14. The system according to claim 10, in, The read bit lines contain results of a NOR operation performed on the complement of data stored in the plurality of storage computing cells when corresponding read word lines for the plurality of storage computing cells are asserted.

15. The system according to claim 10, in, A first plurality of the plurality of storage computing units stores a set of data, and a second plurality of the plurality of storage computing units stores the complement of the set of data, and a search is performed by enabling corresponding read word lines of the first plurality of the plurality of storage computing units according to a search keyword (S) and enabling corresponding read word lines of the second plurality of the plurality of storage computing units according to the complement of the search keyword (Sb), wherein a first value on the read bit line indicates that the search keyword matches the set of data, and a second value on the read bit line indicates that the search keyword does not match the set of data.