Computing Memory Unit and Processing Array Device with Incomparable Write Ports
By using an incomparable write port technology in the SRAM cell, the problem that existing SRAM cells require a high write ratio in the write operation is solved, and a lower cost and more efficient write operation is achieved.
Patent Information
- Application Number
- CN202110166435.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-07
- Filing Date
- 2021-02-05
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-02-05
AI Technical Summary
Existing SRAM cells require strong write transistor strength when performing write operations, resulting in a high write ratio, increasing the complexity and cost of the cell.
Incomparably selective write is achieved by using an incomparable write port during the write operation, using the WBL or WBLb bit lines to write without overcoming the PMOS pull-up strength of the storage latch PMOS.
Reduces the size and complexity of the write port transistor, reduces the area and cost of the cell, while avoiding limitations in write comparison performance.
Smart Images

Figure CN113257305B_ABST
Abstract
Description
[0001] Related Applications
[0002] This application is a partial continuation of U.S. application Ser. No. 15 / 709,401, filed Sep. 19, 2017 (now published as U.S. Pat. No. 10,249,362, issued Apr. 2, 2019), and U.S. application Ser. No. 15 / 709,399, filed Sep. 19, 2017, both of which claim the benefit of U.S. Provisional Application No. 62 / 430,767, filed Dec. 6, 2016, and entitled "Computational Dual Port SRAM Cell And Processing Array Device Using The Dual Port SRAM Cells For Xor And Xnor Computations", the entire contents of all of which are incorporated herein by reference. TECHNICAL FIELD
[0003] The present disclosure generally relates to static random access memory cells that can be used for computing. BACKGROUND OF THE DISCLOSURE
[0004] Arrays of memory cells, such as dynamic random access memory (DRAM) cells, static random access memory (SRAM) cells, content addressable memory (CAM) cells, or non-volatile memory cells, are a well-known mechanism for storing digital bits of data in a variety of computer- or processor-based devices. The various computer- and processor-based devices can include computer systems, smart phone devices, consumer electronics, televisions, Internet switches and routers, and the like. Memory cell arrays are typically encapsulated in integrated circuits, and can also be encapsulated within integrated circuits that also have processing devices. Different types of typical memory cells have different capabilities and characteristics that distinguish each type of memory cell. For example, DRAM cells require longer access times and lose their data content unless refreshed periodically, but are relatively inexpensive to manufacture because the structure of each DRAM cell is relatively simple. On the other hand, SRAM cells have faster access times and do not lose their data content unless the SRAM cell is powered off, but are relatively more expensive because each SRAM cell is more complex than a DRAM cell. CAM cells have the unique function of being able to easily address content within the cell, but are more expensive to manufacture because each CAM cell requires more circuitry to implement this content addressing function.
[0005] A variety of computing devices that can perform calculations on digital binary data are also well known. Computing devices can include microprocessors, CPUs, microcontrollers, and so on. These computing devices are typically fabricated on integrated circuits, but can also be fabricated on integrated circuits that also have a certain amount of memory integrated thereon. In these well-known integrated circuits with computing devices and memory, the computing device performs calculations on digital binary data bits, and the memory is used to store various digital binary data, including for example instructions to be executed by the computing device and data to be operated on by the computing device.
[0006] Recently, devices that perform computing operations using memory arrays or memory cells have been introduced. In some of these devices, a processor array for performing calculations can be formed by memory cells. These devices can be referred to as in-memory computing devices.
[0007] Big data operations are data processing operations in which a large amount of data must be processed. Machine learning uses artificial intelligence algorithms to analyze data and typically requires a large amount of data to execute. Big data operations and machine learning are also often computationally intensive applications and often encounter input / output problems due to bandwidth bottlenecks between the computing device and the memory that stores the data. For example, these big data operations and machine learning applications can use the aforementioned in-memory computing devices because the in-memory computing devices perform calculations within the memory, thereby eliminating the bandwidth bottleneck.
[0008] An SRAM cell can be configured to perform Boolean operations such as AND, OR, NAND, and NOR, XOR, and NOR. Such an SRAM cell can also support selective write operations. However, a typical SRAM cell requires a write transistor that is stronger than the transistors in the storage latch to overwrite the stored data. The transistor strength ratio of the write transistor to the storage transistor can be referred to as the write ratio. For a typical SRAM cell, the write ratio is 2 to 3, which means the write transistor is 2 to 3 times the strength of the storage transistor for a successful write. Therefore, it is desirable to provide a computing memory cell that can perform write-less, which can be an SRAM cell, with a write port and that performs Boolean operations such as AND, OR, NAND, NOR, XOR (exclusive OR), and XNOR (exclusive NOR). BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 Shows a dual-port SRAM cell that can perform Boolean operations;
[0010] Figure 2 Shows an implementation of a processing array having Figure 1 the multiple SRAM cells shown in and performing a logic function;
[0011] Figure 3 Shows the write port truth table of a dual-port SRAM cell with a selective write function; Figure 1 of;
[0012] Figure 4 Shows an embodiment of a dual-port SRAM cell that can perform Boolean operations and write without selectivity.
[0013] Figure 5 Shows an embodiment of a 3-port SRAM cell that can perform basic Boolean operations, XOR and XNOR functions, and write without selectivity; and
[0014] Figure 6 Shows an embodiment of a processing array having multiple Figure 5 SRAM cells as shown in and performing basic Boolean operations, XOR and XNOR functions. DETAILED DESCRIPTION
[0015] The present disclosure is particularly applicable to CMOS-implemented memory cells and processing arrays with multiple such memory cells that are capable of performing logic functions with write ports without selectivity, and the present disclosure will be described in this context. However, it should be understood that since memory cells can be constructed using different processes and can have circuit configurations different from those of the circuits for performing logic functions disclosed below, the memory cells and processing arrays have greater utility and are not limited to the embodiments disclosed below, and are therefore also within the scope of the present disclosure. For illustrative purposes, dual-port SRAMs and 3-port cells are disclosed below and in the figures. However, it should be understood that SRAM computing cells and processing arrays can also be implemented using SRAM cells with more ports, and the present disclosure is not limited to the SRAM cells disclosed below. It should also be understood that SRAM cells with more ports can be constructed slightly differently from the SRAM cells shown in the figures, but those skilled in the art should understand how to construct those SRAM cells with more ports based on the disclosure below.
[0016] In addition, although SRAM cells are used in the following examples, it should be understood that the disclosed memory cells for computing and processing arrays using memory cells can be implemented using various different types of memory cells, including DRAM, CAM, non-volatile memory cells, and non-volatile memory devices, and these embodiments using various types of memory cells are within the scope of the present disclosure.
[0017] Figure 1 Shows a dual-port SRAM cell 10 that can be used for computing. The dual-port SRAM cell can include two cross-coupled inverters (the pair of transistors M17, M19 as one inverter and the pair of transistors M18 and M110 as the other inverter) that form a latch or storage cell, and asFigure 1 The access transistors M11 - M16 shown are coupled together to form an SRAM cell. The SRAM cell can be used as a storage latch and can have a read port and a write port, such that the SRAM cell is a dual - port SRAM cell. The two inverters are cross - coupled because the input of the first inverter is connected to the output of the second inverter, and the output of the first inverter is coupled to the input of the second inverter, as Figure 1 shown.
[0018] The write word line carries a signal and is referred to as WE (see Figure 1 ). The write bit line and its complementary write bit line are referred to as WBL and WBLb, respectively. The write word line (WE) is coupled to the gate of each of the two access transistors M15, M16 that are part of the SRAM cell. The write bit line and its complementary write bit line (WBL and WBLb) are each coupled to the gate of the corresponding access transistors M13, M14, as Figure 1 shown, and M13 is coupled to M15, and M14 is coupled to M16. The source of each of the transistors M13 and M14 is coupled to ground. The drain of each of these access transistors M15, M16 is coupled to each side of the cross - coupled inverter (labeled D and Db in Figure 1 ).
[0019] Figure 1 The circuit in may also have a read word line RE, a read bit line RBL, and a read port formed by the transistors M11, M12 coupled together to form an isolation circuit. The read word line RE can be coupled to the gate of the transistor M11 that forms part of the read port, while the read bit line is coupled to the drain terminal of the transistor M11. The gate of the transistor M12 can be coupled to the Db output of the cross - coupled inverter, and the source of the transistor M12 can be coupled to ground.
[0020] In operation, the dual - port SRAM cell can read the data stored in the latch on a signal on the read word line (RE) used to address / activate the dual - port SRAM cell and on the read bit line (RBL) used to read the data stored in the dual - port SRAM cell. The dual - port SRAM cell can write data into the dual - port SRAM cell by using the signal on the write word line (WE) to address / activate the dual - port SRAM cell and then using the write bit lines (WBL, WBLb) to write data into the dual - port SRAM cell.
[0021] During a read, multiple cells can be turned on (where only one cell is shown in Figure 1 but multiple cells are shown in Figure 2 ) to perform an AND function between the data stored in the turned - on cells. For example, Figure 2Several cells in a column of the processing array 20, such as cell 00, …, cell m0, can be activated by the RE signal of each of those cells. Thus, at the start of a read cycle, the RBL is precharged high, and if the Db signals of all cells turned on by RE are “0”, then the RBL remains high. Although the gate of transistor M11 is turned on by the RE signal, the gate of M12 is not turned on because the Db signal is low. Thus, the RBL line is not connected to the ground connected to the source of transistor M12, and the RBL line is not discharged. The write operation is activated by WE, and data is written by toggling WBL and WBLb. The read operation is activated by RE, and the read data is accessed on the RBL.
[0022] Cell 10 can be further used for calculations where the RBL is also used to perform logic operations. If the Db signals of any and all activated cells are “1”, then the RBL is discharged to 0 because the gate of M12 is turned on and the RBL line is grounded. Thus, RBL = NOR(Db0, Db1, etc.), where Db0, Db1, etc. are the complementary data of the SRAM cells that have been turned on by the RE signal. Alternatively, RBL = NOR(Db0, Db1, etc.) = AND(D0, D1, etc.), where D0, D1, etc. are the true data of the cells that have been turned on by the RE signal.
[0023] As Figure 1 shown, the Db signal of cell 10 can be coupled to the gate of transistor M12 to drive the RBL line. The Db signal is isolated from the RBL line and its signal / voltage level by transistors M11, M12 (which together form an isolation circuit). Since the Db signal / value is isolated from the RBL line and the signal / voltage level, the Db signal is not vulnerable to the low bit line level caused by multiple “0” data stored in multiple cells. Thus, for the cells in Figure 1 the number of cells that can be turned on to drive the RBL is not limited. Thus, since the number of cells that can be turned on to drive the RBL is not limited, the cells (and devices composed of multiple cells) provide more operands for Boolean functions, such as the AND function described above and the NOR / OR / NAND / XOR / XNOR functions described in the following cases: co-pending and co-owned U.S. application Ser. No. 15 / 709,401, filed Sep. 19, 2017 (currently published as U.S. Pat. No. 10,249,362 on Apr. 2, 2019) and U.S. application Ser. No. 15 / 709,399, filed Sep. 19, 2017, and U.S. Provisional Application No. 62 / 430,767, filed Dec. 6, 2016 (incorporated herein by reference). In addition to the AND function described above, Figure 1The SRAM cell 10 in [[ ]] can also perform a NOR function by storing inverted data. Specifically, if D rather than Db is stored at the gate of M12, then RBL = NOR(D0, D1, etc.).
[0024] Figure 2 illustrates a processing array device 20 with dual-port SRAM cells that can be combined Figure 1 where each cell, such as cell 00, ……, cell 0n and cell m0, ……, cell mn, is Figure 1 the cell shown in [[ ]]. These cells form a cell array arranged as Figure 2 shown. The processing array 20 can perform calculations using the computing power of the dual-port SRAM cells described above. The array device 20 can be formed by M word lines (e.g., RE0, WE0, ……, RE m, WE m) and N bit lines (e.g., WBL0, WBLb0, RBL0, ……, WBLn, WBLbn, RBLn). The array device 20 may also include a word line generator 24 (WL generator) that generates word line signals and a plurality of bit line read / write logics 26 (e.g., BL read / write logic 0, ……, BL read / write logic n) that perform read and write operations using the bit lines. Depending on the use of the processing array 20, the array device 20 can be fabricated on an integrated circuit or integrated into another integrated circuit.
[0025] During a read cycle, the word line generator 24 can generate one or more RE signals in one cycle to turn on / activate one or more cells, and the RBL lines of the cells activated by the RE signals form an AND or NOR function, and the output of the function is sent to the corresponding BL read / write logic (26o, ……, 26n). Each BL read / write logic 26 processes the RBL result (the result of the AND or NOR operation) and sends the result back to its WBL / WBLb for use / write back to the same BL, or sends the result to an adjacent BL read / write logic 26 for use / write back to an adjacent BL, or sends it outside the processing array. Alternatively, the BL read / write logic 26 can store the RBL result of its own bit line or an adjacent bit line in a latch within the BL read / write logic, such that during the next or subsequent cycle, the BL read / write logic 26 can perform logic using the latched data as the RBL result.
[0026] During a write cycle, the word line generator 24 generates one or more WE signals for the cells to which data is to be written. The BL read / write logic (26o, …, 26n) processes write data from its own RBL or from an adjacent RBL or from outside the processing array 20. The ability of the BL read / write logic 26 to process data from adjacent bit lines means that data can be shifted from one bit line to an adjacent bit line, and one or more or all of the bit lines in the processing array can be shifted simultaneously. The BL read / write logic 26 can also decide not to perform a write for a selective write operation based on the RBL result. For example, if RBL = 1, then the data on the WBL line can be written to the cell. If RBL = 0, then no write operation is performed.
[0027] Figure 3 Shows Figure 1 The write port truth table of the dual-port SRAM cell. If WE is 0, then no write is performed (as Figure 3 reflected by D(n - 1) shown in). If WE is 1, then the storage node D and its complementary node Db are written through WBL and WBLb. If WBL = 1 and WBLb = 0, then D = 1 and Db = 0. If WBL = 0 and WBLb = 1, then D = 0 and Db = 1. If both WBL and WBLb are 0, then no write is performed. Thus, this cell can perform a selective write function where WBL = WBLb = 0 and WE = 1.
[0028] Now refer to Figure 1 to describe the write operation of the circuit in more detail. When WE = 1 and either WBL or WBLb is 1, a write is performed. To write D from 1 to 0, and thus make WBLb = 1 and WBL = 0, M13 and M15 need to be turned on to overcome the strength of the PMOS transistor M19. In a 16 nm or more advanced process technology using FINFET transistors, PMOS transistors typically have almost the same drive strength as NMOS transistors, and the combined drive strength of M13 and M15 in series must be 3 times or more the drive strength of M19 to successfully perform a write. Therefore, both M13 and M15 must be 6 times the drive strength of M19. Similarly, both M14 and M16 must be 6 times the drive strength of M110. This makes the sizes of the M13, M14, M15, and M16 transistors extremely large, which in turn results in Figure 1 the extremely large size of the cell 10 of.
[0029] Figure 4 The circuit 40 in modifies the circuit shown in during a write operation to be ratio-less to improve the write port transistor size problem. Figure 1 The table in also applies to the circuit 40 in because Figure 3 The table in also applies to Figure 4 the circuit 40 in becauseFigure 4 The circuit 40 in Figure 1 The circuit 10 in FIG. 4 has the same components and operates in the same manner, but the circuit 40 has an incomparable write operation, as described below. The unit 40 can also replace the unit 10 and Figure 2 The processing array 20 in is used seamlessly.
[0030] exist Figure 4 In cell 40, if WE = 0, then Figure 4 The circuit in does not perform any write. If WE is 1, then the storage node D and its complement node Db are written through WBL and WBLb, where if WBL=1 and WBLb=0, then D=1 and Db=0, and if WBL=0 and WBLb=1, then D=0 and D=1. If both WBL and WBLb are 0, then no write is performed, so this unit 40 can perform a selective write function, where WBL=WBLb=0 and WE=1, just like Figure 1 As in circuit 10.
[0031] exist Figure 4 In the example, when WE=1, WBLb=1, and WBL=0, transistors M43 and M45 are turned on and data D is written from 1 to 0, and there is no pull-up strength of the series PMOS transistors M49 and M411 because transistor M411 is turned off when its gate is connected to WBLb which is 1. In addition, because WBL=0, transistor M412 is turned on and data D is pulled down to 0, so that Db is pulled up from 0 to 1 and the writing is completed. Figure 4 In the write operation of the circuit 40, since M43 and M45 pull down D without having to overcome the PMOS pull-up strength of the storage transistor, there is no write ratio in the write operation.
[0032] Similarly, when WE=1, WBLb=0 and 1, transistors M44 and M46 are turned on and data Db is written from 1 to 0 without overcoming the pull-up strength of the series PMOS transistors M410 and M412 because transistor M412 is disconnected when the gate is connected to WBL. Similarly, in this write operation, there is no write ratio in the write operation because M44 and M46 do not need to overcome the storage PMOS pull-up strength.
[0033] In this way, the write port transistors M43, M44, M45, and M46 can have the same minimum transistor size as the PMOS transistors M49, M410, M411, and M412. Thus, the size of cell 40 can be reduced, and the write port is not affected by the write ratio. It should be noted that when WE = 0, no writing is performed, but when WBLb or WBL is 1, M411 or M412 can be turned on. This can keep D or Db floating at 1, which is acceptable because the write cycle only lasts for a very short time, and nodes D and Db have sufficient capacitance to hold the change so that the value in the storage cell remains unchanged in this case. During normal operation in the non-write cycle, both WBLb and WBL are low to keep the cross-coupled transistors M47, M48, M49, and M410 operating as the cross-coupled latch of the SRAM cell 40.
[0034] In Figure 4 the circuit 40 shown, the series transistor pairs M49, M411 and M410, M412 can be swapped to achieve the same function. For example, M49 can have its gate connected to Db and itself coupled to VDD and the source of M411, while M411 has its gate connected to WBLb coupled to D. Similarly, the series transistor pairs M43, M45 and M44, M46 can be swapped to achieve the same function.
[0035] In summary, write without ratio is performed using the write bit line (WBL) or the complementary write bit line (WBLb) to write the "0" node of the storage latch with the pull-up transistor disabled and to write the "1" node of the storage latch with the pull-up transistor enabled. Figure 4 The cell 40 in Figure 1 can be used in the processing array 20 in Figure 2 in the same way as the cell 10 in
[0036] Figure 5 shows an embodiment of a 3-port SRAM cell 50 that can perform basic Boolean operations, XOR and XNOR functions, and write without ratio selectivity. The cell 50 has the same storage latch and write port circuitry as the cell 40, and thus has the same write without ratio selectivity operation as the cell 40. Compared with the cell 40 in Figure 4 the cell 40 in Figure 5Another read port is added to the cell 50 in FIG. 5A. Transistors M513 and M514 are added to form a second read port and an isolation circuit for the second read port. In this circuit 50, a supplementary read word line REb can be coupled to the gate of transistor M513 forming part of the read port, and a supplementary read bit line RBLb is coupled to the drain terminal of transistor M513. The gate of transistor M514 can be coupled to the D output of the cross-coupled inverter, and the source of transistor M514 can be coupled to ground.
[0037] During read, multiple cells can be switched on (where Figure 5 Only one unit is shown, but in Figure 6 Multiple cells are shown in the processing array 60 in the figure) to perform an AND function between the complementary data stored in the turned-on cells. During reading, the RBLb line is precharged to high. If the D signal of any and all cells activated is "1", then RBLb discharges to 0 because the gate of M514 is turned on and the RBLb line is grounded. Therefore, RBLb=NOR(D0, D1, etc.), where D0, D1, etc. are the data of the SRAM cells that have been turned on by the REb signal. Alternatively, RBLb=NOR(D0, D1, etc.)=AND(Db0, Db1, etc.), where Db0, Db1, etc. are the complementary data of the cells that have been turned on by the REb signal. Therefore, cell 50 is a 3-port SRAM cell with one write port (controlled by WE) and 2 read ports (controlled by RE and Reb), where RBL=AND(D0, D1, etc.) and RBLb=(D0b, D1b, etc.).
[0038] Figure 6 An embodiment of a processing array 60 is shown having a plurality of Figure 5 , separate sectors (segment 1 and sector 2, as shown), and each bit line (BL) read / write logic circuit system 64 (BL read / write logic 0, ..., BL read / write logic n for each bit line) located in the middle of each bit line. This processing array has a word line generator 62 that generates control signals (RE0, ..., REm, REb0, ..., REbm and WEO, ..., WEm), and each bit line has the two sectors. In one embodiment, sector 1 has RBLs1 and RBLs1b read bit lines (RBL0s1, ..., RBLns1 and RBL0s1b, ..., RBLns1b), wherein a number of cells (in Figure 6In the example in, units 00, …, unit 0n) are all connected to the BL read / write circuitry 64, and section 2 has RBLs2 and RBLs2b lines (RBL0s2, …, RBLns2 and RBL0s2b, …, RBLns2b), and several units in the RBLs2 and RBLs2b lines (in Figure 6 the example in are units m0, …, unit mn) are all connected to another input of the BL read / write circuitry 64.
[0039] During a read cycle, the word line generator can generate one or more RE, REb signals in one cycle to turn on / activate one or more units, and the RBL, RBLb lines of the units activated by the RE and REb signals form an AND or NOR function, and the output of the function is sent to the corresponding BL read / write logic 64 of each bit line. Each BL read / write logic 64 processes the RBL result (the result of the AND or NOR operation), and sends the result back to its WBL / WBLb for use / write back to the same unit, or sends the result to the adjacent BL read / write logic for use / write back to the adjacent unit, or sends it outside the processing array. Alternatively, the BL read / write logic 64 can store the RBL result of its own bit line or the adjacent bit line in a latch within the BL read / write logic, so that during the next or subsequent cycle, the read / write logic can perform logic using the latched data as the RBL result.
[0040] In using Figure 6 During the write cycle of the processing array in, the word line generator 62 generates one or more WE signals for the one or more units to which data is to be written. The BL read / write logic 64 processes the write data from its own RBL or from an adjacent RBL or from outside the processing array. The ability of the BL read / write logic 64 to process data from an adjacent bit line (note the connection between the bit line and each BL read / write logic 64) means that data can be shifted from one bit line to an adjacent bit line, and one or more or all of the bit lines in the processing array can be shifted simultaneously. The BL read / write logic 64 can also decide not to write for a selective write operation based on the RBL or RBLb result. For example, if RBL = 1, then the data on the WBL line can be written to the unit. If RBL = 0, then no write operation is performed.
[0041] SRAM Ultra-Low VDD Operation SRAM
[0042] The units 40 and 50 described herein are for computing memory applications, but Figure 4 and 5These cells in [the context] can be used as SRAM cells with great noise immunity and ultra-low VDD operation. Specifically, the VDD operation level can be as low as the threshold voltages of the NMOS and PMOS transistors of the cell.
[0043] Isolated storage latch: Read or write operations do not affect the stability of the storage latch. The VDD operation level for storage is as low as the threshold voltages of the NMOS and PMOS transistors to keep the cross-coupled latch in operation.
[0044] Buffered read: The read bit-line voltage level does not affect the stability of the storage node. The read bit-line is pre-charged high and discharged by turning on the read-port access transistor. The VDD operation level is as low as the threshold voltage of the read-port NMOS transistor.
[0045] Ratio-less write: Writing to the storage latch is performed by turning on only the NMOS or PMOS transistor of the write port without a write ratio. The VDD operation level is as low as the threshold voltages of the write-port NMOS and PMOS transistors.
[0046] For purposes of illustration, the foregoing description has been described with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive, nor is it intended to limit the disclosure to the precise forms disclosed. Given the above teachings, many modifications and variations are possible. The embodiments were chosen and described in order to best explain the principles of the disclosure and its practical application, thereby enabling others skilled in the art to best utilize the disclosure and various embodiments and to make various modifications to suit the particular purposes contemplated.
[0047] The systems and methods disclosed herein can be implemented via one or more components, systems, servers, devices, other sub-components, or distributed among these elements. When implemented as a system, such a system can in particular include or involve components such as software modules, general-purpose CPUs, RAM, etc. found in a general-purpose computer. In an embodiment where the innovation resides on a server, such a server can include or involve components such as CPUs, RAM, etc., components found in a general-purpose computer.
[0048] Additionally, in addition to the above, the systems and methods herein can be implemented via embodiments that utilize different or completely different software, hardware, and / or firmware components. Regarding such other components (e.g., software, processing components, etc.) and / or computer-readable media associated with or embodying the present invention, for example, aspects of the innovation herein can be implemented in accordance with many general-purpose or special-purpose computing systems or configurations. Various exemplary computing systems, environments, and / or configurations suitable for use with the innovation herein can include, but are not limited to: software or other components such as routing / connectivity components embodied within or on a personal computer, server, or server computing device, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, consumer electronic devices, network PCs, other existing computer platforms, distributed computing environments including one or more of the above systems or devices, etc.
[0049] In some instances, for example, aspects of the systems and methods can be implemented by or executed by logic and / or logical instructions that include program modules executed in association with such components or circuitry. Generally speaking, program modules can include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific instructions herein. The present invention can also be practiced in the case of distributed software, computers, or circuitry arrangements where the circuitry is connected via a communication bus, circuitry, or link. In a distributed setting, local and remote computer storage media including memory storage devices can have control / instructions.
[0050] The software, circuitry, and components herein may also include and / or utilize one or more types of computer-readable media. Computer-readable media can be any available media that resides on, is associated with, or can be accessed by such circuitry and / or computing components. By way of example and not limitation, computer-readable media may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, digital versatile disks (DVDs) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computing component. Communication media may include computer-readable instructions, data structures, program modules, and / or other components. In addition, communication media may include wired media such as a wired network or direct wired connection, but all such types of media herein do not include transient media. Combinations of any of the above are also included within the scope of computer-readable media.
[0051] In this specification, the terms component, module, device, etc. may refer to any type of logical or functional software element, circuitry, block, and / or process that can be implemented in various ways. For example, the functions of various circuits and / or blocks can be combined with each other into any other number of modules. Each module can even be implemented as a software program stored on a tangible memory (e.g., random access memory, read-only memory, CD-ROM memory, hard disk drive, etc.) for a central processing unit to read to implement the innovative functions herein. Or, a module may include programming instructions transmitted via a transmission carrier to a general-purpose computer or processing / graphics hardware. In addition, a module can be implemented as hardware logic circuitry that implements the functions covered by the innovation herein. Finally, a module can be implemented using special instructions (SIMD instructions), field-programmable logic arrays, or any combination capable of providing the desired level of performance and cost.
[0052] As disclosed herein, features consistent with the present disclosure can be implemented via computer hardware, software, and / or firmware. For example, the systems and methods disclosed herein can be embodied in various forms, including, for example, a data processor, such as a computer that also includes a database, digital electronic circuitry, firmware, software, or a combination thereof. Additionally, although some of the disclosed embodiments describe specific hardware components, the systems and methods consistent with the innovations herein can be implemented with any combination of hardware, software, and / or firmware. Further, the above-described features and other aspects and principles of the innovations herein can be implemented in a variety of environments. Such environments and related applications can be specifically constructed to execute various routines, processes, and / or operations in accordance with the present invention, or they can include general-purpose computers or computing platforms that are selectively activated or reconfigured by code to provide the desired functionality. The processes disclosed herein are not inherently related to any particular computer, network, architecture, environment, or other device, and can be implemented with a suitable combination of hardware, software, and / or firmware. For example, various general-purpose machines can be used with programs written in accordance with the teachings of the present invention, or it may be more convenient to construct a special-purpose device or system to perform the required methods and techniques.
[0053] Aspects of the methods and systems described herein, such as logic, can also be implemented as functionality programmed into any of a variety of circuitry, including programmable logic devices ("PLDs"), such as field programmable gate arrays ("FPGAs"), programmable array logic ("PAL") devices, electrically programmable logic and memory devices, and standard cell-based devices, as well as application specific integrated circuits. Some other possibilities for implementing aspects include: memory devices, microcontrollers with memory (such as EEPROM), embedded microprocessors, firmware, software, etc. Additionally, aspects can be embodied in a microprocessor having software-based circuit simulation, (sequential and combinational) discrete logic, custom devices, fuzzy (neural) logic, quantum devices, and combinations of any of the above device types. The underlying device technology can be provided in a variety of component types, such as metal oxide semiconductor field effect transistor ("MOSFET") technology (such as complementary metal oxide semiconductor ("CMOS")), bipolar technology (such as emitter coupled logic ("ECL")), polymer technology (e.g., silicon conjugated polymers and metal conjugated polymer-metal structures), hybrid analog and digital, etc.
[0054] It should also be noted that the various logics and / or functions disclosed herein, in terms of their behavior, register transfers, logic components, and / or other features, can be enabled using any amount of hardware, firmware combinations, and / or as data and / or instructions embodied in various machine-readable or computer-readable media. The computer-readable media in which such formatted data and / or instructions can be embodied include, but are not limited to, various forms of non-volatile storage media (such as optical, magnetic, or semiconductor storage media), but do not include transient media. Unless the context clearly requires otherwise, throughout the description, the words "comprise / comprising" and the like should be construed in an inclusive sense rather than an exclusive or exhaustive sense; that is, in the sense of "including but not limited to". The use of singular or plural words also correspondingly includes plural or singular. In addition, the words "herein", "hereinafter", "above", "below", and words of similar import refer to the application as a whole and not to any particular part of the application. When the word "or" is used in reference to a list of two or more items, this word covers all of the following interpretations: any one of the items in the list, all of the items in the list, and any combination of the items in the list.
[0055] Although certain presently preferred embodiments of the invention have been specifically described herein, those skilled in the art of the invention will appreciate that changes and modifications can be made to the various embodiments shown and described herein without departing from the spirit and scope of the invention. Therefore, it is desired that the invention be limited only to the extent required by the applicable legal rules.
[0056] Although the foregoing has referred to specific embodiments of the present disclosure, those skilled in the art should understand that changes can be made to these embodiments without departing from the principles and spirit of the present disclosure, the scope of which is defined by the appended claims.
Claims
1. A memory computing unit, which consists of the following components: An SRAM memory cell, which consists of two cross-coupled inverters forming a data port D and a complementary data port Db; A read port, which consists of the following components: (i) A first access transistor, including a first terminal coupled to a read bit line RBL, a gate coupled to a read enable line RE, and a second terminal ; And (ii) A second access transistor, including a first terminal coupled to the second terminal of the first access transistor, a gate coupled to the complementary data port, and a second terminal coupled to ground; And An asymmetric write port, which is coupled to the SRAM memory cell and provides write access to the SRAM memory cell. The asymmetric write port consists of the following components: (i) A third access transistor, including a first terminal coupled to a supply voltage VDD, a gate coupled to a write bit line WBL, and a second terminal coupled to the SRAM memory cell; (ii) A fourth access transistor, including a first terminal coupled to the supply voltage, a gate coupled to a complementary write bit line WBLb, and a second terminal coupled to the SRAM memory cell; (iii) A fifth access transistor, including a first terminal coupled to the data port, a gate coupled to a write enable line WE, and a second terminal; (iv) A sixth access transistor, including a first terminal coupled to the second terminal of the fifth access transistor, a gate coupled to the complementary write bit line, and a second terminal coupled to ground; (v) A seventh access transistor, including a first terminal coupled to the complementary data port, a gate coupled to the write enable line, and a second terminal; And (vi) An eighth access transistor, including a first terminal coupled to the second terminal of the seventh access transistor, a gate coupled to the write bit line, and a second terminal coupled to ground. The asymmetric write port allows data to be written into the SRAM memory cell by disconnecting the third access transistor from the supply voltage.
2. A processing array, which Includes: A plurality of memory cells arranged in an array, where each memory cell has a memory cell including a first PMOS transistor, a read port for reading data from the memory cell, and a write port for writing data into the memory cell. The read port buffers the memory cell according to a signal on at least one read bit line, and the read bit line is configured to provide read access to a piece of data stored in the memory cell; A word line generator, which is coupled to a read word line signal and a write word line signal for each memory cell in the array; And A plurality of bit line read and write logic circuits, which are coupled to the read bit line, write bit line, and complementary write bit line of each memory cell; Each memory cell is coupled to a write word line and a read word line whose signals are generated by the word line generator, and is also coupled to a read bit line, a write bit line (WBL), and a complementary write bit line sensed by one of the plurality of bit line read and write logic circuits; Each write port is a non-comparable write port that provides write access to the memory cell. The non-comparable write port allows data to be written into the memory cell by disconnecting only one transfer PMOS transistor that is coupled to the PMOS memory transistor at a node and is gate-controlled by a write bit line. And wherein two or more of the memory cells are coupled to at least one read bit line and are activated to perform a Boolean operation.
3. The processing array according to claim 2, wherein each non-comparable write port further includes a write bit line and a complementary write bit line, and the gate of the second PMOS transistor is connected to the write bit line.
4. The processing array according to claim 3, wherein the memory cell in each memory unit further includes a first inverter having an input and an output, and a second inverter having an input coupled to the output of the first inverter and an output coupled to the input of the first inverter. The first inverter includes a first PMOS transistor coupled to the second PMOS transistor, and the second inverter includes a third PMOS transistor coupled to a fourth PMOS transistor.
5. The processing array according to claim 3, wherein the read port further includes an isolation circuit that buffers the memory cell according to a signal on the at least one read bit line.
6. The processing array according to claim 2, which is capable of performing a selective write operation.
7. The processing array according to claim 5, wherein each memory unit further includes a second read port connected to a complementary read bit line; and wherein two or more of the memory cells are coupled to the complementary read bit line and are activated to perform another Boolean operation.
8. The processing array according to claim 4, wherein the fourth PMOS transistor is disconnected to cut off the second PMOS transistor.
9. The processing array according to claim 4, wherein the first PMOS transistor, the second PMOS transistor, the third PMOS transistor, and the fourth PMOS transistor have the same size.
10. A memory computing unit (50) comprising: An SRAM memory cell formed by two cross-coupled inverters forming a data port D and a complementary data port Db; a read port comprising: (i) a first access transistor including a first terminal coupled to a read bit line RBL, a gate coupled to a read enable line RE, and a second terminal ; (ii) a second access transistor including a first terminal coupled to the second terminal of the first access transistor, a gate coupled to the complementary data port, and a second terminal coupled to ground; (iii) a third access transistor including a first terminal coupled to a complementary read bit line RBLb, a gate coupled to a complementary read enable line REb, and a second terminal; and (ii) a fourth access transistor including a first terminal coupled to the second terminal of the third access transistor, a gate coupled to the data port, and a second terminal coupled to ground; An asymmetric write port, which is coupled to the SRAM memory cell and provides write access to the SRAM memory cell, the asymmetric write port is composed of the following: (i) a fifth access transistor, including a first terminal coupled to a supply voltage VDD, a gate coupled to a write bit line WBL, and a second terminal coupled to the SRAM memory cell; (ii) a sixth access transistor, including a first terminal coupled to the supply voltage, a gate coupled to a complementary write bit line WBLb, and a second terminal coupled to the SRAM memory cell; (iii) a seventh access transistor, including a first terminal coupled to the data port, a gate coupled to a write enable line WE, and a second terminal; (iv) an eighth access transistor, including a first terminal coupled to the second terminal of the seventh access transistor, a gate coupled to the complementary write bit line, and a second terminal coupled to ground; (v) a ninth access transistor, including a first terminal coupled to the complementary data port, a gate coupled to the write enable line, and a second terminal; and (vi) a tenth access transistor, including a first terminal coupled to the second terminal of the ninth access transistor, a gate coupled to the write bit line, and a second terminal coupled to ground, the asymmetric write port allows data to be written into the SRAM memory cell by disconnecting the fifth access transistor from the supply voltage.
Citation Information
Patent Citations
Computational memory cell and processing array device using the memory cells for XOR and XNOR computations
US10249362B2
Computational memory cell and processing array device using the memory cells for XOR and XNOR computations
US10998040B2
Computational memory cell and processing array device using memory cells
CN110291587A
Memory device having memory cells with write assist functionality
US20120212996A1
Static random access memory cell having improved write margin for use in ultra-low power application
US20160254045A1