Bitline-Storage Node Decoupled Bitline Sense Amplifier, Memory-in-Computing Memory, and In-Situ Matrix Computing Method

Through the bit line-storage node decoupling sensing amplifier and the cascading feedback bit line sensing amplifier, the system bottleneck caused by bit line current accumulation is solved, and efficient vector matrix multiplication calculation is realized in open bit line ReRAM, which improves hardware resource utilization and data transmission rate.

CN118748031BActive Publication Date: 2025-08-05HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410814184.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2025-08-05
Estimated Expiration
2044-06-24

AI Technical Summary

Technical Problem

In the prior art, the area and delay overhead caused by bit line current accumulation become a system bottleneck for vector matrix multiplication (VMM), and the existing sub-array microarchitecture without ADC cannot efficiently support the activation of multiple continuous word line operations at one time, resulting in insufficient hardware utilization and excessive delay.

Method used

The bit line-memory node decoupling sensing amplifier and the cascade feedback bit line sensing amplifier are adopted, combined with the memory microarchitecture of open bit line ReRAM, and the cascade feedback bit line sensing architecture is realized by decoupling bit line and storage nodes, which supports multiple bits to sense in the same row period, and uses the open bit line topology layout between adjacent subarrays to improve hardware resource utilization.

Benefits of technology

The function of "pre-charge once and read out multiple effective bits" is realized, which shortens the overall sensing delay, improves the execution performance of vector matrix multiplication, and improves the data transmission rate and hardware resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118748031B_ABST
    Figure CN118748031B_ABST
Patent Text Reader

Abstract

This invention discloses a bitline sense amplifier with bitline-storage node decoupling, a memory microarchitecture within a storage-computation system, and an in-situ matrix computation method. These methods, belonging to the field of storage-computation integration, include: a sense amplifier with a latch, an equalizer, and an unbalancer to achieve complete decoupling of ReRAM bitlines and storage nodes; a cascaded feedback bitline sensing architecture based on this sense amplifier, enabling single-bit precharge and multi-bit readout within a single row cycle; and a sense amplifier-centric cross-level interleaving mechanism for continuous column VMM access. This mechanism utilizes the open bitline topology between adjacent subarrays to improve underlying hardware resource utilization. The proposed subarray VMM sensing mechanism and memory-level parallel execution mechanism can significantly improve the overall module performance and execution efficiency of modern memory-constrained scientific parallel computing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of integrated storage and computing, and more specifically, relates to a bit line sense amplifier with bit line-storage node decoupling, an integrated storage and computing memory microarchitecture, and an in-situ matrix calculation method. Background Art

[0002] Resistive RAM (ReRAM) can perform vector-matrix multiplication (VMM) in-situ within the cell array and is expected to enable efficient iterative solutions to linear systems derived from mathematical physics equations, including electromagnetics, power networks, semiconductor devices, and circuit simulation. However, when multiple rows are activated simultaneously, the accumulation of bitline currents imposes significant area and latency overhead on VMM result buffering and sensing, which has become a system bottleneck. Previous work has used a sample-and-hold (S&H) to buffer the output results for each bitline and set up multiple analog-to-digital converters (ADCs) to obtain the results from the S&H for sensing, such as Figure 1 (a) in Figure 1 shows this. This design, where the sample-and-hold is separated from the analog-to-digital converter (ADC), increases the latency and energy overhead of bitline sensing. Furthermore, because previous work used a general-purpose ADC, the ADC occupies a significant portion of the bitline side spacing. Consequently, at least eight or more bitlines must share a single ADC for sensing, significantly limiting the parallelism and performance of column access.

[0003] In order to solve the problems of traditional separate design, researchers proposed Figure 1 Figure (b) shows a subarray microarchitecture without an ADC. In this architecture, bitline sense amplifiers (BLSAs) are provided at the top and bottom of the subarray. The storage nodes of the top and bottom bitline sense amplifiers are collectively called subarray row buffers, where even bit lines and odd bit lines are connected to the bottom bitline sense amplifier and the top bitline sense amplifier, respectively.

[0004] By activating multiple consecutive word lines at once, the sub-array operation efficiency can be improved. Accordingly, the number of bit line sense amplifiers connected to a bit line needs to be increased. The above-mentioned ADC-free sub-array micro-architecture can effectively alleviate the problems existing in the traditional micro-architecture with separate sample-and-hold and ADC. However, due to the limitations of the sensing structure and the corresponding sensing mechanism, it cannot efficiently support the operation of activating multiple consecutive word lines at once. Specifically, in the existing ADC-free sub-array micro-architecture, the structure of BLSA is as follows: Figure 2 For ease of display, only the sense amplifier (denoted as MSB-SA) corresponding to the most significant bit (MSB) is shown in the figure. Figure 2It can be seen that the existing BLSA has a structure similar to a latch, in which the bit line and the reference line are directly connected to the storage node SN and the reference line of the BLSA respectively. When the MSB-SA is enabled, it flips. However, the MSB-SA injects charge into the bitline and restores the bitline voltage to twice the reference line voltage or ground based on a comparison between the bitline voltage and the reference line voltage. Simultaneously, the reference line voltage is also restored to the opposite side by the MSB-SA. Therefore, this flip-based restoration process destroys both the established bitline voltage and the constant reference line voltage. This means that by the time the MSB-SA is stable enough to be read from its storage node, the CSB-SA and LSB-SA corresponding to the subsequent valid bits, CSB and LSB, cannot be sensed during the same voltage buildup period because their established bitline voltages have already been destroyed by the MSB-SA. This limits the VMM's "precharge once, read single bit" functionality. Consequently, the tRCD delays for the MSB, CSB, and LSB must be separated and serially completed over multiple row cycles, which underutilizes hardware and increases overall sensing latency. Summary of the Invention

[0005] In response to the shortcomings of the existing technology and the need for improvement, the present invention provides a bitline sense amplifier with bitline-storage node decoupling, a storage-computation-in-one memory microarchitecture, and an in-situ matrix calculation method. Its purpose is to propose a new bitline sense amplifier for open bitline ReRAM to improve the execution efficiency of vector-matrix multiplication in open bitline ReRAM.

[0006] To achieve the above objectives, according to one aspect of the present invention, a sense amplifier with bit line-storage node decoupling is provided, which is applied to sensing bit lines in an open bit line ReRAM subarray, and includes: a PMOS transistor P1, a latch, an equalizer, an unbalancer, and an NMOS transistor N1;

[0007] The drain of P1 is connected to the positive reading voltage V dd_RD , the gate of P1 acts as Signal input terminal;

[0008] The latch includes: PMOS transistors PU1 and PU2, and NMOS transistors PD1 and PD2; the drain of PU1 and the drain of PU2 are connected to form a node SAP, and the source of P1 is connected to the node SAP; the source of PU1 is connected to the drain of PD1 to form a first node; the gate of PU1 is connected to the gate of PD1 to form a second node; the source of PU2 is connected to the drain of PD2 to form a third node; the gate of PU2 is connected to the gate of PD2 to form a fourth node; the first node is connected to the third node as a storage node SN, and the second node is connected to the fourth node as a storage node Storage node SN and The stored signals are opposite to each other;

[0009] The equalizer includes an NMOS transistor EQZ and an NMOS transistor Bridge; the drain and source of EQZ are connected to the gates of PD1 and PD2 respectively, the gate of EQZ serves as an EQL signal input terminal; the gate of Bridge serves as a LOCKL signal input terminal;

[0010] The unbalancer includes: NMOS transistors N2 and N3; the drain of N2 is connected to the source of PD1, forming a node V P ; The drain of N3 is connected to the source of PD2, forming a node V Q The sources of N2 and N3 are connected to form a node SAN; the drain and source of Bridge are connected to the node V P and V Q The gate of N2 is used as the reference voltage input terminal, and the gate of N3 is used as the bit line voltage input terminal;

[0011] The drain of N1 is connected to the node SAN, the source of N1 is grounded, and the gate of N1 serves as the input terminal of the SAEN signal;

[0012] in, The signal is opposite to the SAEN signal. The SAEN signal is used to enable the sense amplifier. The EQL signal is the equalization line signal, and the LOCKL signal is the lock line signal.

[0013] According to another aspect of the present invention, a cascade feedback bit line sense amplifier is provided, comprising: sense amplifiers MSB-SA, CSB-SA, and LSB-SA, global reference voltage generators GRVG2, GRVG1, and GRVG0, a four-way decoder D0 and a two-way decoder D1, selectors S0, S1, and S2, and NMOS transistors SAPG0, SAPG1, and SAPG2;

[0014] The sense amplifiers MSB-SA, CSB-SA and LSB-SA are all the sense amplifiers with the bit line-storage node decoupling provided by the present invention;

[0015] The gates of SAPG0, SAPG1, and SAPG2 are connected to the CSL_LSB signal, CSL_CSB signal, and CSL_MSB signal, respectively. The sources of SAPG0, SAPG1, and SAPG2 are all connected to the local word line LDL. The drain of SAPG2 is connected to the storage node SN of MSB-SA, the drain of SAPG1 is connected to the storage node SN of CSB-SA, and the drain of SAPG0 is connected to the storage node SN of LSB-SA.

[0016] The first input terminals of S0, S1 and S2 are all connected to the SASL signal, the second input terminals of S0, S1 and S2 are all connected to the bit line, and the third input terminals of S0, S1 and S2 are all connected to the complement line; the output terminals of S0, S1 and S2 are respectively connected to the bit line voltage input terminals of LSB-SA, CSB-SA and MSB-SA;

[0017] GRVG2 is used to generate the reference voltage V ref3 , GRVG1 is used to generate the reference voltage V ref1 and V ref5 , GRVG1 is used to generate the reference voltage V ref0 、V ref2 、V ref4 and V ref6 ; V ref0 ~V ref6 Increase successively;

[0018] The output of GRVG2 is connected to the reference voltage input of MSB-SA;

[0019] The first input terminal of D1 is connected to the output terminal of GRVG1, the second input terminal of D1 is connected to the storage node SN of MSB-SA, and the output terminal of D1 is connected to the reference voltage input terminal of CSB-SA; D1 is used to change the voltage signal at the storage node SN of MSB-SA from V ref1 and V ref5 Select one input CSB-SA;

[0020] The first input terminal of D0 is connected to the output terminal of GRVG0, the second input terminal of D0 is connected to the storage node SN of CSB-SA, and the output terminal of D0 is connected to the reference voltage input terminal of LSB-SA; D0 is used to change the voltage signal at the storage node SN of CSB-SA from V ref0 、V ref2 、V ref4 and V ref6 Select one input LSB-SA;

[0021] Among them, the SASL signal is used to select the sense amplifier, the CSL_LSB signal is used to select the LSB-SA, the CSL_CSB signal is used to select the CSB-SA, and the CSL_MSB signal is used to select the MSB-SA.

[0022] Further, the selector includes: a PMOS transistor P4 and an NMOS transistor N4;

[0023] The gate of P4 is connected to the gate of N4, and the connection terminal serves as the first input terminal of the bit line sense amplifier;

[0024] The drain of P4 serves as the second input terminal of the bit line sense amplifier;

[0025] The source of N4 serves as the third input terminal of the bit line sense amplifier;

[0026] The source of P4 is connected to the drain of N4, and the connection terminal serves as the output terminal of the bit line sense amplifier.

[0027] According to another aspect of the present invention, a memory-computing-in-one microarchitecture based on open bitline ReRAM is provided. In a ReRAM subarray, each bitline is connected to a bitline sense amplifier using the cascade feedback bitline sense amplifier provided by the present invention. The bitline sense amplifiers connected to all bitlines constitute a row buffer of the subarray.

[0028] In a subarray, the bit line sense amplifiers connected to odd-numbered bit lines are located at the top of the subarray edge and are called top bit line sense amplifiers. The bit line sense amplifiers connected to even-numbered bit lines are located at the bottom of the subarray edge and are called bottom bit line sense amplifiers. In two adjacent subarrays, the bottom bit line sense amplifier of the previous subarray also serves as the top bit line sense amplifier of the next subarray.

[0029] According to another aspect of the present invention, an in-situ matrix calculation method for the above-mentioned storage-computation-in-one memory micro-architecture based on open bitline ReRAM is provided, comprising:

[0030] Step R1: Sending a row activation command ACTP to the target subarray to activate a word line with a non-zero input in an accessed page in the target subarray; the execution of the row activation command ACTP includes:

[0031] disabling the precharger in the target subarray to float the bit lines;

[0032] The word line voltage of the accessed page with non-zero input is increased from V BBW Rising to V PP , so that the sub-array row buffer senses the content of the row unit and takes it out to the storage node of the row buffer;

[0033] Step R2: Sending a precharge command PREP to the target sub-array to close all word lines in the accessed page in the target sub-array; the execution of the precharge command PREP includes:

[0034] The word line voltage in the accessed page is turned on from V PP Discharge to V BBW , so that the word line returns to a fully closed state; at the same time, the precharger is used to charge the bit line to V dd_RD , so that the bit line returns to the precharge state;

[0035] Step R3: Sending a column VMM command to the target subarray to gradually fetch the sensing results in the row buffer of the target subarray into a memory burst buffer and transmit the results to the memory burst buffer via a global I / O data line to complete data reading. The memory burst buffer includes a fixed number of I / O sense amplifiers, which are connected to bit line sense amplifiers in the I / O sense amplifiers via a shared local data line. The process of transmitting the sensing results from each bit line sense amplifier to the memory burst buffer includes:

[0036] Step R31: Select MSB-SA through the SASL signal, and drive the sensing data of MSB-SA to be transmitted to the corresponding I / O sense amplifier in the memory burst buffer through the local data line, and start to trigger the flip of the I / O sense amplifier, while releasing the local data line;

[0037] Step R32: After the sensing data transmission of MSB-SA is completed, CSB-SA is selected through the SASL signal, and the sensing data of CSB-SA is driven to be transmitted to the corresponding I / O sense amplifier in the memory burst buffer through the local data line, and the flip of the I / O sense amplifier is triggered, and the local data line is released at the same time;

[0038] Step R33: After the sensing data transmission of CSB-SA is completed, CSB-SA is selected through the SASL signal, and the sensing data of LSB-SA is driven to be transmitted through the local data line to the corresponding I / O sense amplifier in the memory burst buffer, and the flip of the I / O sense amplifier is triggered, and the local data line is released at the same time;

[0039] Among them, V BBW is the word line voltage in the fully off state, V PP is the word line voltage in the fully open state, V dd_RD Indicates a positive voltage reading.

[0040] Furthermore, the in-situ matrix calculation method provided by the present invention, that is, the matrix operands are stored in the array without being moved, in step R1, the word line voltage with non-zero input in the accessed page is changed from V BBW Rising to V PP During the read operation, when a read voltage margin of each MSB-SA in the row buffer of the target sub-array is formed, the corresponding word line is turned off.

[0041] Furthermore, the in-situ matrix calculation method provided by the present invention further includes, after step R33, if a subarray adjacent to the target subarray receives a column VMM command, selecting MSB-SA through a SASL signal after the sensing data transmission of LSB-SA is completed to sense the result in MSB-SA.

[0042] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects:

[0043] (1) The bit line-storage node decoupling sense amplifier provided by the present invention is mainly composed of an unbalancer, an equalizer and a latch, wherein the storage node SN and They are located on both sides of the latch, and the bit line reference voltage is input through the gates of the two NMOS units in the unbalancer. The latch and the unbalancer are isolated by an equalizer. Based on this structure, the content stored in the latch will not be disturbed by further attenuation of the bit line voltage. Therefore, there is no need to precharge again, and the bit line voltage can be used for sensing lower significant bits, realizing the function of "precharging once and reading out multiple significant bits of the result". This allows the tRCD delays of multiple significant bits to overlap, effectively shortening the overall sensing delay, improving sensing efficiency, and further improving the execution performance of vector-matrix multiplication in open bit line ReRAM.

[0044] (2) The cascade feedback bit line sensing amplifier provided by the present invention realizes a cascade feedback bit line sensing architecture by cascading three bit line-storage node decoupled sensing amplifiers. The three SAs in the sensing architecture can sense the results of three bits, and the three SAs can be enabled in the same row cycle, and the busy periods of the three SAs overlap, thereby realizing parallelism of the sensing amplifier level, effectively saving overall delay, and improving sensing efficiency.

[0045] (3) The present invention provides a memory-computing integrated memory micro-architecture based on open bit line ReRAM. In the sub-array of ReRAM, the bit line sense amplifier connected to each bit line is a cascade feedback bit line sense amplifier provided by the present invention. Through such an architecture, each bit line is equipped with a three-bit bit line sense amplifier and a total of seven reference voltages, which can activate 2 at a time. 3 = 8 adjacent word lines, which effectively improves the execution efficiency of subsequent operations compared to the traditional open bit line ReRAM near memory processing paradigm that only activates one word line at a time.

[0046] (4) The in-situ matrix calculation method provided by the present invention is based on the storage-computation-in-one memory micro-architecture based on open bit line ReRAM provided by the present invention. When performing vector-matrix multiplication calculation, all word lines with non-zero inputs in the accessed page are activated at one time, and then all opened word lines are closed at one time, and the bit lines are precharged. Since 8 adjacent word lines can be activated at one time in a sub-array of the storage-computation-in-one memory based on open bit line ReRAM provided by the present invention, when the result of the vector-matrix multiplication calculation is burst-read out from the row buffer, the three-bit sensing results in the same bit line sensing amplifier can be continuously transmitted without recharging, thereby realizing a sensing mechanism with effective bit interleaving, saving data lines and improving data bus efficiency while maintaining the same total bandwidth.

[0047] (5) In the preferred embodiment of the in-situ matrix calculation method provided by the present invention, when activating a word line, the word line is not turned off until the word line is charged to a fully open voltage. Instead, the corresponding word line is turned off when the read voltage margin of the MSB-SA is formed, thereby realizing an early word line closing mechanism. Compared with a fully open word line, under the early word line closing mechanism, the word line has a lower voltage when it is turned off, which makes the bit line have a slower decay rate during the word line closing period, thereby allowing subsequent CSB-SA and LSB-SA to use a reference voltage with a similar read margin and share the same voltage decay process in a single activation cycle, thereby reducing the energy consumption of the sub-word line driver used to drive the sub-word line and the power consumption of the local word line and the corresponding access transistor.

[0048] (6) In the preferred embodiment of the in-situ matrix calculation method provided by the present invention, based on the characteristic of sharing some bit line sense amplifiers between adjacent sub-arrays, after the sensing data transmission of LSB-SA in the previous array is completed, the MSB-SA is selected through the SASL signal to sense the result in MSB-SA again, so that the calculation result in the adjacent sub-array can be read out, thereby realizing the page interleaving execution mechanism between sub-arrays, that is, the activation of two pages falling in adjacent sub-arrays is interleaved, and the row cycles of two page accesses in two sub-arrays falling in the same group are overlapped, further improving the data transmission rate and improving the overall execution efficiency of the in-situ matrix calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 Schematic diagram of the subarray microarchitecture in a conventional open bitline ReRAM; (a) shows a subarray microarchitecture with a separate sample-and-hold and ADC, and (b) shows a subarray microarchitecture without an ADC.

[0050] Figure 2 A schematic diagram of a bit line sense amplifier in a conventional sub-array micro-architecture without an ADC;

[0051] Figure 3 A schematic diagram of an existing memory-computing-in-one microarchitecture based on open bitline ReRAM;

[0052] Figure 4 A schematic diagram of a hexagonal layout of an existing open bit line ReRAM;

[0053] Figure 5 Schematic diagram of a bit line-storage node decoupled sense amplifier and a cascade feedback bit line sense amplifier provided by an embodiment of the present invention;

[0054] Figure 6 A schematic diagram illustrating the principle of implementing data sensing by a sense amplifier with bit line-storage node decoupling provided by an embodiment of the present invention;

[0055] Figure 7 A timing diagram of different operations provided by an embodiment of the present invention;

[0056] Figure 8 This is a timing diagram of page access in two sub-arrays under the interleaving execution mechanism of pages between sub-arrays provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0057] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0058] In the present invention, the terms "first", "second", etc. (if any) in the present invention and the drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0059] Before explaining the technical solution of the present invention in detail, the architecture and related operation mechanism of the open bit line ReRAM are first explained.

[0060] The term ReRAM, as used in this disclosure, refers to any memory technology that uses nonvolatile resistance to store information. It is a microarchitectural-level abstraction for devices and circuits. This disclosure uses a subcategory of ReRAM, called metal oxide ReRAM, as an example. This subcategory has a metal-insulator-metal (MIM) device structure, but is not limited to metal oxide ReRAM and also includes PCRAM, MRAM, and other technologies.

[0061] like Figure 3As shown in the figure, ReRAM consists of multiple memory ranks. A rank is a group of chips working simultaneously to serve a single command issued by the controller. Each chip consists of memory banks that can be accessed in parallel and independently. Each memory bank is divided into multiple blocks to reduce the RC delay of long lines. Each block in the memory bank is called a tile. A tile includes peripheral core circuits and cell arrays. The structure of the cell array is as follows: Figure 4 As shown, each cell stores a "1" in a high-conductance state and a "0" in a low-conductance state. The peripheral core includes a sub-wordline driver (SWD), a bitline sense amplifier (BLSA), a precharger, and column access circuitry. A row of tiles within a memory bank is called a subarray, and a row within a subarray refers to a row of cells. The sub-wordline driver can span two rows through the interleaving of local wordlines. Only one subarray within a memory bank can be accessed at a time. Each bitline in a subarray is connected to a bitline sense amplifier. The bitline sense amplifiers connected to the same subarray are distributed at the top and bottom of the subarray, with even and odd bitlines connected to the bottom and top bitline sense amplifiers, respectively. The bitline sense amplifiers connected to all bitlines in a subarray together constitute the subarray row buffer (SRB). Multiple rows (typically 8 in size) within a subarray that are simultaneously activated to execute the VMM are called a page.

[0062] Corresponding to the DDR5 memory bus technology, each sub-array consists of 1024 rows of normal cells and 8192 normal bitlines, as well as dummy global wordlines and corresponding cells on both sides (the number of dummy global wordlines is usually 8 or 16) to tolerate the overdriven insulated gate PMOS transistor in a certain normal sub-wordline driver at V PP A hard error occurs under bias (the voltage when the word line is fully turned on), causing an entire row of cells to fail and be replaced. In order to save the area of the decoder, the sub-word line driver adopts a two-stage decoding method to parse the row address. Specifically, for the sub-word line driver, the sub-array row address is hierarchically decoded into 128 global word lines (7 bits). ) and 8 pre-decoded row address lines (3-bit ) to trigger SWD to drive the sub-word line.

[0063] The SWD is essentially an AND gate with inverted inputs. Within a subarray, a column select line (CSL) is shared between the SA storage nodes of 128 consecutive bit lines. A column consists of 128 consecutive BLSAs, meaning the column size, or word size, is 128. The storage nodes of these BLSAs are selected simultaneously by a single CSL to transmit data to the 128 I / O sense amplifiers (IOSAs) in the memory bank. These 128 IOSAs form the subarray's bank burst buffer (BBB). On both the rising and falling edges of the clock, the BBB outputs data in burst mode on eight global I / O data lines (GDLs), with a burst length (BL) of 16.

[0064] Figure 4 The hexagonal layout of the open bitline ReRAM is shown, where cylindrical ReRAM is stacked on top of the storage element contacts on the bitline, with the top metal electrode connected to the common ground plane. The open bitline architecture is characterized by an interdigital layout of odd and even bitlines connected to the top and bottom bitline sense amplifiers at the edge of the subarray, respectively, and each BLSA is connected to two bitlines from two adjacent subarrays. For the open bitline organization, there are two additional redundant dummy subarrays at the top and bottom edges of the memory bank, which serve as reference bitlines. To match the normal bit line parasitic capacitance in the 0th and (N-1)th sub-arrays, N represents the number of sub-arrays. The open bit line architecture can achieve 6F 2 The cell area is 8F, where F is the feature size. The folded cell array architecture has an area to accommodate bit line twisting and twice the word line length, with a minimum of 8F. 2 Therefore, at the same feature size, the open bit line architecture is 25% denser than the folded bit line architecture.

[0065] In open bitline ReRAM, for bipolar write access, the bitline is biased at a negative SET voltage and a positive RESET voltage, and 1 and 0 are written respectively by enabling the write driver. For read access, the bitline is first precharged to a positive read voltage (V dd_RD ) to prepare for voltage development by floating the bit line.

[0066] To ensure analog signal integrity, the timing parameters of the memory row access and column access commands on the controller side are as follows:

[0067] 1) Row Activation (ACT) on a subarray. The ACT command accompanied by the row address first disables the precharger to float the bit line and pulls the word line from V BBW Open to V PP, and as the bit line parasitic capacitance discharges to ground through the open cell, the precharged bit line voltage begins to drop, where V PP is the voltage when the word line is fully open, V BBW The voltage when the word line is completely closed corresponds to the DDR5 bus transmission standard. Typically, V PP =2.5V, V BBW = -0.2V. The sub-array row buffer (BLSA) can then sense the contents of the entire row of cells and fetch them into the row buffer storage nodes. The delay from the issuance of the ACT command to the stabilization of the sub-array row buffer is called Delay time (tRCD).

[0068] 2) Row Precharge (PRE) on the sub-array. The PRE command with row address starts at V PP The word line that is already turned on in the state is turned off back to the fully off state (V BBW The PRE command also enables the precharger to bias the bit line back to V dd_RD The PRE command is a sub-array initialization command, and the delay of the whole process is called Precharge time (tRP). The time interval from the issuance of the ACT command to the issuance of the PRE command is called Time (tRAS), tRAS plus tRP are together called the row cycle time (tRC).

[0069] 3) Column read access (RD) on the row buffer storage node. When the row buffer storage node and the bit line are sufficiently stable during activation, the RD command issues a column address to enable a specific column select line (CSL) and transfers a strip of data (that is, reading "SRAM", typically 128 bits) from the sub-array row buffer storage node through the local data line (LDL) and writes to the bank burst buffer (BBB) composed of 128 I / O sense amplifiers (IOSA). Then, at the rising and falling edges of the clock, the BBB outputs data with 8 global I / O data lines (GDL) in burst mode with a burst length (BL) of 16. Therefore, the data rate of the external GDL is 32 times the data rate of the internal core clock frequency. The delay from issuing the RD command to the BBB being stable and outputting the first data bit to the GDL is called the delay of the read operation. Delay time (tAA), the burst delay of GDL is calculated as BL / 2*tCK, where tCK is the clock period of GDL. Note that tAA is larger than the busy cycle time of the sub-array row buffer because the bit line read only occupies one stage in the memory bank pipeline. The delay is CL, the write operation The delay is CWL, CL is longer than CWL. The delay between two adjacent column commands to the same bank is called Delay time (tCCD_L).

[0070] For the vector matrix multiplication operation, each bit line performs a column-by-column dot product calculation. The bit line is first precharged to a positive read voltage V by the PRE command. dd_RD , similar to the read-bitline in a dual-port single-ended 3T DRAM. When the wordline in a page is turned on, the charge in the parasitic capacitance of the bitline leaks through the turned-on cell. Using binary inputs to the wordline gate, the decay of the bitline voltage in a given time By dot product ∑ i V i G ij Determine, where if V wordline_i =V PP , then V i =V dd_RD ; If V wordline_i =V BBW , then V i = 0. Here, G ij is the cell conductance at the i-th word line and j-th bit line, V wordline_i Represents the input voltage of the word line in row i, V i represents the bias voltage value for the cells in row i. Different calculation results generate different bitline voltages. If the bitline dot product current is larger, the bitline voltage decays faster (i.e., the bitline parasitic capacitance discharges), resulting in a lower post-decline bitline voltage.

[0071] Based on open bit line ReRAM, it consists of 7 bits and 3 digits The present invention finds that if a two-stage decoding mechanism consisting of a The corresponding eight adjacent sub-word lines can effectively improve the computational efficiency of the tiles in the sub-array. To this end, each bit line needs to be equipped with a three-bit BLSA and a total of seven reference voltages. Based on binary search hierarchical sensing, the most significant bit (MSB) is first sensed given the third reference potential. Then, the center significant bit (CSB) is sensed given the first and fifth reference levels. Finally, the least significant bit (LSB) is read out given the 0th, 2nd, 4th, and 6th reference levels.

[0072] However, in Figure 2 In the conventional BLSA structure shown in FIG, the bit line and the reference line are directly connected to the storage node SN and the storage node SN of the BLSA respectively. The aforementioned scheme of activating eight word lines at once results in a long sensing delay. Specifically, when the MSB-SA is enabled, it flips, injecting charge into the bitline and restoring the bitline voltage to twice the reference line voltage or ground based on a comparison between the bitline voltage and the reference line voltage. Simultaneously, the reference line voltage is also restored to the opposite side by the MSB-SA. Therefore, this flip-based recovery process destroys both the established bitline voltage and the constant reference line voltage. This means that when the MSB-SA is sufficiently stable and ready for reading from its storage node, the subsequent CSB-SA and LSB-SA cannot be sensed during the same voltage buildup period because the established bitline voltage has already been destroyed by the MSB-SA. This limits the VMM's "precharge once, read single bit" functionality. In this case, after reading the MSB-SA result, a precharge operation should be performed to initialize the bitline voltage before issuing the next activate command to reset the bitline voltage and perform CSB-SA sensing. As a result, the tRCD delays for the MSB, CSB, and LSB must be separated in a serial manner over multiple line cycles, which underutilizes the hardware and increases the overall delay.

[0073] In addition, for traditional designs where the bit lines are directly connected to the SA storage nodes, the sub-array row buffers should be disabled when the bit line precharge begins. In traditional bit line-storage node connection designs, there is a column conflict between bit line precharge and VMM column access. Since the precharger and the SA storage node are both directly connected to the bit lines and they drive the bit lines to different target voltages, the precharger and the SA storage node are mutually exclusive and cannot work at the same time. During bit line precharge, the SA storage node voltage changes with the bit line and becomes invalid. Precharging the bit line will destroy the data in the SA storage node, so the CSL cannot read the data from the SA storage node to the local data line by issuing a VMM column access. In this case, the bit line precharge command should be issued after the CSL signal of the LSB falls, and the delay overhead is significant.

[0074] Overall, efficient and low-cost subarray sensing is a key hardware challenge for high-performance in-situ vector-matrix multiplication computations performed on ReRAM cell arrays. To address this issue, this paper first proposes a bitline-to-storage node decoupled sense amplifier to completely isolate the bitline voltage buildup process from SA flips, while simultaneously implementing two-in-one buffering and sensing, enabling conflict-free column access. Furthermore, a cascaded feedback bitline sensing architecture is proposed, in which each bitline sense amplifier (BLSA) is composed of a cascade of the aforementioned bitline-to-storage node decoupled sense amplifiers. This architecture enables precharge-once, read-out-of-multiple-bits (PORM) functionality within a single row cycle, overlapping the tRCD delays of different valid result bits sensed on the same bitline. Furthermore, a sense amplifier-centric interleaving mechanism is proposed for consecutive column VMM accesses, improving underlying hardware resource utilization by leveraging the open bitline topology between adjacent subarrays. The proposed subarray VMM sensing mechanism and memory-level parallelization technology can significantly improve the overall module performance and execution efficiency of modern memory-constrained scientific parallel computing.

[0075] The following are examples.

[0076] Example 1:

[0077] A sense amplifier with bit line-storage node decoupling, such as Figure 5 As shown, it is applied to sensing of a bit line in an open bit line ReRAM subarray, including: a PMOS transistor P1, a latch, an equalizer, an unbalancer and an NMOS transistor N1;

[0078] The drain of P1 is connected to the positive reading voltage V dd_RD , the gate of P1 acts as Signal input terminal;

[0079] The latch includes: PMOS transistors PU1 and PU2, and NMOS transistors PD1 and PD2; the drain of PU1 and the drain of PU2 are connected to form a node SAP, and the source of P1 is connected to the node SAP; the source of PU1 is connected to the drain of PD1 to form a first node; the gate of PU1 is connected to the gate of PD1 to form a second node; the source of PU2 is connected to the drain of PD2 to form a third node; the gate of PU2 is connected to the gate of PD2 to form a fourth node; the first node is connected to the third node as a storage node SN, and the second node is connected to the fourth node as a storage node Storage node SN and The stored signals are opposite to each other;

[0080] The equalizer includes an NMOS transistor EQZ and an NMOS transistor Bridge; the drain and source of EQZ are connected to the gates of PD1 and PD2 respectively, the gate of EQZ serves as an EQL signal input terminal; the gate of Bridge serves as a LOCKL signal input terminal;

[0081] The unbalancer includes: NMOS transistors N2 and N3; the drain of N2 is connected to the source of PD1, forming a node V P ; The drain of N3 is connected to the source of PD2, forming a node V Q The sources of N2 and N3 are connected to form a node SAN; the drain and source of Bridge are connected to the node V P and V Q The gate of N2 is used as the reference voltage input terminal, and the gate of N3 is used as the bit line voltage input terminal;

[0082] The drain of N1 is connected to the node SAN, the source of N1 is grounded, and the gate of N1 serves as the input terminal of the SAEN signal;

[0083] in, The EQL signal is the inverse of the SAEN signal, which is used to enable the sense amplifier; the EQL signal is the equalization line signal, and the LOCKL signal is the lock line signal.

[0084] This embodiment is proposed to completely isolate the line voltage establishment process from the SA flip. The sense amplifier SA provided in this embodiment is VMM sensing, which is inspired by the Wheatstone bridge and mainly includes three parts: unbalancer, equalizer and latch. Its sensing process is as follows Figure 6 As shown, when both the EQL and LOCKL signals are disabled, a noise voltage (i.e., the difference between the bitline voltage and the reference voltage) is injected into the right leg of the latch through the unbalancer, triggering a latch flip. When the storage node of the SA reaches a sufficiently stable threshold state, the resulting bit is stably stored in the SA latch. At this point, the corresponding latch line (LOCKL) is enabled, decoupling the latch from the unbalancer. The SA's conformation changes, with the N2 and N3 transistors now connected in parallel through a balancing mechanism. Therefore, the contents stored in the SA latch are not disturbed by further attenuation of the bitline voltage, and the bitline voltage can be used for sensing by the less significant SA. This embodiment implements a folded differential layout by forming a differential pair input with N2 and N3, allowing the proposed SA to fit into an open bitline layout with twice the bitline pitch of a 6F. This ensures that every bitline in the subarray can be connected to the BLSA, making the bitlines fully parallelizable.

[0085] The decoupling function implemented by the SA provided in this embodiment has three aspects:

[0086] First, when the SA is on, the floating bit line is not biased by the storage node of the flipping SA, so the generated bit line voltage is not restored and destroyed by the SA. The SA of the less significant bit can compare the bit line voltage with its reference voltage for sensing in the same row cycle, and the tRCD delays of different significant bits sensed on the same bit line can overlap, enabling the charge-one-read-many (PORM) function.

[0087] Second, because the bitline voltage continues to decay over time during the bitline voltage buildup period, if the resulting bit to be sensed is zero, the bitline voltage may fall below the corresponding reference voltage during this period. Without a built-in equalizer, changes in the unbalanced gate inputs of the SA could reverse the positions of the opposing CMOS inverter arms in the latch and potentially cause unwanted flipping. By introducing a Wheatstone bridge as a built-in equalizer, the continuous voltage decay of the floating bitline does not affect the voltage of the SA's storage node when the bridge is enabled.

[0088] Third, the SA is completely decoupled from the precharger through the unbalanced gate input. The bitline precharge process does not affect the read process of the SA storage node. Therefore, the row buffer storage node can be enabled during the bitline precharge period. Now the bitline precharge process can be parallel with the VMM column access within the subarray in a conflict-free manner, such as Figure 7 shown.

[0089] This embodiment decouples the bit line from the SA storage node, thereby optimizing the bit line voltage buildup for result sensing and the SA flipped final state for storage node readout. Decoupling the bit line from the SA storage node reduces parasitic capacitance connected to the SA storage node, thereby increasing the SA storage node flipping speed.

[0090] Example 2:

[0091] A cascade feedback bit line sense amplifier, such as Figure 5 As shown, it includes: sense amplifiers MSB-SA, CSB-SA and LSB-SA, global reference voltage generators GRVG2, GRVG1 and GRVG0, four-way decoder D0 and two-way decoder D1, selectors S0, S1 and S2, and NMOS transistors SAPG0, SAPG1 and SAPG2;

[0092] The sense amplifiers MSB-SA, CSB-SA and LSB-SA are all the sense amplifiers with the bit line-storage node decoupling provided by the present invention;

[0093] The gates of SAPG0, SAPG1, and SAPG2 are connected to the CSL_LSB signal, CSL_CSB signal, and CSL_MSB signal, respectively. The sources of SAPG0, SAPG1, and SAPG2 are all connected to the local word line LDL. The drain of SAPG2 is connected to the storage node SN of MSB-SA, the drain of SAPG1 is connected to the storage node SN of CSB-SA, and the drain of SAPG0 is connected to the storage node SN of LSB-SA.

[0094] The first input terminals of S0, S1 and S2 are all connected to the SASL signal, the second input terminals of S0, S1 and S2 are all connected to the bit line, and the third input terminals of S0, S1 and S2 are all connected to the complement line; the output terminals of S0, S1 and S2 are respectively connected to the bit line voltage input terminals of LSB-SA, CSB-SA and MSB-SA;

[0095] GRVG2 is used to generate the reference voltage V ref3 , GRVG1 is used to generate the reference voltage V ref1 and V ref5 , GRVG1 is used to generate the reference voltage V ref0 、V ref2 、V ref4 and V ref6 ; V ref0 ~V ref6 Increase successively, such as Figure 7 As shown;

[0096] The output terminal of GRVG2 is connected to the reference voltage input terminal of MSB-SA; the connection line between the output terminal of GRVG2 and the reference voltage input terminal of MSB-SA serves as the global reference line (GRL);

[0097] The first input terminal of D1 is connected to the output terminal of GRVG1, the second input terminal of D1 is connected to the storage node SN of MSB-SA, and the output terminal of D1 is connected to the reference voltage input terminal of CSB-SA; D1 is used to change the voltage signal at the storage node SN of MSB-SA from V ref1 and V ref5 Select one input CSB-SA;

[0098] The first input terminal of D0 is connected to the output terminal of GRVG0, the second input terminal of D0 is connected to the storage node SN of CSB-SA, and the output terminal of D0 is connected to the reference voltage input terminal of LSB-SA; D0 is used to change the voltage signal at the storage node SN of CSB-SA from V ref0 、V ref2 、V ref4 and V ref6 Select one input LSB-SA;

[0099] Among them, the SASL signal is used to select the sense amplifier, the CSL_LSB signal is used to select the LSB-SA, the CSL_CSB signal is used to select the CSB-SA, and the CSL_MSB signal is used to select the MSB-SA.

[0100] In this embodiment, the selector is used to select the bit line or the complementary bit line voltage as the input of SA under the control of SASL; optionally, Figure 5 As shown, the selector includes: a PMOS transistor P4 and an NMOS transistor N4;

[0101] The gate of P4 is connected to the gate of N4, and the connection terminal serves as the first input terminal of the bit line sense amplifier;

[0102] The drain of P4 serves as the second input terminal of the bit line sense amplifier;

[0103] The source of N4 serves as the third input terminal of the bit line sense amplifier;

[0104] The source of P4 is connected to the drain of N4, and the connection terminal serves as the output terminal of the bit line sense amplifier.

[0105] The cascade feedback bit line sense amplifier provided in this embodiment is proposed to construct a bit-accurate row buffer and store the sensed result bit directly in the SA, which includes three cascaded sense amplifiers, and each sense amplifier is a sense amplifier with bit line-storage node decoupling provided in the above-mentioned embodiment 1. Each SA latches a result, and the latch output of each SA is fed back to select the reference voltage input of the lower-level SA. Since the SA realizes the decoupling of the bit line and the storage node, the MSB-SA, CSB-SA and LSB-SA can be enabled successively in the same BLSA. By constructing a feedback link structure, the precise reference value of the subsequent lower-significant bit SA can be determined, so that bit-by-bit buffering can be achieved by using cascade feedback in the SA structure. The VMM column access for obtaining the MSB-SA sensing result can be parallel with the CSB-SA sensing, and the busy periods of the MSB-SA, CSB-SA and LSB-SA overlap, thereby saving the overall delay. Figure 7 As shown, parallelism of the sense amplifier stage is achieved. In the case of a traditional architecture without cascade feedback, the MSB result bit can only be output after all bits are sensed. In contrast, this embodiment can effectively reduce the overall delay and improve the sensing efficiency.

[0106] Example 3:

[0107] A memory-computing-in-one microarchitecture based on open bitline ReRAM, such as Figure 5As shown, in the sub-array of the ReRAM, the bit line sense amplifier connected to each bit line is the cascade feedback bit line sense amplifier provided in the above embodiment 2, and the bit line sense amplifiers connected to all the bit lines constitute a row buffer of the sub-array;

[0108] Based on the architectural characteristics of open bit line ReRAM, in a subarray, the bit line sense amplifiers connected to the odd-numbered bit lines are located at the top edge of the subarray, called top bit line sense amplifiers, while the bit line sense amplifiers connected to the even-numbered bit lines are located at the bottom edge of the subarray, called bottom bit line sense amplifiers. In two adjacent subarrays, the bottom bit line sense amplifier of the previous subarray also serves as the top bit line sense amplifier of the next subarray.

[0109] Based on the cascade feedback bit line sense amplifier provided in the above embodiment 2, each bit line is equipped with a three-bit bit line sense amplifier and a total of seven reference voltages, and the sensing results in the three SAs can be read out in sequence after one precharge, so that two bit lines can be activated at a time. 3 =8 adjacent word lines, compared to traditional open bit line ReRAM that activates one word line at a time, effectively improving the execution efficiency of subsequent operations.

[0110] Example 4:

[0111] The above-mentioned in-situ matrix calculation method for the memory-computation-in-one memory micro-architecture based on open bitline ReRAM is proposed to improve the bitline VMM sensing efficiency, and mainly includes a PROM mechanism for VMM access and a cross-level interleaved execution scheme.

[0112] Based on a cascade feedback bitline sensing architecture, with binary voltages used as wordline inputs to the gates of access transistors and a one-bit-per-cell memory configuration, VMM operations are effectively read accesses. This embodiment modifies the row ACT and PRE commands and introduces VMM as a column command to support bank VMM access.

[0113] This embodiment specifically includes:

[0114] Step R1: Send a row activation command ACTP to the target subarray to activate the word lines with non-zero input in the accessed page in the target subarray; in this embodiment, the ACTP command is similar to the native ACT command, except that the word lines with non-zero input in the accessed page are turned on to V PP Specifically, the execution of the row activation command ACTP includes:

[0115] disabling the precharger in the target subarray to float the bit lines;

[0116] The word line voltage of the accessed page with non-zero input is increased from V BBW Rising to VPP , so that the sub-array row buffer senses the content of the row unit and takes it out to the storage node of the row buffer;

[0117] Step R2: Send a precharge command PREP to the target subarray to close all word lines in the accessed page in the target subarray. In this embodiment, the precharge command PREP is similar to the native command PRE, except that all opened word lines in the accessed page are closed to V BBW Specifically, the execution of the precharge command PREP includes:

[0118] The word line voltage in the accessed page is turned on from V PP Discharge to V BBW , so that the word line returns to a fully closed state; at the same time, the precharger is used to charge the bit line to V dd_RD , so that the bit line returns to the precharge state;

[0119] Step R3: Send a column VMM command to the target subarray to gradually fetch the sensing results in the row buffer of the target subarray into the memory burst buffer and transmit them outward through the global I / O data line to complete the data reading. The memory burst buffer includes a fixed number of I / O sense amplifiers, which are connected to the bit line sense amplifiers in the I / O sense amplifiers through shared local data lines. In this embodiment, the column VMM command is similar to the native column RD command, except that the column VMM command fetches the sensing results of 128 bit lines from the subarray row buffer storage node to the BBB and transmits them out in a burst for global shift accumulation. During the column VMM access operation, the SA storage node must be enabled. This embodiment introduces three versions of VMM commands, namely VMMM, VMMC, and VMML, for obtaining the sensing results of the MSB-SA, CSB-SA, and LSB-SA, respectively. In a designed memory bank, there are a total of 64 CSL groups, one of which includes CSL_M, CSL_C, and CSL_L. Specifically, the process of transmitting the sensing results from each bit line sense amplifier to the memory bank burst buffer includes:

[0120] Step R31: Select MSB-SA through the SASL signal, drive the sensed data of MSB-SA to be transmitted to the corresponding I / O sense amplifier in the memory bank burst buffer through the local data line, and start triggering the flip of the I / O sense amplifier, while releasing the local data line; Step R31 corresponds to the execution of the VMMM command;

[0121] Step R32: After the MSB-SA sensing data transmission is completed, CSB-SA is selected through the SASL signal, and the sensing data of CSB-SA is driven to be transmitted through the local data line to the corresponding I / O sense amplifier in the memory bank burst buffer, and the flip of the I / O sense amplifier is triggered, and the local data line is released at the same time; Step R32 corresponds to the execution of the VMMC command;

[0122] Step R33: After the sensing data transmission of CSB-SA is completed, CSB-SA is selected through the SASL signal, and the sensing data of LSB-SA is driven to be transmitted through the local data line to the corresponding I / O sense amplifier in the memory burst buffer, and the flip of the I / O sense amplifier is triggered, and the local data line is released at the same time; Step R33 corresponds to the execution of the VMML command;

[0123] Among them, V BBW is the word line voltage in the fully off state, V PP is the word line voltage in the fully open state, V dd_RD Indicates a positive voltage reading.

[0124] SA plus SA pass gate transistor (SAPG) is actually an SRAM cell. In this embodiment, it is observed that the column select line CSL is actually the word line of SA, and the local data line LDL is actually the bit line of SA. LDL is first precharged to V dd_RD . When CSL is enabled, in order to prevent read interference and maintain the read static noise margin (RSNM) of the SA storage node, the SA storage node should be more decoupled from the local data line LDL, so the resistance of the SAPG transistor should be much larger than the resistance of the SA pull-down (PD) transistor. To ensure this, in this embodiment, the width-to-length ratio of the SAPG transistor is set to four times the width-to-length ratio of the pull-down transistor in the SA, that is, the β ratio of the "SRAM cell" is four. The large channel resistance of the SAPG transistor causes a large delay. During the "SRAM read access" driven by the CSL signal, the turn-on / off delay of the small SAPG is 3.75ns, which suppresses the column cycle time tCCD_L and tAA of the VMM column access. Although the SA access circuit can be designed as in a dual-port 8T SRAM cell without read interference, this will increase the layout area.

[0125] In order to hide the latency of the frequently enabled SAPG and improve the data bus efficiency, in this embodiment, when executing a column VMM command to burst-read the VMM result bits from the BLSA, the result bits sensed in the MSB-SA, CSB-SA, and LSB-SA are transferred from the row buffer to the memory bank burst buffer in an interleaved manner, thereby implementing a valid bit interleaving scheme. Based on this valid bit interleaving scheme, VMM column accesses of different valid bits are partially parallelized, and the peripheral column access latency overlaps, as shown in FIG. Figure 7 shown.

[0126] like Figure 5 As shown, a bit-interleaved column access connector (BIC) is designed in the BLSA. The MSB is first transmitted outward via the local data line LDL to the corresponding MSB IOSA in the bank burst buffer, triggering the toggling of the single-ended IOSA. At this point, the shared local data line is released by disconnecting the IOSA, preparing the CSB-SA for transmission outward to the corresponding IOSA. Finally, after the LSB is transferred to the bank burst buffer, reading out the MSB SA result can begin again. The result bits stored in the MSB-SA, CSB-SA, and LSB-SA stabilize sequentially, allowing them to be bursted using a single local data line. Data transmission from different active SAs to the bank burst buffer is independently controlled by different CSLs. The data rate of the local data line LDL is three times the CSL activation frequency. For column accesses, this burst of data from different active SAs saves data lines and improves data bus efficiency compared to a design that allocates a data line for each SA, while maintaining the same total bandwidth.

[0127] Open bitline ReRAM has odd-even staggered bitlines without bitline isolation transistors. In the traditional SA design where the bitlines are directly connected to the storage nodes, when the jth SA is enabled, the (2j+1)th complement bitline in the (2k+1)th subarray The bit line is used as the voltage reference line for the 2jth bit line in the 2kth subarray, and both the bit line voltage and the corresponding complement bit line voltage are restored by the SA. The complement bit line cannot then be sensed again by the same SA in the same row cycle. This dependency prevents the SA from switching to sensing the complement bit line's MSB result once the corresponding bit line's LSB result has been sensed.

[0128] Taking into account the above characteristics of the open bit line ReRAM architecture, in order to further improve the efficiency of the data bus, this embodiment further proposes a page interleaving scheme between subarrays. Specifically, a pair of bit lines and complement lines are decoupled by introducing a global reference voltage generator (GRVG) and a global reference line (GRL) between two adjacent subarrays. The global reference line GRL is input to the gate of the unbalancer of the SA to cooperate with the proposed design of decoupling the bit line from the SA storage node; in this embodiment, the local data line (LDL) is shared between two adjacent subarrays. Through the decoupling mechanism, the GRL voltage will not be destroyed by the flip of the SA. In this way, two dummy edge subarrays are not required to construct the complement line for reference. After the LSB of the bit line result is sensed and the LSB SA becomes idle, SASL is switched to connect the complement line to the MSB SA to sense its MSB result. In this way, the activation of the two pages falling in adjacent subarrays is interleaved, and the row cycles of the two page accesses in the two subarrays falling in the same group are overlapped, for example, Figure 8 In the example, the row periods of subarray [2k] and its adjacent subarray [2k+1] overlap. The inter-subarray page interleaving scheme proposed in this embodiment enables the bit lines and bit complement lines to share the same BLSA and enable burst transmission, further doubling the data rate of the local data line LDL.

[0129] In the traditional design, the accessed word line starts to shut down immediately when the PRE command is issued to the sub-array, and the accessed word line is completely closed to V BBW After that, bit line precharge begins. For ReRAM bit line sensing, the decay rate of the bit line voltage during the word line off period is slower than when the word line is kept at V PP This is because the word line is directly connected to the gate of a row of access transistors. The transistor channel resistance increases as the word line is turned off, and the charge leakage path through the bit line of the ReRAM cell being turned off is gradually cut off, so the rate of bit line voltage decay slows down as the word line is turned off. Finally, when the word line is completely turned off to V BBW When the word line is turned off, the bit line voltage is clamped to a constant value. This word line shutdown process prevents the bit line voltage from falling further, preventing the maximum voltage drop (for the "111" sensing result) from reaching the GND potential. If the result potential of "110" also reaches GND, the two result states are indistinguishable.

[0130] In order to reduce power consumption, this embodiment proposes an Early Sub-Wordline Closing (ESWC) mechanism for the execution of the row activation command ACTPD. Specifically, in step R1, the voltage of the sub-wordline with non-zero input in the accessed page is increased from V BBW Rising to V PPDuring the process, when the read voltage margin of each MSB-SA in the row buffer of the target sub-array is formed, the corresponding word line is turned off. PP When the MSB-SA read voltage margin is established, the opened sub-wordline is closed without waiting for a subsequent precharge command. Taking the MSB SA as an example, the early sub-wordline closing mechanism facilitates subsequent CSB and LSB sensing. This bitline voltage clamping based on sub-wordline closing allows the subsequent CSB-SA and LSB-SA to use a reference voltage with similar read margins, sharing the same voltage decay process in a single activation cycle. Furthermore, when the MSB-SA read voltage margin is established, meaning the SA latch is about to flip, the sub-wordline driver closes the wordline, thereby reducing the power consumption of the sub-wordline driver as well as the power consumption of the sub-wordline and the corresponding access transistor.

[0131] In general, this embodiment can implement the PORM function within a single row cycle, improve the efficiency of vector-matrix multiplication calculations, and enhance hardware resource utilization.

[0132] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A sense amplifier with bit line-storage node decoupling, applied to sensing bit lines in an open bit line ReRAM subarray, characterized in that: include: PMOS transistor P1, latch, equalizer, unbalancer and NMOS transistor N1; The drain of P1 is connected to the positive reading voltage V dd_RD , the gate of P1 acts as Signal input terminal; The latch includes: PMOS transistors PU1 and PU2, and NMOS transistors PD1 and PD2; the drain of PU1 and the drain of PU2 are connected to form a node SAP, and the source of P1 is connected to the node SAP; the source of PU1 and the drain of PD1 are connected to form a first node; the gate of PU1 and the gate of PD1 are connected to form a second node; the source of PU2 and the drain of PD2 are connected to form a third node; the gate of PU2 and the gate of PD2 are connected to form a fourth node; the first node and the third node are connected to serve as a storage node SN, and the second node and the fourth node are connected to serve as a storage node ; Storage node SN and The stored signals are opposite to each other; The equalizer includes: an NMOS transistor EQZ and an NMOS transistor Bridge; the drain and source of EQZ are connected to the gates of PD1 and PD2 respectively, the gate of EQZ serves as an EQL signal input terminal; the gate of Bridge serves as a LOCKL signal input terminal; The unbalancer includes: NMOS transistors N2 and N3; the drain of N2 is connected to the source of PD1 to form a node V P ; The drain of N3 is connected to the source of PD2, forming a node V Q The sources of N2 and N3 are connected to form a node SAN; the drain and source of Bridge are connected to the node V P and V Q The gate of N2 is used as the reference voltage input terminal, and the gate of N3 is used as the bit line voltage input terminal; The drain of N1 is connected to the node SAN, the source of N1 is grounded, and the gate of N1 serves as the input terminal of the SAEN signal; in, The signal is opposite to the SAEN signal, which is used to enable the sense amplifier; the EQL signal is the equalization line signal, and the LOCKL signal is the lock line signal.

2. A cascade feedback bit line sense amplifier, characterized in that: include: Sense amplifiers MSB-SA, CSB-SA, and LSB-SA, global reference voltage generators GRVG2, GRVG1, and GRVG0, four-way decoder D0 and two-way decoder D1, selectors S0, S1, and S2, and NMOS transistors SAPG0, SAPG1, and SAPG2; The sense amplifiers MSB-SA, CSB-SA and LSB-SA are all the sense amplifiers with bit line-storage node decoupling according to claim 1; The gates of SAPG0, SAPG1, and SAPG2 are connected to the CSL_LSB signal, CSL_CSB signal, and CSL_MSB signal, respectively. The sources of SAPG0, SAPG1, and SAPG2 are all connected to the local word line LDL. The drain of SAPG2 is connected to the storage node SN of MSB-SA, the drain of SAPG1 is connected to the storage node SN of CSB-SA, and the drain of SAPG0 is connected to the storage node SN of LSB-SA. The first input terminals of S0, S1 and S2 are all connected to the SASL signal, the second input terminals of S0, S1 and S2 are all connected to the bit line, and the third input terminals of S0, S1 and S2 are all connected to the complement line; the output terminals of S0, S1 and S2 are respectively connected to the bit line voltage input terminals of LSB-SA, CSB-SA and MSB-SA; GRVG2 is used to generate the reference voltage V ref3 , GRVG1 is used to generate the reference voltage V ref1 and V ref5 , GRVG1 is used to generate the reference voltage V ref0 、V ref2 、V ref4 and V ref6 ; V ref0 ~ V ref6 Increase successively; The output of GRVG2 is connected to the reference voltage input of MSB-SA; The first input terminal of D1 is connected to the output terminal of GRVG1, the second input terminal of D1 is connected to the storage node SN of MSB-SA, and the output terminal of D1 is connected to the reference voltage input terminal of CSB-SA; D1 is used to change the voltage signal at the storage node SN of MSB-SA from V ref1 and V ref5 Select one path to input to CSB-SA; The first input terminal of D0 is connected to the output terminal of GRVG0, the second input terminal of D0 is connected to the storage node SN of CSB-SA, and the output terminal of D0 is connected to the reference voltage input terminal of LSB-SA; D0 is used to change the voltage signal at the storage node SN of CSB-SA from V ref0 、V ref2 、V ref4 and V ref6 Select one input to LSB-SA; Among them, the SASL signal is used to select the sense amplifier, the CSL_LSB signal is used to select the LSB-SA, the CSL_CSB signal is used to select the CSB-SA, and the CSL_MSB signal is used to select the MSB-SA.

3. The bit line sense amplifier with cascade feedback according to claim 2, wherein the selector include: A PMOS transistor P4 and an NMOS transistor N4; The gate of P4 is connected to the gate of N4, and the connection terminal serves as the first input terminal of the bit line sense amplifier; The drain of P4 serves as the second input terminal of the bit line sense amplifier; The source of N4 serves as the third input terminal of the bit line sense amplifier; The source of P4 is connected to the drain of N4, and the connection terminal serves as the output terminal of the bit line sense amplifier.

4. A memory-in-memory microarchitecture based on open bitline ReRAM, characterized in that: In the subarray of the ReRAM, the bit line sense amplifier connected to each bit line is the cascade feedback bit line sense amplifier according to claim 2 or 3, and the bit line sense amplifiers connected to all the bit lines constitute a row buffer of the subarray; In a subarray, the bit line sense amplifiers connected to odd-numbered bit lines are located at the top of the subarray edge and are called top bit line sense amplifiers. The bit line sense amplifiers connected to even-numbered bit lines are located at the bottom of the subarray edge and are called bottom bit line sense amplifiers. In two adjacent subarrays, the bottom bit line sense amplifier of the previous subarray also serves as the top bit line sense amplifier of the next subarray.

5. The in-situ matrix calculation method of the memory-computation-in-one memory micro-architecture based on open bitline ReRAM according to claim 4, characterized in that: include: Step R1: sending a row activation command ACTP to the target subarray to activate a word line with a non-zero input in an accessed page in the target subarray; The execution of the row activation command ACTP includes: disabling the precharger in the target subarray to float the bit line; The word line voltage of the accessed page with non-zero input is increased from V BBW Rising to V PP , so that the sub-array row buffer senses the content of the row unit and takes it out to the storage node of the row buffer; Step R2: Sending a precharge command PREP to the target sub-array to close all word lines in the accessed page in the target sub-array; the execution of the precharge command PREP includes: The word line voltage in the accessed page is turned on from V PP Discharge to V BBW , so that the word line returns to a fully closed state; at the same time, the precharger is used to charge the bit line to , so that the bit line returns to the precharge state; Step R3: Sending a column VMM command to the target subarray to gradually fetch the sensing results in the row buffer of the target subarray into a memory burst buffer and transmit the results to the outside through a global I / O data line to complete data reading; the memory burst buffer includes a fixed number of I / O sense amplifiers, which are connected to bit line sense amplifiers in the I / O sense amplifiers through a shared local data line; the process of transmitting the sensing results in each bit line sense amplifier to the memory burst buffer includes: Step R31: Select MSB-SA through the SASL signal, and drive the sensing data of MSB-SA to be transmitted to the corresponding I / O sense amplifier in the memory burst buffer through the local data line, and start to trigger the flip of the I / O sense amplifier, while releasing the local data line; Step R32: After the sensing data transmission of MSB-SA is completed, CSB-SA is selected through the SASL signal, and the sensing data of CSB-SA is driven to be transmitted to the corresponding I / O sense amplifier in the memory burst buffer through the local data line, and the flip of the I / O sense amplifier is triggered, and the local data line is released at the same time; Step R33: After the sensing data transmission of CSB-SA is completed, CSB-SA is selected through the SASL signal, and the sensing data of LSB-SA is driven to be transmitted through the local data line to the corresponding I / O sense amplifier in the memory burst buffer, and the flip of the I / O sense amplifier is triggered, and the local data line is released at the same time; in, V BBW is the word line voltage in the fully off state, V PP is the word line voltage in the fully open state, Indicates a positive voltage reading.

6. The in-situ matrix calculation method according to claim 5, wherein: In step R1, the word line voltage with non-zero input in the accessed page is changed from V BBW rise to V PP During the process, when the read voltage margin of each MSB-SA in the row buffer of the target sub-array is formed, the corresponding word line is turned off.

7. The in-situ matrix calculation method according to claim 6, wherein: After step R33, the method further includes: if a subarray adjacent to the target subarray receives a column VMM command, after the sensing data transmission of LSB-SA is completed, selecting MSB-SA through a SASL signal to sense the result in MSB-SA.

Citation Information

Patent Citations

  • Memory and control method

    CN118038934A

  • Latched sense amplifier

    KR1020020085952A