Super-large-scale in-memory computing core circuit for solving IR-Drop problem
By analyzing the physical equations of the line resistance network, using collaborative design and fully differential ADC technology, the IR-Drop problem caused by line resistance in the memristor array is solved, and efficient calculation and energy efficiency of ultra-large-scale arrays are achieved.
Patent Information
- Application Number
- CN202510095240.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-21
AI Technical Summary
As the scale of the memristor array expands, the IR-Drop problem caused by line resistance between memory cells is gradually significant, affecting the voltage distribution in the array, limiting the application of in-memory computing technology and the improvement of scale.
By analyzing the internal physical equations of the line resistance network, determining the accuracy drop caused by line resistance is an approximate linear error. The coordinated design of devices, layouts and circuits is adopted to adjust the line resistance of rows and columns to maximize the linearization of errors. A fully differential ADC is designed to achieve isolated current sampling and error compensation for line resistance.
It realizes efficient computing of ultra-large-scale arrays, with computing energy efficiency of more than 13TOPS/W and computing throughput of more than 0.41TOPS, significantly improving computing accuracy and efficiency.
Smart Images

Figure CN120045511A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of microelectronics technology and artificial intelligence, and more specifically, relates to a very large scale in-memory computing core circuit for solving the IR-Drop problem. Background Art
[0002] With the rapid development of artificial intelligence technology, especially driven by deep learning algorithms, computing hardware is facing increasing performance requirements. The Computing-In-Memory (CIM) technology has attracted much attention because it integrates storage and computing in the same location. Due to its significant advantages in providing high computing power, high energy efficiency, and low latency, it has become one of the key technologies to solve the AI computing power demand.
[0003] The memristor, as a non-volatile memory device, has the dual functions of storage and computing. Its unique physical properties make it show great potential in CIM applications. By combining data storage and processing operations in the same location, the memristor can significantly reduce the data transmission distance, lower power consumption, and improve computing efficiency. Therefore, the memristor is usually used as the core component of CIM technology, and its unique data storage and processing capabilities make it possible to achieve high-performance, high-energy-efficiency, and low-latency AI hardware.
[0004] However, with the expansion of the memristor array scale, considering the existence of the mutual resistance between wires, the voltage drop IR-Drop between the nodes of the memristor array accumulates continuously, and its influence on the voltage distribution in the array becomes more and more obvious, which hinders the application of in-memory computing technology. Among them, increasing the scale of the array can directly reduce the power consumption and area ratio of the peripheral circuit, which brings an improvement in computing throughput and computing efficiency. However, due to the IR-drop problem caused by the line resistance between the storage units in the array, the improvement of scale and accuracy is limited.
[0005] Currently, there are mainly two technical methods to solve the line resistance problem. One method is the hardware solution, including the redistribution of memristor conductance and circuit compensation. The redistribution scheme generates the redistribution of memristor conductance related to the line resistance through an iterative algorithm, so that the matrix calculation result is consistent with the memristor array without line resistance. However, the redistribution of conductance requires additional computational load and a continuously distributed resistance state range. The hardware compensation scheme uses additional digital modules in the peripheral digital circuit to correct the output result of the ADC, which increases additional circuit consumption and time. The dual-power supply scheme alleviates the problem caused by line resistance by inputting positive and negative voltages to two adjacent rows, but its single-core scale is still only 144k. This scheme does not fundamentally solve the line resistance problem of larger-scale arrays. The other method is hardware-software cooperation. This method introduces a line resistance network during the neural network training process. This method highly depends on the co-training of algorithms and hardware, which hinders the rapid deployment of algorithms and requires more durability of memristors. Summary of the Invention
[0006] In view of the above defects or improvement requirements of the prior art, the present invention provides a very large-scale in-memory computing core circuit for solving the IR-Drop problem, thereby solving the IR-drop problem caused by the line resistance between storage units in the array.
[0007] To achieve the above object, according to the first aspect of the present invention, there is provided a line resistance network circuit for a non-volatile memory array, including:
[0008] An ADC module, including M ADCs and voltage drivers connected thereto, forming M ADC channels;
[0009] A two-dimensional 1T1R array, whose word line control terminal WL is parallel to the column current output terminal SL, and the bit line voltage signal input terminal BL is perpendicular to WL and SL. Two adjacent columns of 1T1R form a differential pair column for storing neural network weights;
[0010] A DAC module, including N DAC channels, and each DAC channel includes a differential readout circuit and an ADC connected in sequence;
[0011] The ADC module is connected to the BL terminal of the two-dimensional 1T1R array, and the SL terminal of the two-dimensional 1T1R array is connected to the DAC module; the size of the two-dimensional 1T1R array is m×n. One DAC channel is shared by i rows in the array and is controlled by a switch array in a time-sharing manner. One ADC is shared by j columns in the array and is controlled by a switch array in a time-sharing manner, and j > i.
[0012] According to the second aspect of the present invention, there is provided an electronic device, including: a computer-readable storage medium and a processor;
[0013] The computer-readable storage medium is used to store executable instructions;
[0014] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the method described in the first aspect.
[0015] According to the third aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to execute the method described in the first aspect.
[0016] According to the fourth aspect of the present invention, there is provided a computer program product including a computer program or instructions, which when executed by a processor implement the method described in the first aspect.
[0017] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0018] By analyzing the internal physical equations of the wire resistance network, the present invention determines that the accuracy degradation caused by wire resistance is an approximate linear error. Through the analysis of the approximate analytical formula, it is found that the linear error is only related to the parameters of the wire resistance network. This means that different input voltage vectors produce the same approximate linear error in the column output current. Therefore, through the collaborative design of devices, layouts, and circuits, the present invention adjusts the wire resistance of rows and columns to maximize the linearization of the error, and eliminates these errors through adjacent differential pair readout design and sampling resistor trimming; at the same time, in combination with the adjacent differential readout circuit, a fully differential ADC is designed and applied to achieve isolated sampling of current and error compensation of wire resistance. Through the above circuit design, taking the two-dimensional 1T1R array with M = 128 and N = 64 and a scale of 128 rows and 128 columns as an example, through simulation verification, the present invention realizes a super-large-scale array of 5 Mb level, with a computing energy efficiency of more than 13 TOPS / W and a computing throughput of more than 0.41 TOPS. Description of the Drawings
[0019] Figure 1 Among them, (a), (b), (c), and (d) are respectively one of the in-memory computing schematic diagrams, the second in-memory computing schematic diagram, the first upper electrode voltage distribution diagram, and the second upper electrode voltage distribution diagram of the non-volatile memory array provided by the embodiments of the present invention;
[0020] Figure 2 Among them, (a), (b), (c), and (d) are respectively the wire resistance output current fitting schematic diagram of the memristor array with a scale of 128 rows and 128 columns provided by the embodiments of the present invention, the error schematic diagram of linear fitting, the wire resistance network model schematic diagram ignoring the column wire resistance RSL, and the wire resistance network model being a 1D π-type wire resistance network model schematic diagram;
[0021] Figure 3 Among them, (a), (b), (c), and (d) are respectively the schematic diagrams of the differential current output results when RBL = RSL provided by the embodiments of the present invention, the schematic diagrams of the differential current output results after linear compensation, the schematic diagrams of the differential current output results when RBL > RSL, and the schematic diagrams of the differential current output results after linear compensation;
[0022] Figure 4 is the schematic diagram of the in-memory analog matrix calculation core circuit provided by the embodiments of the present invention;
[0023] Figure 5 Among them, (a), (b), (c), and (d) are respectively the schematic diagram of the 1T1R array in the in-memory analog matrix calculation core circuit provided by the embodiments of the present invention, one of the schematic diagrams of the readout circuit, another schematic diagram of the readout circuit, and the trimming resistor trimming circuit;
[0024] Figure 6 Among them, (a), (b), (c), and (d) are respectively one of the circuit-layout co-optimization layout design diagrams provided by the embodiments of the present invention, another layout design diagram, the schematic diagram of the memristor unit device structure in 1T1R, and the schematic diagram of the simulation results. Detailed implementation manners
[0025] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0026] The memristor array (m rows and n columns) with wire resistance is a large-scale wire resistance network with 2mn nodes, which can be divided into two dimensions: rows and columns. The Kirchhoff's current law (KCL) equation for the node (i, j) on the storage unit is:
[0027] Ip(i, j) = Im(i, j + 1) + Ip(i, j + 1) (1)
[0028] Among them, Ip(i,j) is the wire resistance current of the i-th row and the j-th column, Im(i,j + 1) is the storage unit current of the i-th row and the (j + 1)-th column, and Ip(i,j + 1) is the wire resistance current of the i-th row and the (j + 1)-th column. The KCL equations of all the top nodes in the i-th row are combined. After simulating the wire resistance network using the Spectre simulator, the voltage distribution of the top electrode shows a similar decreasing trend with different input voltages, which means that the output current of the array shows high linearity.
[0029] Randomly input multiple different voltage vectors Vin into the memristor array and perform linear fitting on the column output current. The linear fitting results of the output current are as follows:
[0030]
[0031] where Iact j is the current of the j-th column under different input voltage vectors, and Iideal j is the current of the j-th column without line resistance. Through the analysis of the linear fitting results of formula (2), the relative errors of all column output currents are less than 1%, which indicates that the column current error has high linearity. To maximize the linearization of the error, a line resistance network with column line resistance ignored is analyzed. Therefore, the two-dimensional line resistance network is reduced to a one-dimensional row line resistance network. Among them, the row line resistance network is a one-dimensional π-type resistance network. For this one-dimensional π-type resistance network, the voltage solution of the upper electrode of the storage cell can be obtained by solving the following second-order difference equation and boundary conditions.
[0032] G BL ·(V (j) -V (j-1) )-G BL ·(V (j) -V (j+1) )-G mean ·(V (j) )=0 (3)
[0033] where G mean represents the average conductance of the memristor array. The analytical solution can be expressed as:
[0034]
[0035] In formula (4), h and the function F are polynomial functions related to GBL and Gmean. The voltage solution Vt(i,j) of the i-th row reveals that all the top node voltages are the results related to Vin(i) and the coefficient kj, where kj is independent of Vin(i). Therefore, the voltage solutions of all the upper electrodes of the storage cells are:
[0036] Vt=Vin T ·K (5)
[0037] Among them, Vt is the voltage vector of all top nodes, Vin is the input vector, and K is the coefficient vector derived from kj in formula (4). The derivation result of the above top node voltage shows that under different Vin inputs, the column output current error exhibits a high degree of linearity. At the same time, it can also be concluded that the column line resistance RSL only affects the bias b of the linear error in formula (2). This analysis result provides a method for circuit compensation. By reducing the column line resistance RSL, higher linearity of the error can be achieved. In addition, by using adjacent columns as a differential pair to store weights, the current output result can be expressed as:
[0038] Iideal = Ioutp - Ioutn ≈ K × (Ioutp - Ioutn) (6)
[0039] In formula (6), Ioutp and Ioutn represent the positive and negative column output currents of the differential pair respectively. Iideal is the ideal current without line resistance after linear compensation. Due to the adjacent differential pair structure, the bias b of the linear error is eliminated, and the remaining linear error coefficient k can be corrected by the sampling resistor. Map the neural network weights to the memristor array with line resistance, where adjacent columns store weights in the form of a differential pair. When RBL > RSL, the accuracy after linear compensation reaches 92.4%, which proves the feasibility of the differential linear compensation scheme.
[0040] Based on this, an embodiment of the present invention provides a line resistance network circuit for a non-volatile memory array, including:
[0041] An ADC module, including M ADCs and voltage drivers connected thereto, forming M ADC channels;
[0042] A two-dimensional 1T1R array, whose word line control terminal WL is parallel to the column current output terminal SL, and the bit line voltage signal input terminal BL is perpendicular to WL and SL. Adjacent two columns of 1T1R form a differential pair column for storing neural network weights;
[0043] A DAC module, including N DAC channels, and each DAC channel includes a differential readout circuit and an ADC connected in sequence;
[0044] The ADC module is connected to the BL terminal of the two-dimensional 1T1R array, and the SL terminal of the two-dimensional 1T1R array is connected to the DAC module; the size of the two-dimensional 1T1R array is m×n, one DAC channel is shared by i rows in the array and is controlled by a switch array in a time-sharing manner, and one ADC is shared by j columns in the array and is controlled by a switch array in a time-sharing manner, and j > i.
[0045] Specifically, the circuit provided by the present invention includes:
[0046] M input DACs and drive circuit channels are used to convert a digital-domain input vector into an analog voltage input vector.
[0047] A memristor-based non-volatile memory array with a 1T1R structure, where the word line control terminal WL is parallel to the column current output terminal SL, and the bit line voltage signal input terminal BL is perpendicular to WL and SL. Two adjacent columns of 1T1R form a differential pair column for storing neural network weights. The layout of the array adopts a layout-circuit co-design, solves the IR-drop problem caused by line resistance, realizes ultra-large-scale circuit design, and has a computing power greater than 0.4 TOPS@int8 and a computing energy efficiency greater than 13 TOPS / W@int8; that is, the 1T1R array adopts a structure where WL is parallel to the column current output SL, and the bit line voltage signal input terminal BL is perpendicular to WL and SL. The voltage signal is pre-loaded onto the BL terminal, and then a pulse signal is applied to WL for matrix operations in the analog domain. Two adjacent columns of 1T1R serve as a differential pair column for storing neural network weights.
[0048] An N-channel differential readout circuit and ADC compensate for the voltage drop error caused by line resistance on the circuit through the adjustment of trimming resistors.
[0049] M DACs are connected to the BL terminals of the storage array in item b, and voltage is loaded onto the array. The SL terminals of the array described in item b are connected to the ADC and readout circuit described in item c for signal readout.
[0050] The storage array has m rows and n columns. One DAC and drive circuit channel is shared by i rows in the array and is controlled by a switch array in a time-sharing manner. One ADC and readout circuit channel is shared by j columns in the array and is controlled by a switch array in a time-sharing manner. To ensure the compensation effect, it is necessary to make j greater than i, so that the line resistance of the row is greater than that of the column.
[0051] Preferably, the two-dimensional 1T1R array is set as a rectangle on the layout, and the length of the rectangle is greater than the width.
[0052] Specifically, using design technology co-optimization (DTCO), the 1T1R memory cell is set as a rectangle on the layout, and the length of the rectangle is greater than the width, so that the BL line resistance is greater than the SL line resistance. This design linearizes the error caused by line resistance to the maximum extent, which is beneficial for subsequent peripheral circuits to compensate for the error.
[0053] Preferably, the differential readout circuit uses a folded cascode as the differential input stage and NMOS as the common-source amplification output stage to perform voltage clamping and current isolation sampling on two adjacent columns of 1T1R. The sampled current is converted into a voltage by a trimming resistor trimming circuit.
[0054] Specifically, the dual - ended differential readout circuit performs voltage clamping and current isolation sampling on adjacent two columns of 1T1R. Then the sampled current is converted into voltage through the trimming resistor, which not only compensates for the line resistance error but also samples for the ADC.
[0055] The dual - ended differential readout circuit uses a folded cascode as the differential input stage and NMOS as the common - source amplifier output stage to clamp the voltage at the SL end of the array for differential readout. In addition, the output stage uses NMOS mirror transistors for current isolation sampling, and then outputs to the ADC for quantization. The current after isolation sampling passes through a trimming resistor trimming circuit to compensate for the line resistance error. The offset term of the linear error is eliminated by subtracting the differential pair, and the proportional term of the linear error is compensated by the trimming resistor trimming circuit.
[0056] Preferably, the DAC is a C - 2C DAC, and the voltage driver is a class - AB output operational amplifier.
[0057] A C - 2C DAC is used to convert the original digital signal into an analog voltage signal, and then a class - AB output operational amplifier is used as the voltage driver to drive the BL end of the array.
[0058] Preferably, SL is implemented with the top metal layer, and its sheet resistance is lower than that of the bottom metal layer used for BL; BL is implemented with M2 or M3 metal layer; SL is implemented with the top M6 metal layer.
[0059] Figure 1 The in - memory computing based on the non - volatile memory array is shown. Among them, the array size is 128 rows and 128 columns, and the row line resistance RBL and the column line resistance RSL are both 1 ohm. As Figure 1 shown in (a) and (b) of [], a voltage excitation is applied to BL at the left end of the cross - array (i.e., the bit - line voltage signal input end), and then SL at the lower end of the array (i.e., the column - current output end) is clamped to a fixed level. According to Kirchhoff's theorem, matrix multiplication in the analog domain is performed, and the final result is the output current at the SL end. Figure 1 Shown in (c) and (d) of [] is the upper - electrode voltage distribution. After mapping the upper - electrode voltage to a two - dimensional plane, it can be found that there is a similar voltage distribution trend for the upper - electrode voltage. This distribution trend reveals the existence of the output current.
[0060] Figure 2 The linear fitting result of the line - resistance output current is shown. Figure 2 Shown in (a) of [] is a memristor array with a size of 128 rows and 128 columns, and the output current of this array is given by the formula (2) Iideal j =K×Iact j+b fitting, where K is the linear coefficient of the linear fitting and b is the linear offset of the linear fitting. Figure 2 (b) shows the error of the linear fitting, and the result shows that the linear fitting error is less than 1%. This fully proves the linearity of the output current. Figure 2 (c) in shows the line resistance network model that ignores the column line resistance RSL. By Figure 2 separating the network model of (c) in Figure 2 (d) in shows that the line resistance network model is a one-dimensional π-type line resistance network model. By solving the difference equation formula (3) of the one-dimensional π-type line resistance network, the analytical solution of the voltage distribution on the upper electrode of the storage unit can be obtained. The result of the analytical solution is formula (4), which proves that the current result can achieve high linearity, but its core lies in how to reduce the column line resistance.
[0061] Figure 3 shows the comparison of the current results of the linear compensation. From Figure 2 the results solved in, it can be seen that increasing the row line resistance RBL and reducing the column line resistance RSL can achieve the maximum linearization of the line resistance error. Figure 3 (a) in shows the differential current output result of RBL = RSL, and the error caused by the line resistance reaches 60%. After the linear compensation scheme, the error caused by the line resistance still reaches 50%. Figure 3 (c) in shows the layout optimization result of RBL > RSL. Without the line resistance compensation scheme, the error caused by the line resistance reaches 43%. Figure 3 (d) in shows the result after the linear compensation scheme, and the error caused by the line resistance is only 7.6%. This simulation result proves the feasibility of the linear error compensation scheme after layout optimization.
[0062] Figure 4 shows the in-memory analog matrix calculation core circuit provided by the present invention based on the above-described linear compensation scheme. Taking M = 128, N = 64, and the two-dimensional 1T1R array scale of 128 rows and 128 columns as an example, through layout-circuit co-optimization, the present invention designs and verifies a matrix operation core of a 5Mb array. While meeting the calculation accuracy requirements, this matrix operation core achieves excellent calculation throughput and calculation efficiency. Based on the SIMC 55nm technology, the design of the CIM macro is as Figure 4 shown. In order to meet the condition of RBL being greater than RSL, the design adopts M DAC input channels and N ADC and readout circuit channels, as Figure 4 shown. Since the area of the DAC and ADC is significantly larger than that of the 1T1R unit, one DAC input channel is shared by i rows. The sub-array selected by the switch control circuit participates in the single-cycle calculation, and the word lines (WLs) of other columns are turned off. As Figure 5As shown in (a) of [reference], two adjacent columns of 1T1R form a differential pair, and a total of j / 2 differential pairs are formed. The readout circuit for the output current is designed as shown in Figure 5 (b) of [reference]. Two clamping operational amplifiers are used to read the differential pair current composed of two columns, then convert it into a voltage through two Rtriming circuits, and sample it by an ADC. As shown in Figure 5 (c) of [reference], the Vout port of the clamping operational amplifier is used to clamp the column output of the array to Vclamp, where its output stage NMOS is mirrored to the current Iout. Therefore, the column output current is isolated and sampled. As shown in Figure 5 (d) of [reference], by adjusting the sampling Rtriming, linear compensation of the column output current is achieved.
[0063] Through Design Technology Co-optimization (DTCO), co-design is carried out on circuit-layout-device, Figure 6 showing the results of circuit-layout co-optimization of the present invention. Based on the error analysis results, in order to maximize the linearity of the error, the present invention adopts the layout design shown in Figure 6 (a) and (b) of [reference]. By designing the 1T1R unit as a rectangle, the width of the source line (SL) is made greater than the width of the bit line (BL). At the same time, SL uses the top metal layer, whose sheet resistance is lower than that of the bottom metal layer used by BL. As shown in Figure 6 (c) of [reference], M2&M3 metal layers are used for BL in the array, while the top M6 metal layer with a lower sheet resistance is used for SL in the array. The memristor unit in 1T1R uses the device structure shown in Figure 6 (c) of [reference]. The circuit is post-simulated, and the extracted parasitic parameters are brought into the neural network simulation model based on the memristor in-memory computing core. The results show that the recognition rate of the neural network is restored.
[0064] It is easy for those skilled in the art to understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A line resistance network circuit of a non-volatile memory array, characterized in that: include: An ADC module, including M ADCs and voltage drivers connected thereto, forming M ADC channels; A two-dimensional 1T1R array, wherein the word line control terminal WL is parallel to the column current output terminal SL, the bit line voltage signal input terminal BL is perpendicular to WL and SL, and two adjacent columns of 1T1R form a differential pair column for storing neural network weights; A DAC module, including N DAC channels, each DAC channel including a double-ended differential readout circuit and an ADC connected in sequence; The ADC module is connected to the BL end of the two-dimensional 1T1R array, and the SL end of the two-dimensional 1T1R array is connected to the DAC module; the size of the two-dimensional 1T1R array is m×n, one DAC channel is shared by i rows in the array and is time-sharing controlled by the switch array, one ADC is shared by j columns in the array and is time-sharing controlled by the switch array, and j>i.
2. The circuit according to claim 1, characterized in that The two-dimensional 1T1R array is arranged as a rectangle on the layout, and the length of the rectangle is greater than the width.
3. The circuit according to claim 1, characterized in that The double-terminal differential readout circuit uses a folded common source and common gate as a differential input stage and an NMOS as a common source amplifier output stage, and is used to perform voltage clamping and current isolation sampling on two adjacent columns of 1T1R. The sampled current is converted into voltage through a trimming resistor adjustment circuit.
4. The circuit according to claim 1, characterized in that The DAC is a C-2C DAC, and the voltage driver is a class AB output operational amplifier.
5. The circuit according to claim 1, characterized in that SL is implemented using a top metal layer, which has a lower sheet resistance than the bottom metal layer used in BL.
6. The circuit according to claim 1, characterized in that BL is implemented using the M2 or M3 metal layer.
7. The circuit according to claim 1, characterized in that SL is implemented using the top M6 metal layer.
Citation Information
Patent Citations
Equation set solver based on memristor linear neural network and operation method of equation set solver
CN111460365A
In-memory calculation circuit and compensation method thereof, memory device and chip
CN115731994A
High-energy-efficiency in-memory computing circuit
CN117520261A
Crossbar circuits for analog computing
WO2024263809A2
Cited By
Read-write multiplexing circuit applied to RRAM storage and calculation array and RRAM storage and calculation system
CN120612986A