In-memory computing architecture based on nonvolatile memory

By using a 2-transistor + 1 memory cell structure and nonvolatile memory in the in-memory computing unit, the existing in-memory computing unit has solved the problems of low storage density and complex operation, and efficient in-memory multiplication and adding calculation and low power consumption computing capabilities are achieved.

CN120108464APending Publication Date: 2025-06-06ZHEJIANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510166017.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing in-memory computing units have low storage density and complex operations, making it difficult to meet the needs of the fields of deep learning and artificial intelligence for efficient computing and low power consumption.

Method used

A in-memory computing architecture based on a 2-transistor + 1 memory cell (2T1M) structure is proposed, and the in-memory multiplication and addition calculation is realized using non-volatile memory. Through partial multiplexing of read and write and calculation structures, the area overhead of chip design is reduced.

Benefits of technology

It greatly improves the storage density of data, reduces computing power consumption and data handling overhead, and improves the energy efficiency and computing power of in-memory computing modules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108464A_ABST
    Figure CN120108464A_ABST
Patent Text Reader

Abstract

The invention discloses an in-memory computing architecture based on a nonvolatile memory. The in-memory computing architecture comprises an in-memory computing unit array, a bit line input and output control logic module, a word line input driving module, an output control logic module, a DAC (Digital-to-Analog Converter) and an ADC (Analog-to-Digital Converter), each in-memory calculation unit is structurally composed of two transistors and one storage unit, and read-write and operation functions of the circuit can be realized by switching on and switching off the gating transistors; each storage unit stores the weight in a conductive form, point multiplication operation of an input vector and a weight matrix is completed by utilizing an Ohm law and a Kirchhoff law, input data is applied in an analog signal form after being subjected to digital-to-analog conversion, and a calculation result is given after being subjected to analog-to-digital conversion. A single calculation unit only comprises two transistors and one nonvolatile device, the storage density is further improved through source line multiplexing, compared with a traditional storage calculation framework based on an SRAM and a DRAM, the area and power consumption of a storage calculation circuit are greatly reduced, and the calculation efficiency of the deep network is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of novel storage and computing, and specifically relates to an in-memory computing architecture based on non-volatile memory. Background Art

[0002] As a key technology in the field of deep learning and artificial intelligence, neural networks have been widely used in image recognition, autonomous driving, large language models, etc. With the increasing amount of data and model size, the overhead of moving data between different modules and different levels of memory has seriously restricted the performance of the system. Computing in Memory (CiM), as an emerging technology that has received widespread attention, stores weights in a memory array to implement matrix multiplication, avoiding the repeated movement of weight data, and can greatly improve the computing efficiency and speed of the system.

[0003] The traditional in-memory computing architecture based on static random access memory (SRAM) is limited by its standard 6-transistor unit structure, and the storage density is difficult to improve, which seriously limits its possible application scenarios. The in-memory computing architecture based on dynamic random access memory (DRAM) requires continuous refreshing of storage weights, which is contrary to the low power consumption requirements in the field of edge computing.

[0004] In recent years, the implementation of in-memory computing based on new non-volatile memory (NVM) has been widely studied. The main feature of non-volatile memory is that it has extremely high data retention capability, and the stored data will not be lost after power failure. Compared with SRAM cells, cells based on non-volatile memory only require one device and a gate transistor to cooperate to implement, and their area has obvious advantages; compared with DRAM cells, the static power consumption of non-volatile memory devices is greatly reduced, and the energy efficiency has significant advantages. Currently, the non-volatile storage structures that have received more attention include resistive random access memory (ReRAM), magnetic tunnel junction (MTJ), phase change memory (PCM) and ferroelectric tunnel junction (FTJ). However, most of the existing storage and computing units based on non-volatile memory still have the characteristics of a large number of transistors and complex read and write operations. Summary of the invention

[0005] In order to overcome the shortcomings of low storage density and complex operation of existing in-memory computing units, the present invention proposes a new in-memory computing architecture based on non-volatile memory. This architecture can efficiently implement in-memory multiplication and addition calculations, and can greatly improve data storage density and reduce computing power consumption and data handling overhead compared to existing technical solutions.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] An in-memory computing architecture based on non-volatile memory includes: an in-memory computing unit array, a bit line (BL) input and output control logic module, a word line (WL) input driver module, an output control logic module, a digital-to-analog converter (DAC) and an analog-to-digital converter (ADC);

[0008] The in-memory computing unit array comprises a plurality of rows and columns of in-memory computing units, each of which comprises two transistors and a storage unit, and each storage unit may be a non-volatile storage device with two terminals (including but not limited to MTJ, FTJ, ReRAM, etc.), one terminal of the storage unit is connected to a bit line, and the other terminal is respectively connected to the source terminal and the drain terminal of two transistors; each row of in-memory computing units is connected via two adjacent source lines (SL), and two adjacent rows of in-memory computing units share one source line, and the transistors in the same row of in-memory computing units are controlled by two word lines as gates; each column of in-memory computing units is connected via the same bit line;

[0009] The bit line input and output control logic module is connected to each column of the bit line of the in-memory computing unit array, and is used to control the columns that need to participate in calculation or reading and writing in a specific cycle and apply specific input signals;

[0010] The word line input driving module is connected to each row of word lines of the in-memory computing unit array, and is used to control the opening or closing of the rows involved in the calculation or reading and writing in a specific cycle;

[0011] The output control logic module is connected to each row source line of the in-memory computing unit array, and is used to control the output source lines involved in calculation or reading and writing in a specific cycle to be grounded or suspended;

[0012] The digital-to-analog converter is used to convert the input value to be calculated into an analog value (including but not limited to a voltage pulse, a voltage pulse sequence, a current pulse or a current pulse sequence, etc.) in the in-memory calculation operation;

[0013] The analog-to-digital converter is used to convert the output analog quantity of the output control logic module into a digital result in the in-memory calculation operation.

[0014] Furthermore, the in-memory computing architecture also includes a control module for coordinating data communication and clock control among modules in the in-memory computing architecture.

[0015] Further, the in-memory computing architecture circuit includes a read-write mode and an operation mode;

[0016] When the circuit is in the read-write mode, the word line input drive module will set the word line of the corresponding row to a high potential according to the address to be read or written, so that the two transistors inside each in-memory computing unit of the row are in the on state; the bit line input and output control logic module turns on the specific bit line according to the address to be read or written; the output control logic module grounds the source line of the corresponding row according to the address to be read or written, and the other rows are left hanging. At this time, there is only one storage cell in the in-memory computing unit array that is in a readable and writable state; applying a voltage or current greater than the voltage or current required for switching the state of the storage cell to the bit line can switch the high and low resistance states of the gating device, that is, realize the programmability of the device; when a read operation voltage is applied to the bit line, the current resistance state of the gating device can be known by measuring the transient current on the source line;

[0017] When the circuit is in operation mode, the in-memory computing unit array stores weights in the form of conductance. If each storage unit has q distinguishable resistance states, then the weights in the same column of the array are Each in-memory computing unit stores a Q-bit weight; the word line input driving module enables one transistor in each in-memory computing unit in the corresponding row to be in an on state and the other to be in an off state according to the array size required by the operation scale; the input data is applied to each column determined by the bit line input and output control logic module in an analog quantity, and multiple columns can be simultaneously turned on to participate in the operation in the same clock cycle; the output control logic module transmits the output analog result corresponding to the row selected to participate in the calculation to the analog-to-digital converter, and obtains the digital calculation result after conversion, that is, completing a dot multiplication operation of the input vector and the Q-bit weight matrix.

[0018] The beneficial effects of the present invention are as follows: an in-memory computing architecture based on a 2-transistor + 1 storage unit (2T1M) structure is proposed, which realizes partial reuse of read-write and computing structures in the circuit, reducing the area overhead of chip design; in addition, the design of in-memory computing using non-volatile memory can greatly reduce the time cost caused by weight transfer and improve the energy efficiency and computing power of the in-memory computing module. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A schematic diagram of an in-memory computing architecture based on a non-volatile memory in the present invention;

[0020] Figure 2 A schematic diagram of the working state of the in-memory computing architecture circuit in the read-write mode of the present invention;

[0021] Figure 3 A schematic diagram of the waveforms of the in-memory computing architecture circuit read and write modes in the present invention;

[0022] Figure 4 A schematic diagram of the working state of the in-memory computing architecture circuit operation mode in the present invention;

[0023] Figure 5 A schematic diagram of waveforms of the operation mode of the in-memory computing architecture circuit in the present invention;

[0024] Figure 6 It is a schematic diagram of the result of using the in-memory computing architecture circuit in the present invention for matrix multiplication. DETAILED DESCRIPTION

[0025] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments, which is intended to provide a basic understanding of the present invention, and is not intended to confirm the key or decisive elements of the present invention or the scope to be protected. It is easy to understand that without changing the essential spirit of the present invention, various replacements, changes and modifications are possible by those skilled in the art without departing from the spirit and scope of the present invention and the attached claims. Therefore, the following specific implementation methods and the accompanying drawings are only exemplary descriptions of the technical solution of the present invention, and should not be regarded as the entirety of the present invention or as a limitation or restriction of the technical solution of the present invention.

[0026] The purpose of the present invention is to provide an in-memory computing architecture based on a non-volatile memory. The present invention is further described below in conjunction with the accompanying drawings and specific embodiments.

[0027] Figure 1 The schematic diagram of the in-memory computing architecture based on non-volatile memory is shown. The in-memory computing architecture includes an in-memory computing unit array CIM Array, a bit line input and output control logic module MUX BL_IO, a word line input driver module WLDrv, an output control logic module MUX CIM_IO, a digital-to-analog converter DAC, an analog-to-digital converter ADC and a control module CTRL. The working mode (read-write mode or operation mode) of the circuit can be controlled by the control module, and the word line input driver module selects the two word lines WL corresponding to the row according to the working mode and the operation address. i and WLX i Set to high or low potential; the bit line input and output control logic module selects the bit line BL to be turned on according to the operation address in the read and write mode j And apply the read operation voltage or the write operation voltage, and in the operation mode, apply the analog signal obtained by the digital-to-analog converter to the turned-on bit line BL j The output control logic module selects the source line SL according to the working mode and operation address. k The digital calculation result is obtained by setting it to a low potential or converting it through an analog-to-digital converter.

[0028] Figure 2The schematic diagram of the working state of the in-memory computing architecture circuit read and write mode is shown. Without loss of generality, here we take the storage value of the in-memory computing unit in the first row and first column of the array as an example. In the current state, the two word lines WL in the first row 0 and WLX 0 The source line SL on the output side is set to high potential, and the other word lines are set to low potential, so that the two transistors in each memory calculation unit in the first row are turned on; at the same time, the source line SL on the output side is turned on. 0 and SL 1 Ground, other source lines are left floating; input signal is applied to the bit line BL of the first column 0 , and the other bit lines are left hanging. If a read signal is applied at this time, the SL 0 and SL 1 The storage state of the storage unit is determined by the voltage or current signal amplitude detected; if a write signal is applied, it should be ensured that the amplitude of the signal is greater than the voltage or current required for switching the state of the storage unit.

[0029] Figure 3 The schematic diagram of the waveform of the in-memory computing architecture circuit read and write mode is shown. Without loss of generality, assuming that the transistors in each in-memory computing unit are NMOS, the array size is (N+1)×(M+1), and here we take the storage value of the in-memory computing unit in the first row and first column of the array as an example. When reading and writing, WL <0> and WLX <0> Put it at high potential, WL <n:1>and WLX <n:1>Put it in low potential, at BL <0> Apply read and write operation pulses, BL <m:1>Hanging, at this time in SL <0> and SL <1> The output pulse signal can be observed on the SL<N+1:2> In suspended state.

[0030] Figure 4 The schematic diagram of the working state of the in-memory computing architecture circuit operation mode is shown. Without loss of generality, here we take the input of 2×1 vector and 2×2 matrix dot multiplication as an example. <0> and WL <1> Set to high potential, and other word lines to low potential; at the same time, the output side SL 0 and SL 1 Ground, other source lines are left floating; the analog input signal converted by DAC is applied to BL 0 and BL 1 , and the other bit lines are left hanging. 0 and SL 1 The voltage or current signal detected on the sensor is converted into a digital signal by ADC to obtain the calculation result. This example shows the following calculation process:

[0031]

[0032] Figure 5 The schematic diagram of the waveform of the in-memory computing architecture circuit operation mode is shown. Without loss of generality, assuming that the transistors in each in-memory computing unit are NMOS, the array size is (N+1)×(M+1), and if all N+1 rows participate in the calculation, then WL <n:0>Set to high potential, and WLX <n:0>Set to low potential; at the same time, the output side SL <n:0>Ground and at BL <m:0>Apply the analog signal converted by DAC to SL <n:0>The voltage or current signal on the ADC is converted into a digital signal, and the calculation result of the vector matrix dot product can be obtained.

[0033] Figure 6 The schematic diagram of the in-memory computing architecture circuit used for matrix multiplication is shown. Without loss of generality, here we take a 64×64 in-memory computing unit array as an example, where the storage unit uses MTJ to store weights. In the operation mode, keep the input on each BL unchanged, change the number of BLs selected, and measure the <0> The current results are shown in the figure. It can be seen that the calculation results increase linearly with the number of BLs turned on, indicating that the in-memory computing architecture has good calculation accuracy and feasibility.

[0034] The above is only a preferred embodiment of the present invention. Although the present invention has been disclosed as a preferred embodiment, it is not intended to limit the present invention. Any technician familiar with the art can make many possible changes and modifications to the technical solution of the present invention by using the above disclosed methods and technical contents without departing from the scope of the technical solution of the present invention, or modify it into an equivalent embodiment of equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the scope of protection of the technical solution of the present invention.

Claims

1. An in-memory computing architecture based on non-volatile memory, characterized in that: It includes an in-memory computing unit array, a bit line input and output control logic module, a word line input drive module, an output control logic module, a digital-to-analog converter, and an analog-to-digital converter; The in-memory computing unit array comprises in-memory computing units of multiple rows and columns, each in-memory computing unit comprises two transistors and a storage unit, the storage unit is a non-volatile memory device with two ends, one end is connected to a bit line, and the other end is respectively connected to the source end and the drain end of two transistors, each row of in-memory computing units is connected via two source lines, two adjacent rows of in-memory computing units share one source line, the transistors in the same row of in-memory computing units are controlled by two word lines, and each column of in-memory computing units is connected via one bit line; The bit line input and output control logic module is used to control whether the bit lines of each column in the in-memory computing unit array are enabled or not, select the enabled column by column address decoding, and apply the read / write signal or operation signal to a specific bit line; The word line input driving module is used to control the potential of each row of word lines in the in-memory computing unit array, and select the turned-on row through row address decoding; The output control logic module is used to control whether the source lines of each row of the in-memory computing unit array are enabled or not, and selects the enabled row by decoding the row address; The digital-to-analog converter is used to convert the input data in digital form during the operation into analog quantity and then apply it to the bit line selected by the bit line input and output control logic module; The analog-to-digital converter is used to convert the operation analog quantity on the source line selected by the output control logic module into a digital result.

2. The in-memory computing architecture based on non-volatile memory according to claim 1, characterized in that: The storage unit includes a resistive random access memory, a magnetic tunnel junction, a phase change memory, and a ferroelectric tunnel junction.

3. The in-memory computing architecture based on non-volatile memory according to claim 1, characterized in that: The two transistors in the in-memory computing unit are connected in series, the common end is connected to the negative electrode of the non-volatile memory device, the other end of each transistor is connected to a source line respectively, and the positive electrode of the non-volatile memory device is connected to the bit line.

4. The in-memory computing architecture based on non-volatile memory according to any one of claims 1 to 3, characterized in that: The architecture circuit includes a read / write mode and an operation mode.

5. The in-memory computing architecture based on non-volatile memory according to claim 4, characterized in that: In the read-write mode, the bit line input-output control logic module selects a column of in-memory computing units, the word line input drive module and the output control logic module jointly select a row of in-memory computing units, and both selection transistors in the row of in-memory computing units are in the on state, thereby determining an in-memory computing unit to be read or written.

6. The in-memory computing architecture based on non-volatile memory according to claim 5, characterized in that: In the read-write mode, if a read signal is applied to the bit line, the storage state of the memory cell is determined by the voltage or current signal amplitude detected on the source line; If a write signal is applied to the bit line, it should be ensured that the amplitude of the write signal is greater than the voltage or current required for switching the state of the memory cell.

7. The in-memory computing architecture based on non-volatile memory according to claim 4, characterized in that: In the operation mode, the bit line input and output control logic module selects multiple columns of in-memory computing units, the word line input drive module and the output control logic module jointly select multiple rows of in-memory computing units, and only one transistor in each in-memory computing unit of the selected row is in the on state, thereby determining the in-memory computing units participating in the vector-matrix dot multiplication operation.

Citation Information

Cited By

  • Pushing and training integrated SRAM digital in-memory computing architecture

    CN121349958A