Storage and calculation processing method and device based on capacitance-free DRAM (Dynamic Random Access Memory)
Through the combination of heterogeneous integrated capacitive DRAM memory cells and convolutional kernel logic circuits, the high cost and complexity of the DRAM in-memory processing architecture is solved, and efficient memory processing is achieved.
Patent Information
- Application Number
- CN202510606703.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-19
AI Technical Summary
The existing technology DRAM in-memory processing architecture has high cost, incompatible preparation processes, high circuit complexity, and poor memory processing speed.
Write transistors and read transistors are prepared using oxide semiconductors and silicon transistors, and capacitive-free DRAM memory cells are stacked vertically through heterogeneous integration strategies, and an in-memory processing architecture is constructed in combination with memory cell peripheral circuits and convolutional core logic circuits.
Significantly reduce preparation costs, improve in-memory processing performance and reliability, simplify circuit structure, and improve memory processing speed.
Smart Images

Figure CN120510889A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of electronic device manufacturing, and in particular to a storage and calculation processing method and device based on a capacitor-free DRAM memory. Background Art
[0002] A convolutional neural network (CNN) is primarily composed of an input layer, convolutional layers, activation layers, pooling layers, fully connected layers, and an output layer. The input layer receives raw image data, typically consisting of three color channels. The convolutional layer performs convolution operations on the input image data to extract distinct features from the image. The activation layer introduces nonlinearity into the neural network, accelerating training and alleviating the vanishing gradient problem. The pooling layer is used to reduce the size of the feature map, thereby reducing computational complexity. Finally, the fully connected layer integrates global features at the end of the network and maps them into the network's final output for classification or regression tasks. Due to its high parameter efficiency and translation invariance, CNNs are widely used in tasks such as image classification, object detection, and video analysis. Classic models such as LeNet-5 and ResNet have driven technological innovation in computer vision and are highly efficient and versatile.
[0003] Heterogeneous integration (HI) refers to the integration of components made of different materials (such as silicon, oxides, and two-dimensional materials), different process nodes (such as CMOS and emerging memories), or different storage mechanisms (such as charge storage, resistive random access memory, and phase change memory) into the same chip or package. Common heterogeneous integration solutions include: vertical stacking (3D Integration), such as stacking multiple layers of DRAM or new memory (such as RRAM and MRAM) on a logic chip, interconnected by through-silicon vias (TSVs) to increase bandwidth and reduce latency; hybrid memory cell design, integrating volatile (such as SRAM) and non-volatile (such as ReRAM) memory cells into the same chip, taking into account both high-speed caching and data persistence; material heterogeneous integration, such as using two-dimensional materials (such as graphene and MoS) combined with traditional silicon-based CMOS to manufacture ultra-low power or ultra-high-density memory cells.
[0004] Compute in Memory (CiM) performs computational operations directly within the storage unit, deeply integrating storage and computing. Instead of moving data to a separate computing unit (such as a CPU or GPU), computations are performed directly using the physical properties of the storage medium (such as resistance, capacitance, and charge). For example, matrix multiplication and addition operations can be implemented using a ReRAM crossbar array (leveraging Ohm's law and Kirchhoff's laws), breaking the traditional "storage-computation separation" model and eliminating data transfer bottlenecks. The main challenges are: limited accuracy (analog computing), storage unit reliability (such as the durability of resistive switching devices), and high design complexity (the need for customized circuits and algorithms).
[0005] Process Near Memory (PNM / Compute Near Memory, CNM) is defined as placing computing units in close proximity to memory (as in the same package or chip), while maintaining physical separation between storage and computing. High-bandwidth interconnects (such as through-silicon vias (TSVs)) are used to shorten data transmission distances. For example, 3D stacking of HBM (High Bandwidth Memory) and GPUs / CPUs primarily focuses on digital computing and relies on traditional digital logic circuits, but physical proximity reduces latency. The main challenges are heat dissipation (3D stacking increases heat density); high costs (due to advanced packaging technologies such as TSVs); and theoretically, computational efficiency is not as high as in-memory computing due to data handling. Its key feature is that the memory devices are fabricated using storage processes, while the computing and logic devices are fabricated using corresponding Si CMOS logic processes, and the two are then integrated using packaging technology.
[0006] Process in Memory (PiM) deploys computing units near the location where data is stored (such as inside the storage controller or SSD), offloading some computing tasks from the host CPU to the storage end. For example, some computing circuits are designed directly outside the peripheral circuits of DRAM, and all are manufactured using storage processes. Its main advantage is that some data computing and processing are completed directly near the storage location, which saves some power consumption of data transportation compared to traditional architectures; it has lower design complexity than the architecture of in-memory computing; and compared with the near-memory processing / computing architecture, computing and storage have a more direct and closer physical connection relationship, higher computing efficiency, lower data transportation costs, and no need for various additional packaging technologies. The problem is that the computing circuits manufactured using storage processes are more oriented towards low leakage rather than high speed, so their processing speed is naturally not as high as the computing circuits manufactured using pure logic processes.
[0007] Traditional dynamic random access memory (DRAM) technology uses one transistor and one capacitor (1T1C) to store one bit of information. The transistor switches control the reading and writing of information, while the capacitor stores the information. To address applications with larger storage capacities, DRAM storage density has significantly increased over the past few decades, with the size of the transistors and capacitors in DRAM memory cells becoming smaller per unit area. However, as devices continue to shrink, they pose increasing challenges to process and reliability.
[0008] 2T0C DRAM is composed of two transistors (Transistor) and does not require additional capacitors (Capacitor) to store information. The charge is stored on the gate capacitance of the read transistor. This DRAM technology is an efficient, compact and low-power storage solution. People have explored the integration of 2T0C DRAM on silicon in the early days. However, the charge retention time of traditional silicon-based 2T0C DRAM has always been low because the turn-off current of the write transistor is relatively high and the gate capacitance of the read transistor is smaller than that of traditional capacitors. This has become the biggest factor limiting its application.
[0009] 2T0C DRAM memory cells made of oxide semiconductors (OS), such as IGZO, enable efficient storage in small form factors. Compared to traditional silicon-based devices, thin-film transistors made from these oxide semiconductors have extremely low leakage. As a result, the retention time of small-scale 2T0C devices can be increased by orders of magnitude compared to silicon devices, meeting the requirements for data caching.
[0010] At present, there are some in-memory computing structures based on capacitor-less DRAM memory, which regard the storage unit as the computing unit, focus on using the storage unit itself for calculations, and use peripheral circuits for data processing; relatively speaking, its circuit structure is more complex, and the accuracy of the analog calculations implemented by the unit itself is low.
[0011] The existing technology at this stage mainly has the following shortcomings:
[0012] 1. DRAM memory transistors require low leakage, while logic and computing transistors require high speed, so their manufacturing processes are incompatible. The existing DRAM in-memory processing architecture has high process costs. DRAM write transistors are Si transistors, which have certain retention time limitations.
[0013] 2. The in-memory computing architecture based on capacitor-less DRAM memory treats the storage unit as the computing unit and focuses mainly on in-memory computing. This requires additional design units or processing circuits, resulting in high circuit complexity.
[0014] 3. Currently, the storage and computing processing based on capacitor-free DRAM memory is generally full-back-end capacitor-free DRAM, which reads to the peripheral and computing circuits and has a poor speed.
[0015] In summary, the existing DRAM in-memory processing architecture has high process costs, incompatible manufacturing processes, high circuit complexity, and poor storage and computing processing speed, which urgently need to be solved. Summary of the Invention
[0016] The present application provides a storage and computing processing method and device based on a capacitor-free DRAM memory to solve the problems of high process cost, incompatible preparation process, high circuit complexity, and poor storage and computing processing speed in the existing DRAM in-memory processing architecture.
[0017] The first aspect of the present application provides a storage and computing processing method based on a capacitor-less DRAM memory, including the following steps: preparing corresponding write transistors and read transistors using preset oxide semiconductors and silicon transistors; vertically stacking the write transistors on the read transistors based on a preset heterogeneous integration strategy to establish a capacitor-less DRAM memory unit; establishing a heterogeneous integrated capacitor-less DRAM memory macro based on a pre-built memory unit peripheral circuit and the capacitor-less DRAM memory unit, and constructing a corresponding in-memory processing architecture based on the heterogeneous integrated capacitor-less DRAM memory macro and the pre-built convolution kernel logic circuit to perform corresponding storage and computing operations through the in-memory processing architecture.
[0018] Optionally, in one embodiment of the present application, the preparation of corresponding write transistors and read transistors using preset oxide semiconductors and silicon transistors includes: determining an oxide semiconductor that meets preset ultra-low leakage requirements, and constructing the read transistor based on a preset CMOS process and the silicon transistor.
[0019] Optionally, in one embodiment of the present application, the pre-built memory cell peripheral circuit and the capacitor-free DRAM memory cell are used to establish a heterogeneous integrated capacitor-free DRAM memory macro, including: based on a preset sensitive amplifier SA, latch Latch, selector MUX, array row driver, array column driver, register, address decoder and clock controller, and combined with the CMOS process, constructing the memory cell peripheral circuit; constructing the heterogeneous integrated capacitor-free DRAM memory macro according to the memory cell peripheral circuit and the capacitor-free DRAM memory cell.
[0020] Optionally, in one embodiment of the present application, a corresponding in-memory processing architecture is constructed based on the heterogeneous integrated capacitor-less DRAM memory macro and the pre-built convolution kernel logic circuit to perform corresponding storage operations through the in-memory processing architecture, including: constructing the convolution kernel logic circuit based on the preset multiplication and accumulation logic circuit, selector, pooling circuit, activation circuit and the CMOS process; and establishing the in-memory processing architecture using the heterogeneous integrated capacitor-less DRAM memory macro and the convolution kernel logic circuit.
[0021] The second aspect of the present application provides a storage and computing processing device based on a capacitor-less DRAM memory, including: a preparation module for preparing corresponding write transistors and read transistors using preset oxide semiconductors and silicon transistors; a heterogeneous integration module for vertically stacking the write transistors on the read transistors based on a preset heterogeneous integration strategy to establish a capacitor-less DRAM memory unit; an in-memory processing module for establishing a heterogeneously integrated capacitor-less DRAM memory macro based on a pre-built storage unit peripheral circuit and the capacitor-less DRAM memory unit, and constructing a corresponding in-memory processing architecture based on the heterogeneously integrated capacitor-less DRAM memory macro and the pre-built convolution kernel logic circuit to perform corresponding storage and computing operations through the in-memory processing architecture.
[0022] Optionally, in one embodiment of the present application, the preparation module includes: a determination unit, configured to determine an oxide semiconductor that meets a preset ultra-low leakage requirement, and construct the read transistor based on a preset CMOS process and the silicon transistor.
[0023] Optionally, in one embodiment of the present application, the in-memory processing module includes: a first construction unit, used to construct the memory cell peripheral circuit based on a preset sensitive amplifier SA, latch Latch, selector MUX, array row driver, array column driver, register, address decoder and clock controller, and in combination with the CMOS process; a second construction unit, used to construct the heterogeneous integrated capacitor-less DRAM memory macro based on the memory cell peripheral circuit and the capacitor-less DRAM memory cell.
[0024] Optionally, in one embodiment of the present application, the in-memory processing module further includes: a third construction unit, used to construct the convolution kernel logic circuit based on a preset multiplication and accumulation logic circuit, a selector, a pooling circuit, an activation circuit and the CMOS process; and an establishment unit, used to establish the in-memory processing architecture using the heterogeneous integrated capacitor-less DRAM memory macro and the convolution kernel logic circuit.
[0025] The third aspect of the present application provides a storage and computing processing architecture for implementing the above-mentioned storage and computing processing method based on capacitor-less DRAM memory.
[0026] The fourth aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the storage and computing method based on the capacitor-less DRAM memory as described in the above embodiment.
[0027] The fifth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above-mentioned storage and calculation processing method based on capacitor-less DRAM memory.
[0028] Therefore, the embodiments of the present application have the following beneficial effects:
[0029] The embodiments of the present application can prepare corresponding write transistors and read transistors by utilizing preset oxide semiconductors and silicon transistors; based on a preset heterogeneous integration strategy, the write transistor is vertically stacked on the read transistor to establish a capacitor-free DRAM memory unit; based on the pre-built memory unit peripheral circuit and the capacitor-free DRAM memory unit, a heterogeneous integrated capacitor-free DRAM memory macro is established, and a corresponding in-memory processing architecture is constructed according to the heterogeneous integrated capacitor-free DRAM memory macro and the pre-built convolution kernel logic circuit to perform corresponding storage and computing operations through the in-memory processing architecture. The present application can avoid the special process requirements of traditional DRAM for low-leakage Si transistors, significantly reduce the preparation cost, and improve the in-memory processing performance and reliability. As a result, the problems of high process cost, incompatible preparation process, high circuit complexity, and poor storage and computing processing speed of the prior art DRAM in-memory processing architecture are solved.
[0030] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0032] Figure 1 A flowchart of a storage and calculation processing method based on a capacitor-less DRAM memory provided according to an embodiment of the present application;
[0033] Figure 2 A schematic diagram of a capacitor-less DRAM memory cell structure provided in one embodiment of the present application;
[0034] Figure 3 A schematic diagram of a heterogeneously integrated capacitor-less DRAM memory provided in one embodiment of the present application;
[0035] Figure 4 A schematic diagram of an overall in-memory processing architecture based on heterogeneous integrated capacitor-less DRAM memory provided in one embodiment of the present application;
[0036] Figure 5 1 is an exemplary diagram of a storage and computing device based on a capacitor-less DRAM memory according to an embodiment of the present application;
[0037] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0038] Among them, 10-storage and computing processing device based on capacitor-free DRAM memory; 100-preparation module, heterogeneous integration-computing module, 300-in-memory processing module; 601-memory, 602-processor, 603-communication interface. DETAILED DESCRIPTION
[0039] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0040] The following describes the storage and calculation processing method and device based on the capacitor-free DRAM memory of the embodiment of the present application with reference to the accompanying drawings. In response to the problems mentioned in the above background technology, the present application provides a storage and calculation processing method based on the capacitor-free DRAM memory, in which the corresponding write transistor and read transistor are prepared by using a preset oxide semiconductor and silicon transistor; based on a preset heterogeneous integration strategy, the write transistor is vertically stacked on the read transistor to establish a capacitor-free DRAM memory unit; based on the pre-built memory unit peripheral circuit and the capacitor-free DRAM memory unit, a heterogeneous integrated capacitor-free DRAM memory macro is established, and a corresponding in-memory processing architecture is constructed based on the heterogeneous integrated capacitor-free DRAM memory macro and the pre-built convolution kernel logic circuit to perform the corresponding storage and calculation operations through the in-memory processing architecture. The present application can avoid the special process requirements of traditional DRAM for low-leakage Si transistors, significantly reduce the preparation cost, and improve the in-memory processing performance and reliability. As a result, the problems of high process cost, incompatible preparation process, high circuit complexity, and poor storage and calculation processing speed of the existing DRAM in-memory processing architecture are solved.
[0041] Specifically, Figure 1 A flowchart of a storage and calculation processing method based on a capacitor-less DRAM memory provided in an embodiment of the present application.
[0042] like Figure 1 As shown, the storage and calculation processing method based on the capacitor-free DRAM memory includes the following steps:
[0043] In step S101 , corresponding write transistors and read transistors are prepared using preset oxide semiconductors and silicon transistors.
[0044] In step S102 , based on a predetermined heterogeneous integration strategy, a write transistor is vertically stacked on a read transistor to create a capacitor-less DRAM memory cell.
[0045] Figure 2 This is a schematic diagram of the structure of a capacitor-free DRAM memory unit. Figure 2 As shown, the embodiment of the present application can use OS (such as IGZO) as a write transistor and a silicon transistor as a read transistor; Figure 3 Schematic diagram of heterogeneous integrated capacitor-free DRAM memory. Figure 3 As shown, embodiments of the present application can utilize heterogeneous integration technology to vertically stack an IGZO write transistor on a Si read transistor, directly fabricating a back-end compatible oxide semiconductor channel transistor (such as an IGZO transistor) on the Si-based read transistor. The IGZO write transistor improves retention time, while the Si read transistor balances speed. Furthermore, Si transistors only need to focus on high speed, not low leakage, making them compatible with normal SiCMOS logic processes.
[0046] It can be understood that the write tube in the embodiment of the present application can use OS transistors to achieve low leakage, and the read tube can use Si transistors to achieve high-speed reading; Si transistors have higher mobility than IGZO transistors and are suitable for higher-speed reading operations; at the same time, since the process of Si transistors themselves is compatible with the logic process, the peripheral circuit of the storage array can be prepared at the same time as the read tube, and the calculation circuit can be prepared adjacently, so that calculations can be performed directly after the storage Macro is read out, saving the cost and power consumption of data movement; therefore, the embodiment of the present application can realize an in-memory processing circuit based on a capacitor-free DRAM memory using OS write transistors and Si read transistors.
[0047] Therefore, the embodiment of the present application integrates the IGZO layer and the Si layer in 3D, which significantly improves the storage density. Each storage unit only occupies the area of one transistor and is compatible with existing CMOS design rules.
[0048] Optionally, in one embodiment of the present application, corresponding write transistors and read transistors are prepared using preset oxide semiconductors and silicon transistors, including: determining an oxide semiconductor that meets preset ultra-low leakage requirements, and constructing a read transistor based on a preset CMOS process and silicon transistors.
[0049] It should be noted that, in the actual implementation process, the embodiments of the present application can separate the process requirements of storage and logic units through heterogeneous integration technology (IGZO write transistors stacked vertically on Si read transistors). That is to say, the embodiments of the present application can utilize the ultra-low leakage characteristics of IGZO transistors (<1e-20A / μm) to meet the long retention time of storage cells without the need for the refresh circuit of traditional DRAM; in addition, the embodiments of the present application can use standard CMOS logic processes to manufacture high-speed read transistors, peripheral circuits and computing circuits, avoiding the special process requirements of traditional DRAM for low-leakage Si transistors, and significantly reducing preparation costs.
[0050] It should be noted that the channel material used in the embodiment of the present application is IGZO channel, which has the advantage of low leakage. During the specific implementation process, those skilled in the art can select a variety of N-type oxide semiconductor channel materials with low leakage characteristics such as IWO, IZO, ZnO, ITO, or two-dimensional materials with low leakage characteristics such as MoS2 as the corresponding channel materials according to actual conditions. No specific limitation is made here.
[0051] Therefore, the embodiment of the present application can utilize the ultra-low off-state current of the IGZO write tube to reduce static power consumption, thereby effectively improving overall energy efficiency.
[0052] In step S103, a heterogeneous integrated capacitor-less DRAM memory macro is established based on the pre-built storage unit peripheral circuit and the capacitor-less DRAM memory unit, and a corresponding in-memory processing architecture is constructed according to the heterogeneous integrated capacitor-less DRAM memory macro and the pre-built convolution kernel logic circuit to perform corresponding storage and computing operations through the in-memory processing architecture.
[0053] Furthermore, the embodiments of the present application can directly prepare computing circuits on the periphery of Si CMOS, that is, the existing Si CMOS logic circuit design can be directly used to complete the calculation after data readout, and an in-memory processing architecture is adopted without the need to introduce independent analog computing circuits (such as DAC / ADC of ReRAM), thereby making storage reliability higher, reducing area overhead, and the circuit is relatively simple without bringing additional complexity.
[0054] It can be understood that the embodiments of the present application eliminate the destructive reading problem of 1T1C DRAM and improve data reliability by separating the read and write paths (IGZO write transistor and Si read transistor are independently controlled); in addition, in the embodiments of the present application, the Si read transistor is directly integrated with high-speed CMOS logic circuits (such as adders and shift units), and the in-memory processing delay is significantly reduced compared to traditional peripheral calculations.
[0055] It should be noted that the 2T capacitor-free DRAM memory cell constructed in the embodiment of the present application can also be replaced by 3T or 4T according to actual conditions by those skilled in the art, and no specific limitation is made here.
[0056] In addition, the embodiments of the present application can support hybrid integration with advanced process logic circuits (such as FinFET, GAAFET, CFET and other future advanced process nodes), and are suitable for scenarios such as AI accelerators and high-density edge computing chips.
[0057] Optionally, in one embodiment of the present application, a heterogeneous integrated capacitor-less DRAM memory macro is established based on a pre-built memory cell peripheral circuit and a capacitor-less DRAM memory cell, including: building a memory cell peripheral circuit based on a preset sensitive amplifier SA, a latch Latch, a selector MUX, an array row driver, an array column driver, a register, an address decoder and a clock controller, and combining the CMOS process; and building a heterogeneous integrated capacitor-less DRAM memory macro based on the memory cell peripheral circuit and the capacitor-less DRAM memory cell.
[0058] It should be noted that Figure 4 The figure is a schematic diagram of the overall architecture of in-memory processing based on heterogeneous integrated capacitor-free DRAM memory. Figure 4 As shown in the figure, the left half is a heterogeneous integrated capless DRAM memory macro, which includes an IGZO write / Si read capless DRAM memory cell array and its peripheral circuits. The peripheral circuits include sense amplifiers (SA), latches (Latch), selectors (MUX), array row drivers, array column drivers, registers, address decoders, and clock controllers, realizing the complete memory macro array functionality.
[0059] In the embodiment of the present application, the number of data bits stored in the storage unit is 0 or 1, and digital calculations are mainly performed after reading out; but if the storage unit can store multiple bits, in theory, it can also perform analog calculations, and the peripheral circuit design needs to include ADC, DAC and other digital-to-analog conversion modules, which are also compatible with the storage and computing processing architecture of the embodiment of the present application.
[0060] Optionally, in one embodiment of the present application, a corresponding in-memory processing architecture is constructed based on a heterogeneous integrated capacitor-less DRAM memory macro and a pre-built convolution kernel logic circuit to perform corresponding storage operations through the in-memory processing architecture, including: constructing a convolution kernel logic circuit based on a preset multiplication and accumulation logic circuit, a selector, a pooling circuit, an activation circuit and a CMOS process; and establishing an in-memory processing architecture using a heterogeneous integrated capacitor-less DRAM memory macro and a convolution kernel logic circuit.
[0061] like Figure 4As shown in the figure, the right half contains the convolution kernel logic circuitry, primarily including the multiply-accumulate logic circuitry, selectors, pooling circuitry, activation circuitry, and all other logic required for convolution calculations. The read transistors, memory peripheral circuitry, and convolution logic circuitry are all fabricated using a complete SiCMOS process, enabling tight integration of memory and logic circuitry, thus realizing an in-memory processing architecture based on heterogeneously integrated capacitor-free DRAM memory.
[0062] During the actual execution process, the embodiment of the present application mainly performs matrix-vector multiplication operations for CNN neural network applications after reading out the data; and in the face of different network requirements, such as large language models LLM and diffusion models DM, those skilled in the art can also design logic circuits according to different deep neural network computing requirements, which can also be carried out under the architectural framework of the embodiment of the present application.
[0063] According to the storage-computing processing method based on the capacitor-free DRAM memory proposed in the embodiment of the present application, the corresponding write transistor and read transistor are prepared by using a preset oxide semiconductor and silicon transistor; based on a preset heterogeneous integration strategy, the write transistor is vertically stacked on the read transistor to establish a capacitor-free DRAM memory cell; based on the pre-built memory cell peripheral circuit and the capacitor-free DRAM memory cell, a heterogeneous integrated capacitor-free DRAM memory macro is established, and a corresponding in-memory processing architecture is constructed based on the heterogeneous integrated capacitor-free DRAM memory macro and the pre-built convolution kernel logic circuit to perform corresponding storage-computing operations through the in-memory processing architecture. This application can avoid the special process requirements of traditional DRAM for low-leakage Si transistors, significantly reduce preparation costs, and improve in-memory processing performance and reliability.
[0064] Secondly, the storage and computing processing device based on the capacitor-less DRAM memory proposed in accordance with the embodiment of the present application is described with reference to the accompanying drawings.
[0065] Figure 5 It is a block diagram of a storage and computing processing device based on a capacitor-less DRAM memory according to an embodiment of the present application.
[0066] like Figure 5 As shown, the storage and computing processing device 10 based on the capacitor-less DRAM memory includes: a preparation module 100, a heterogeneous integration module 200 and an in-memory processing module 300.
[0067] The preparation module 100 is used to prepare corresponding write transistors and read transistors using preset oxide semiconductors and silicon transistors.
[0068] The heterogeneous integration module 200 is used to vertically stack a write transistor on a read transistor based on a preset heterogeneous integration strategy to establish a capacitor-less DRAM memory cell.
[0069] The in-memory processing module 300 is used to establish a heterogeneous integrated capacitor-less DRAM memory macro based on pre-built storage unit peripheral circuits and capacitor-less DRAM memory cells, and to construct a corresponding in-memory processing architecture based on the heterogeneous integrated capacitor-less DRAM memory macro and the pre-built convolution kernel logic circuit to perform corresponding storage and computing operations through the in-memory processing architecture.
[0070] Optionally, in one embodiment of the present application, the preparation module 100 includes: a determination unit, configured to determine an oxide semiconductor that meets a preset ultra-low leakage requirement, and construct a read transistor based on a preset CMOS process and a silicon transistor.
[0071] Optionally, in one embodiment of the present application, the in-memory processing module 300 includes: a first construction unit and a second construction unit.
[0072] Among them, the first construction unit is used to construct the storage unit peripheral circuit based on the preset sense amplifier SA, latch Latch, selector MUX, array row driver, array column driver, register, address decoder and clock controller, and combined with CMOS technology.
[0073] The second construction unit is used to construct a heterogeneous integrated capacitor-less DRAM memory macro according to the memory cell peripheral circuit and the capacitor-less DRAM memory cell.
[0074] Optionally, in one embodiment of the present application, the in-memory processing module 300 further includes: a third construction unit and an establishment unit.
[0075] Among them, the third construction unit is used to construct a convolution kernel logic circuit based on a preset multiplication and accumulation logic circuit, selector, pooling circuit, activation circuit and CMOS process.
[0076] Establish a unit for building an in-memory processing architecture using heterogeneously integrated capacitor-less DRAM memory macros and convolution kernel logic circuits.
[0077] It should be noted that the above explanation of the embodiment of the storage and calculation processing method based on the capacitor-less DRAM memory is also applicable to the storage and calculation processing device based on the capacitor-less DRAM memory of this embodiment, and will not be repeated here.
[0078] According to the embodiment of the present application, the storage and computing processing device based on the capacitor-free DRAM memory proposed includes a preparation module 100 for preparing corresponding write transistors and read transistors using preset oxide semiconductors and silicon transistors; a heterogeneous integration module 200 for vertically stacking the write transistors on the read transistors based on a preset heterogeneous integration strategy to establish a capacitor-free DRAM memory unit; and an in-memory processing module 300 for establishing a heterogeneous integrated capacitor-free DRAM memory macro based on pre-built memory unit peripheral circuits and capacitor-free DRAM memory units, and constructing a corresponding in-memory processing architecture based on the heterogeneous integrated capacitor-free DRAM memory macro and the pre-built convolution kernel logic circuit to perform corresponding storage and computing operations through the in-memory processing architecture. This application can avoid the special process requirements of traditional DRAM for low-leakage Si transistors, significantly reduce preparation costs, and improve in-memory processing performance and reliability.
[0079] The embodiment of the present application also provides a storage and computing processing architecture for implementing the above-mentioned storage and computing processing method based on capacitor-less DRAM memory.
[0080] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0081] A memory 601 , a processor 602 , and a computer program stored in the memory 601 and executable on the processor 602 .
[0082] When the processor 602 executes the program, the storage and calculation processing method based on the capacitor-less DRAM memory provided in the above embodiment is implemented.
[0083] Furthermore, the electronic device further includes:
[0084] The communication interface 603 is used for communication between the memory 601 and the processor 602 .
[0085] The memory 601 is used to store computer programs that can be run on the processor 602 .
[0086] The memory 601 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0087] If the memory 601, processor 602, and communication interface 603 are implemented independently, the communication interface 603, memory 601, and processor 602 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0088] Optionally, in a specific implementation, if the memory 601, the processor 602 and the communication interface 603 are integrated on a chip, the memory 601, the processor 602 and the communication interface 603 can communicate with each other through an internal interface.
[0089] The processor 602 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0090] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned storage and calculation processing method based on capacitor-less DRAM memory.
[0091] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0092] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0093] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0094] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or N wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0095] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0096] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0097] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0098] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A storage and calculation processing method based on capacitor-less DRAM memory, characterized in that: The following steps are involved: Using the preset oxide semiconductor and silicon transistor to prepare the corresponding write transistor and read transistor; Based on a preset heterogeneous integration strategy, vertically stacking the write transistor on the read transistor to establish a capacitor-less DRAM memory cell; Based on the pre-built storage unit peripheral circuit and the capacitor-less DRAM memory unit, a heterogeneous integrated capacitor-less DRAM memory macro is established, and a corresponding in-memory processing architecture is constructed according to the heterogeneous integrated capacitor-less DRAM memory macro and the pre-built convolution kernel logic circuit to perform corresponding storage and computing operations through the in-memory processing architecture.
2. The method according to claim 1, characterized in that The method of using a preset oxide semiconductor and a silicon transistor to prepare a corresponding write transistor and a read transistor includes: An oxide semiconductor that meets a preset ultra-low leakage requirement is determined, and the read transistor is constructed based on a preset CMOS process and the silicon transistor.
3. The method according to claim 2, characterized in that The method of establishing a heterogeneous integrated capacitor-less DRAM memory macro based on the pre-built memory cell peripheral circuit and the capacitor-less DRAM memory cell comprises: Based on the preset sense amplifier SA, latch Latch, selector MUX, array row driver, array column driver, register, address decoder and clock controller, and in combination with the CMOS process, the memory cell peripheral circuit is constructed; The heterogeneous integrated capacitor-less DRAM memory macro is constructed based on the memory cell peripheral circuit and the capacitor-less DRAM memory cell.
4. The method according to claim 3, characterized in that The method includes constructing a corresponding in-memory processing architecture based on the heterogeneous integrated capacitor-less DRAM memory macro and the pre-built convolution kernel logic circuit to perform corresponding storage and computing operations through the in-memory processing architecture, including: Constructing the convolution kernel logic circuit based on the preset multiplication-accumulation logic circuit, selector, pooling circuit, activation circuit and the CMOS process; The in-memory processing architecture is established using the heterogeneous integrated capacitor-less DRAM memory macro and the convolution kernel logic circuit.
5. A storage and calculation processing device based on capacitor-less DRAM memory, characterized in that: include: A preparation module, used to prepare corresponding write transistors and read transistors using preset oxide semiconductors and silicon transistors; a heterogeneous integration module for vertically stacking the write transistor on the read transistor based on a preset heterogeneous integration strategy to establish a capacitor-less DRAM memory cell; An in-memory processing module is used to establish a heterogeneous integrated capacitor-less DRAM memory macro based on a pre-built storage unit peripheral circuit and the capacitor-less DRAM memory unit, and to construct a corresponding in-memory processing architecture based on the heterogeneous integrated capacitor-less DRAM memory macro and the pre-built convolution kernel logic circuit, so as to perform corresponding storage and computing operations through the in-memory processing architecture.
6. The device according to claim 5, characterized in that The preparation module includes: The determining unit is configured to determine an oxide semiconductor that meets a preset ultra-low leakage requirement, and construct the read transistor based on a preset CMOS process and the silicon transistor.
7. The device according to claim 6, characterized in that The in-memory processing module includes: A first construction unit is configured to construct the memory cell peripheral circuit based on a preset sense amplifier SA, latch Latch, selector MUX, array row driver, array column driver, register, address decoder, and clock controller, in combination with the CMOS process; A second construction unit is configured to construct the heterogeneous integrated capacitor-less DRAM memory macro according to the memory cell peripheral circuit and the capacitor-less DRAM memory cell.
8. A storage and computing processing architecture, characterized in that: Used to implement the storage and calculation processing method based on the capacitor-free DRAM memory as described in any one of claims 1-4.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the storage and calculation processing method based on the capacitor-less DRAM memory as described in any one of claims 1 to 4.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the storage and calculation processing method based on the capacitor-less DRAM memory as described in any one of claims 1 to 4.
Citation Information
Cited By
Multi-model fusion memory macro-cell design quality prediction system and method
CN121659892A