Monolithic three-dimensional integration implementation method and device for hybrid precision in-memory computing architecture

By using silicon-based CMOS process and back-channel integrated process in the in-memory computing architecture to achieve vertical stacking multi-layer functional layers, the problems of large chip area and low computing efficiency in the in-memory computing architecture are solved, and efficient computing and high parallelism are achieved.

CN120086485APending Publication Date: 2025-06-03TSINGHUA UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411941400.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In-memory computing architectures usually require a large portion of the chip area and have low computing efficiency.

Method used

The control logic circuit is manufactured using silicon-based CMOS process, and multi-layer functional layers are stacked vertically on the logic circuit through the rear-channel integrated process. Communication between multi-layer functional layers is achieved through interlayer dielectric vias, integrating analog in-memory arrays, digital in-memory arrays and caches.

Benefits of technology

It improves the computing efficiency of the chip, reduces the chip area, enhances the computing parallelism, and achieves high communication bandwidth through efficient on-chip interconnection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086485A_ABST
    Figure CN120086485A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of monolithic integration, in particular to a monolithic three-dimensional integration implementation method and device for a hybrid precision in-memory computing architecture, and the method comprises the steps: obtaining a silicon-based CMOS technology and a subsequent integration technology in a monolithic three-dimensional integration technology; manufacturing a control logic circuit on the computing architecture in the mixed precision memory by adopting a silicon-based CMOS (Complementary Metal Oxide Semiconductor) process; a plurality of functional layers are vertically stacked on a logic circuit by adopting a back-end integration process, communication among the functional layers is realized through interlayer dielectric via holes in the vertical direction, an analog in-memory array, a digital in-memory array and a cache are arranged on the functional layers, the analog in-memory array realizes matrix and vector multiplication with fixed weight, and the digital in-memory array realizes matrix and vector multiplication with fixed weight. And the array in the digital memory realizes matrix multiplication of refresh weight. Therefore, the problems that an in-memory computing architecture usually needs to occupy a relatively large part of chip area, the computing efficiency is relatively low and the like in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of monolithic integration technology, and particularly relates to a method and device for realizing monolithic three-dimensional integration of a mixed-precision in-memory computing architecture. Background Art

[0002] In-memory computing technology adopts a non-Von Neumann computing architecture and can efficiently implement the core operation in neural networks: matrix-vector multiplication. Taking the ViT (Vision Transformer) network as an example, the core operation of the ViT network includes the multiplication operations of three matrices Q, K, and V. These three matrices are the matrices after linear mapping of the input image through the convolutional layer and are related to the input image. Simply put, QKV are all non-fixed matrices, and the elements in the matrices will change according to the input content. When performing matrix multiplication, the matrices need to be refreshed continuously.

[0003] The RRAM-based in-memory computing technology is suitable for matrix operations with fixed weights. Because the principle of RRAM in-memory computing is to map the weight matrix to the conductance value of RRAM, and since the time required to regulate the RRAM conductance value is much greater than the computing time, the RRAM in-memory computing technology is suitable for ordinary convolutional layer operations. The SRAM-based in-memory computing technology is more suitable for performing Attention operations because the SRAM write speed is very fast and the precision of digital operations is high. The disadvantage is that SRAM occupies a relatively large part of the chip area compared to RRAM.

[0004] In the neural network operations of the traditional two-dimensional architecture, the main bottleneck is the transfer of the weight matrix between the storage module and the computing module. The use of in-memory computing eliminates this part of data transfer, resulting in the data transfer efficiency between the input / output and different computing modules becoming the decisive factor for computing efficiency. Summary of the Invention

[0005] This application provides a method and device for realizing monolithic three-dimensional integration of a mixed-precision in-memory computing architecture to solve the problems in related technologies that the in-memory computing architecture usually occupies a relatively large part of the chip area and has low computing efficiency, etc.

[0006] The first aspect of the present application provides a monolithic three-dimensional integration implementation method for a mixed-precision in-memory computing architecture, including the following steps: obtaining the silicon-based CMOS process and the back-end integration process in the monolithic three-dimensional integration technology; manufacturing a control logic circuit on the mixed-precision in-memory computing architecture using the silicon-based CMOS process; vertically stacking multiple functional layers on the fabricated logic circuit using the back-end integration process, and realizing communication between the multiple functional layers through interlayer dielectric vias in the vertical direction. Among them, an analog in-memory array, a digital in-memory array, and a cache are provided on the multiple functional layers. The analog in-memory array realizes the matrix-vector multiplication operation with fixed weights, and the digital in-memory array realizes the matrix multiplication operation for refreshing weights.

[0007] Optionally, vertically stacking multiple functional layers on the control logic circuit of the mixed-precision in-memory computing architecture using the back-end integration process includes: obtaining the target process step sequence of the back-end integration process; determining the target stacking sequence of the multiple functional layers according to the target process step sequence; vertically stacking multiple functional layers on the control logic circuit through the back-end process and the target stacking sequence.

[0008] Optionally, the multiple functional layers include a functional layer for analog in-memory, a functional layer for digital in-memory, and a functional layer for cache. Among them, an analog in-memory array is provided on the functional layer for analog in-memory, a digital in-memory array is provided on the functional layer for digital in-memory, and a cache is provided on the functional layer for cache.

[0009] Optionally, the functional layer for digital in-memory and the functional layer for cache are integrated on the same functional layer, or the functional layer for digital in-memory and the functional layer for cache are independent functional layers.

[0010] Optionally, the analog in-memory array and the digital in-memory array adopt one or more of hafnium oxide-based resistive random access memory, phase change memory, magnetic memory, ferroelectric memory, and conductive bridge memory.

[0011] Optionally, the control logic circuit adopts any one of the back-end CFET structures of carbon nanotubes and IGZO, the back-end CFET structure based on tungsten diselenide and molybdenum disulfide, and the back-end CFET structure based on P-type silicon and N-type silicon.

[0012] Optionally, the number of layers of the respective functional layers of the analog in-memory array and the digital in-memory array is configured according to the algorithm requirements of the mixed-precision in-memory computing architecture.

[0013] In the second aspect of the present application, an embodiment provides a monolithic three-dimensional integrated implementation device for a mixed-precision in-memory computing architecture, including: an acquisition module for acquiring the silicon-based CMOS process and the back-end integration process in monolithic three-dimensional integration technology; a manufacturing module for manufacturing control logic circuits on the mixed-precision in-memory computing architecture using the silicon-based CMOS process; a stacking module for vertically stacking multiple functional layers on the manufactured logic circuits using the back-end integration process, and realizing communication between the multiple functional layers through interlayer dielectric vias in the vertical direction. Among them, an analog in-memory array, a digital in-memory array, and a cache are provided on the multiple functional layers. The analog in-memory array realizes matrix-vector multiplication operations with fixed weights, and the digital in-memory array realizes matrix multiplication operations for refreshing weights.

[0014] In the third aspect of the present application, an embodiment provides a mixed-precision in-memory computing architecture, including: control logic circuits; multiple functional layers vertically stacked on the manufactured logic circuits, where the multiple functional layers communicate through interlayer dielectric vias in the vertical direction. An analog in-memory array, a digital in-memory array, and a cache are provided on the multiple functional layers. The analog in-memory array realizes matrix-vector multiplication operations with fixed weights, and the digital in-memory array realizes matrix multiplication operations for refreshing weights.

[0015] In the fourth aspect of the present application, an embodiment provides an electronic device, characterized by including the mixed-precision in-memory computing architecture of the first aspect.

[0016] Therefore, the present application has the following beneficial effects:

[0017] In the embodiments of the present application, the control logic circuits are manufactured on the mixed-precision in-memory computing architecture using the silicon-based CMOS process, and multiple functional layers are vertically stacked on the manufactured logic circuits using the back-end integration process. Communication between the multiple functional layers is realized through interlayer dielectric vias in the vertical direction, enabling data communication between RRAM, SRAM, and the silicon-based control module to be carried out through efficient on-chip interconnection, having a very high communication bandwidth, improving the computing efficiency of the chip, and at the same time reducing the chip area and increasing the computing parallelism. Thus, the problems in the related art that the in-memory computing architecture usually occupies a large part of the chip area and has low computing efficiency are solved.

[0018] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. Description of the Drawings

[0019] The above-mentioned and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0020] Figure 1 It is the structural diagram of the ViT model;

[0021] Figure 2 Flow chart of a monolithic three-dimensional integration implementation method for a mixed-precision in-memory computing architecture provided according to an embodiment of the present application;

[0022] Figure 3 Schematic diagram of a mixed-precision in-memory computing architecture and array optical micrograph provided according to an embodiment of the present application;

[0023] Figure 4 Transmission electron micrograph of a chip sample provided according to an embodiment of the present application;

[0024] Figure 5 ViT algorithm mapping of a mixed-precision in-memory computing architecture provided according to an embodiment of the present application;

[0025] Figure 6 Comparison chart of inference accuracy rates provided according to an embodiment of the present application;

[0026] Figure 7 Comparison chart of energy efficiency evaluation provided according to an embodiment of the present application;

[0027] Figure 8 Comparison chart of operation throughput evaluation provided according to an embodiment of the present application;

[0028] Figure 9 Block diagram of a monolithic three-dimensional integration implementation device for a mixed-precision in-memory computing architecture provided according to an embodiment of the present application;

[0029] Figure 10 Block diagram of a mixed-precision in-memory computing architecture provided according to an embodiment of the present application;

[0030] Figure 11 Schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. Detailed implementation manners

[0031] The embodiments of the present application are described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described by referring to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.

[0032] Traditional computer vision tasks usually use CNN (Convolutional Neural Network) to extract image features, while the ViT (Vision Transformer) model uses the Transformer model to process computer vision tasks and captures long-range dependencies in images through a global attention mechanism. Traditional Transformer models have achieved great success in natural language processing. However, image processing tasks are different from natural language models. The structure and features of image data determine that the ViT model needs to divide the image data into blocks (image patches) and then input the image blocks as a sequence into the Transformer model. As Figure 1 shown, the ViT model is mainly divided into the following three parts:

[0033] 1. Input Encoding: After the input image is divided into a series of image blocks, each image block is linearly mapped through a convolutional layer and represented as a vector sequence.

[0034] 2. Encoder: Use multiple layers of Transformer encoders to process the input image block vector sequence. Each encoder layer consists of a multi-head self-attention mechanism and a feed-forward neural network. The core part is the multi-head self-attention mechanism, which can capture the correlations between image blocks and perform context-aware feature representation.

[0035] 3. Classification Head: Perform classification prediction through an additional linear output layer, synthesize the features of the image block sequence into a global feature vector, and classify through a linear layer.

[0036] The core operation of the ViT model lies in the multi-head self-attention mechanism. The self-attention operation involves matrix multiplication operations, and the main calculation steps are as follows:

[0037] 1. Multiply the input vector matrix by three weight matrices respectively to obtain three matrices: Query, Key, and Value.

[0038] 2. Calculate the attention scores, that is, calculate the product of the Query matrix and the transposed matrix of the Key matrix.

[0039] 3. Divide the calculation result by d k which is the dimension of Key, and perform a Softmax operation on the result for normalization.

[0040] 4. Multiply the calculation result after the Softmax operation by the Value matrix.

[0041] 5. Sum the corresponding weighted Value matrices after multiplication and output after linear transformation.

[0042] (2) In-memory computing technology based on RRAM (Resistive Random Access Memory) arrays

[0043] In-memory computing is a new architecture that integrates computing functions into storage units at the device level, which can fundamentally reduce the repeated transfer of a large amount of data between the data storage and computing modules during the computing process. In-memory computing technology generally calculates the product of a vector and a matrix. The matrix is stored in the memory, and the product result can be output when the vector is input to the memory. The generally recognized advantage of in-memory computing technology is low power consumption.

[0044] According to the different types of stored weight values, in-memory computing architectures can be divided into analog in-memory computing and digital in-memory computing. Analog in-memory computing means that the information stored in each memory-computation unit is a multi-bit value (such as 0, 0.125, 0.25,..., 1), which can only be implemented by non-volatile memories with multi-bit storage functions. The advantages are high weight storage density, and multiple bits of information can be stored in a single device, thus reducing the area and power consumption overhead of a single memory-computation array; but the disadvantage is that it is difficult for current new non-volatile memories to store correct information for a long time, so the calculation accuracy rate will decrease. Digital in-memory computing means that the information stored in each memory-computation unit is discrete digital information (i.e., 0 or 1), which is reflected as a high resistance state / low resistance state in a resistor and as a low level / high level in SRAM (Static Random Access Memory). Compared with the analog in-memory computing architecture, the information storage in the memory-computation units in the digital in-memory computing array is more reliable and the calculation accuracy rate is higher. The disadvantage is that the storage density becomes smaller. Each memory-computation unit in the digital in-memory computing array can only store 1 bit of information, so the array area and power consumption required to complete the same computing task are larger.

[0045] A resistive random access memory can be regarded as a resistor with variable conductance. If we map a number to the conductance of the resistive random access memory and another number to the input voltage of the resistive random access memory, we can rely on Ohm's law to obtain that the output current is the product of the input voltage and the conductance of the resistive random access memory. Resistive random access memories have the advantages of simple structure, CMOS process compatibility, high integration density, etc. At the same time, they have the advantages of low power consumption and multi-bit storage in terms of performance. A typical RRAM consists of upper and lower electrodes and an oxide as the resistive change layer in the middle. Under the action of an externally applied voltage pulse, the movement of oxygen ions in the oxide forms a conductive filament based on oxygen vacancies to connect the upper and lower electrodes. By regulating the morphology of the conductive filament, the conductance of the memristor can be continuously adjusted, showing an analog resistive change characteristic. Currently, it has been used to manufacture large-scale storage / memory-computation integrated arrays.

[0046] RRAM devices with different resistive switching layers have different conductance state characteristics. Therefore, devices made of different materials can be selected according to application requirements. For example, RRAM with HfAlO x as the resistive switching layer has good continuous change characteristics of the conductance state and is suitable for multi-bit storage functions; while RRAM with Ta 2 O 5 as the resistive switching layer has a higher on / off ratio (i.e., the conductance ratio between the high-resistance state and the low-resistance state). Therefore, it is more suitable for digital storage functions.

[0047] (3) In-memory computing technology based on the back-end SRAM memory array

[0048] SRAM is a type of random access memory. The so-called "static" means that as long as this memory remains powered on, the data stored in it can be constantly maintained and does not require a refresh circuit to lock the data stored inside. In contrast, DRAM (Dynamic Random Access Memory) needs to be updated periodically and refreshed and charged every once in a while, otherwise the data stored inside will disappear.

[0049] An SRAM cell usually consists of 4 - 6 transistors. New functional SRAMs also have 8 - 10 transistor structures. Taking the most commonly used 6-transistor SRAM (6T SRAM, where T is the abbreviation for transistor) cell as an example, each storage cell in SRAM can store one bit of data. The storage structure is composed of a flip-flop, and the input and output of two CMOS inverters are cross-connected, that is, the output of the first inverter is connected to the input of the second inverter, and the output of the second inverter is connected to the input of the first inverter, realizing the latching of the output states of the two inverters. After this SRAM cell is given a state of 0 or 1, it will maintain this state until it is given a new state next time or the power is cut off before it will change or disappear.

[0050] In addition to its static storage characteristics, SRAM has very fast storage and read speeds, so it is often used as a cache. A cache is a temporary memory used to improve data access efficiency. It is located between the CPU (Central Processing Unit) and the main memory and is used to temporarily store frequently accessed data and instructions to quickly respond to the processor's read requests. SRAM is a common cache device in current general computer architectures. The main advantages of SRAM as a cache are its fast read and write speeds and high reliability, usually operating at a speed of 10 ns or faster. At the same time, the bistable flip-flop structure in SRAM can store data persistently, which gives SRAM high stability and reliability. The disadvantage of the SRAM structure is that its storage capacity is relatively small, usually expressed in bytes or smaller units. Because compared with the single-transistor single-capacitor structure of DRAM, SRAM has more transistors in a single storage cell, occupying a larger area, making it difficult and expensive to achieve large-capacity storage and reducing the integration degree of chip storage.

[0051] The SRAM prepared by traditional silicon-based transistors has fast operation speed and high calculation accuracy, but the area of the operation unit is 6 times larger than that of new memories such as RRAM (considering that additional transistors are required for in-memory computing, the multiple of the SRAM area is even larger). The back-end SRAM prepared using new semiconductor devices such as carbon nanotubes, oxide semiconductors, and two-dimensional materials can integrate the back-end cache on a silicon-based chip. Through monolithic three-dimensional integration technology, the back-end cache and high-precision digital in-memory computing functional modules can be realized, which can greatly save the chip area while increasing the bandwidth.

[0052] (4) Monolithic three-dimensional integration technology

[0053] Traditional silicon-based semiconductor processes require high-temperature processes (>1000 °C) such as active layer growth, ion implantation, and annealing. Therefore, after preparing a layer of transistors on a silicon wafer, it is impossible to continue using the same high-temperature process to prepare the second-layer device on the same chip. Using processes such as epitaxy and bonding is also limited by factors such as temperature and yield. Therefore, in traditional manufacturing technologies, modules such as CPUs, GPUs, and memories need to be manufactured, packaged separately, and then soldered on the PCB board, relying on the circuit leads on the PCB board to achieve data exchange. However, with the sharp increase in data transmission volume, the high parasitics and low density of the interconnections on the PCB board have led to inefficient communication and can no longer meet the requirements.

[0054] The advantages of monolithic three-dimensional integration technology lie in that by selecting semiconductor materials that are compatible with the back-end process and can be mass-produced, and using a low-temperature back-end process (≤400 °C), on a chip where silicon-based transistors have already been fabricated, multiple layers of new logic, memory, and memory-computation-in-one devices are vertically stacked. While greatly reducing the chip area, it increases the data transmission bandwidth between layers. Different from the TSV (Through-Silicon Vias) process with a millimeter-level diameter, the monolithic three-dimensional integration process does not require drilling holes and interconnecting multiple silicon wafers, but uses ILV (Inter-layer Vias) at the nanometer scale to achieve ultra-high-bandwidth interconnection of multiple device functional layers on a single silicon wafer, greatly improving the efficiency of data exchange.

[0055] The following describes the monolithic three-dimensional integration implementation method and device of the hybrid-precision in-memory computing architecture according to the embodiments of the present application with reference to the accompanying drawings. Regarding the problem mentioned in the above background technology that the regulation of the RRAM conductance value takes a long time, SRAM has a faster write speed and higher computing precision compared to RRAM, but occupies a large part of the chip area. In this method, by obtaining the silicon-based CMOS process and the back-end integration process in the monolithic three-dimensional integration technology, a control logic circuit is fabricated on the hybrid-precision in-memory computing architecture using the silicon-based CMOS process; multiple functional layers are vertically stacked on the fabricated logic circuit using the back-end integration process, and communication between the multiple functional layers is achieved through vertical interlayer dielectric vias, enabling data communication between RRAM, SRAM, and the silicon-based control module to be carried out through efficient on-chip interconnection, having a very high communication bandwidth, improving the computing efficiency of the chip, and at the same time reducing the chip area and increasing the computing parallelism. Thus, the problems in the related technology that the in-memory computing architecture usually occupies a large part of the chip area and has low computing efficiency are solved.

[0056] Specifically, Figure 2 FIG. is a schematic flowchart of a monolithic three-dimensional integration implementation method of a hybrid-precision in-memory computing architecture provided by an embodiment of the present application.

[0057] As Figure 2 shown, the monolithic three-dimensional integration implementation method of the hybrid-precision in-memory computing architecture includes the following steps:

[0058] In step S101, the silicon-based CMOS process and the back-end integration process in the monolithic three-dimensional integration technology are obtained.

[0059] Among them, CMOS (Complementary Metal-Oxide-Semiconductor) is a widely used technology for manufacturing microelectronic devices, especially suitable for digital logic circuits. The silicon-based CMOS process specifically refers to the process of using silicon as the substrate material to manufacture CMOS devices; the back-end integration process refers to the processing steps for forming interconnects, vias, and realizing key technologies such as vertical interconnection after the front-end transistor manufacturing is completed.

[0060] It can be understood that in the embodiments of the present application, the silicon-based CMOS process and the back-end integration process in the monolithic three-dimensional integration technology are first obtained. Silicon can be used as the substrate material to manufacture CMOS devices. After the front-end transistor manufacturing is completed, the back-end integration process is used to realize key technologies such as forming interconnects, vias, and vertical interconnection.

[0061] In step S102, a control logic circuit is manufactured on the hybrid-precision in-memory computing architecture using the silicon-based CMOS process.

[0062] Among them, the hybrid-precision in-memory computing architecture includes three functional layers: the first layer is a control logic circuit manufactured using the standard silicon-based CMOS process, the second layer is an in-memory computing module implemented using resistive random-access memory, which can perform multi-bit analog in-memory computing, has a high device density, and can be stacked in multiple layers. The third layer is an SRAM layer implemented using the back-end CMOS process, which can perform high-precision digital in-memory computing and can also be used as a cache module, and can be stacked in multiple layers to reduce the chip area. Manufacturing a control logic circuit on the hybrid-precision in-memory computing architecture using the silicon-based CMOS process has high performance, good reliability, and a mature process. Since the process temperature allows only one layer of silicon-based CMOS to be made for a single chip, the part with the highest reliability requirements in the circuit needs to be implemented using this process.

[0063] It can be understood that in the embodiments of the present application, a control logic circuit is manufactured on the hybrid-precision in-memory computing architecture using the silicon-based CMOS process, and the part with the highest reliability requirements in the circuit is selected to be implemented using this process.

[0064] In step S103, multiple functional layers are vertically stacked on the logic circuit using the back-end integration process, and communication between the multiple functional layers is realized through the interlayer dielectric vias in the vertical direction. Among them, an analog in-memory array, a digital in-memory array, and a cache are provided on the multiple functional layers. The analog in-memory array realizes the matrix-vector multiplication operation with fixed weights, and the digital in-memory array realizes the matrix multiplication operation for refreshing weights.

[0065] Among them, the interlayer dielectric via is a conductive path connecting different functional layers to achieve vertical signal transmission; the analog in-memory array can directly execute computing tasks inside the memory cell and uses analog signals to represent data; the digital in-memory array is similar to the analog in-memory array and works based on digital signals.

[0066] It can be understood that the embodiments of this application use a back-end integration process to stack multiple functional layers vertically in the vertical direction of the logic circuit. An analog in-memory array, a digital in-memory array, and a cache are provided on the multiple functional layers. The analog in-memory array implements matrix-vector multiplication operations with fixed weights, the digital in-memory array implements matrix multiplication operations for refreshing weights, and the vertical interlayer dielectric vias are used to ensure unobstructed communication between layers.

[0067] In the embodiments of this application, a back-end integration process is used to vertically stack multiple functional layers on the control logic circuit of the mixed-precision in-memory computing architecture, including: obtaining the target process step sequence of the back-end integration process; determining the target stacking sequence of the multiple functional layers according to the target process step sequence; and vertically stacking the multiple functional layers on the control logic circuit through the back-end integration process and the target stacking sequence.

[0068] Among them, the target process steps of the integration process include thin film deposition, photolithography, or dry etching, etc., which are selected and sorted according to actual needs and are not specifically limited here.

[0069] It can be understood that the embodiments of this application first need to determine the target process step sequence of the back-end integration process. Based on these process steps, determine the most suitable target stacking sequence between the multiple functional layers, and use the back-end process technology and the determined stacking sequence to accurately vertically stack these functional layers on the control logic circuit.

[0070] In the embodiments of this application, the multiple functional layers include a functional layer for analog in-memory, a functional layer for digital in-memory, and a functional layer for cache. Among them, an analog in-memory array is provided on the functional layer for analog in-memory, a digital in-memory array is provided on the functional layer for digital in-memory, and a cache is provided on the functional layer for cache.

[0071] It can be understood that the multiple functional layers of the embodiments of this application include a functional layer for analog in-memory, a functional layer for digital in-memory, and a functional layer for cache. An analog in-memory array is provided on the functional layer for analog in-memory, which can be used to execute matrix-vector multiplication operations with fixed weights; a digital in-memory array is provided on the functional layer for digital in-memory, which can be responsible for matrix multiplication operations for refreshing weights to adapt to dynamically changing data requirements; a cache is provided on the functional layer for cache, which can be used to temporarily store common data or intermediate calculation results.

[0072] In the embodiments of the present application, the functional layer of in-memory digital and the functional layer of the cache are integrated into the same functional layer, or the functional layer of in-memory digital and the functional layer of the cache are independent functional layers.

[0073] It can be understood that the functional layer of in-memory digital and the functional layer of the cache in the embodiments of the present application can be integrated into the same functional layer or can be independent functional layers from each other.

[0074] In the embodiments of the present application, the analog in-memory array and the digital in-memory array adopt one or more of hafnium oxide-based resistive random access memory (RRAM), phase change memory (PCM), magnetic random access memory (MRAM), ferroelectric random access memory (FeRAM), and conductive bridge random access memory (CBRAM).

[0075] Among them, the hafnium oxide-based RRAM is a non-volatile memory that can achieve multi-bit storage and is suitable for analog in-memory computing; the PCM can store information by changing the resistivity; the MRAM uses the spin direction of electrons to store binary information, featuring non-volatility and high-speed read and write, and is suitable for digital in-memory computing; the FeRAM stores information based on the polarization state of ferroelectric materials, with the advantages of fast write speed and high durability; the CBRAM changes the resistance state by forming or breaking conductive filaments under an applied voltage to achieve information storage.

[0076] It can be understood that the analog in-memory array and the digital in-memory array in the embodiments of the present application adopt a variety of advanced non-volatile storage technologies. The analog in-memory array can select hafnium oxide-based RRAM, which can efficiently support matrix-vector multiplication operations with fixed weights; for the digital in-memory array, one or more of PCM, MRAM, FeRAM, or CBRAM can be selected.

[0077] In the embodiments of the present application, the control logic circuit adopts any one of the back-end complementary FET (CFET) structures of carbon nanotubes and indium gallium zinc oxide (IGZO), the back-end CFET structure based on tungsten diselenide and molybdenum disulfide, and the back-end CFET structure based on P-type silicon and N-type silicon.

[0078] Among them, the back-end CFET structure refers to a complementary metal-oxide-semiconductor field-effect transistor (CMOSFET) structure constructed in the back-end integration stage after the completion of the traditional front-end process.

[0079] It can be understood that there are three back-end CFET structures for the control logic circuit in the embodiments of the present application, namely the back-end CFET structure of carbon nanotubes and IGZO, the back-end CFET structure based on tungsten diselenide and molybdenum disulfide, and the back-end CFET structure based on P-type silicon and N-type silicon.

[0080] In the embodiments of the present application, the number of layers of the respective functional layers of the analog in-memory array and the digital in-memory array is configured according to the algorithm requirements of the hybrid-precision in-memory computing architecture.

[0081] It can be understood that the embodiments of the present application can meet the algorithm requirements of the hybrid-precision in-memory computing architecture, such as which parts are suitable for the analog in-memory array and which parts are suitable for the digital in-memory array. Based on these requirements, the number of functional layers corresponding to each type of in-memory array is determined.

[0082] According to the monolithic three-dimensional integration implementation method of the hybrid-precision in-memory computing architecture proposed by the embodiments of the present application, a control logic circuit is fabricated on the hybrid-precision in-memory computing architecture using a silicon-based CMOS process, and multiple functional layers are vertically stacked on the fabricated logic circuit using a back-end integration process. Communication between the multiple functional layers is achieved through interlayer dielectric vias in the vertical direction, enabling data communication between RRAM, SRAM, and the silicon-based control module to be carried out through an efficient on-chip interconnection, with a very high communication bandwidth, improving the computing efficiency of the chip, reducing the chip area, and increasing the computing parallelism.

[0083] The monolithic three-dimensional integration implementation method of the hybrid-precision in-memory computing architecture is further described below through a specific embodiment.

[0084] (1) Hybrid-precision in-memory computing architecture implemented through monolithic three-dimensional integration

[0085] This embodiment proposes a hybrid-precision in-memory computing architecture that includes logic control, RRAM analog in-memory computing, and SRAM digital in-memory computing, and realizes the vertical stacking of different functional layers through monolithic three-dimensional integration. Each functional layer relies on high-density and low-parasitic-effect interlayer dielectric vias for communication. The RRAM and SRAM layers are integrated on the silicon-based control circuit through a back-end process, and the stacking order of the back-end functional layers can be swapped by adjusting the process step sequence.

[0086] Taking the implementation of the ViT algorithm as an example, the mixed-precision in-memory computing architecture includes three functional layers: The first layer is the control logic circuit fabricated using the standard silicon-based CMOS process. It features high performance, good reliability, and mature technology. However, due to the process temperature, only one layer of silicon-based CMOS can be fabricated on a single chip. Therefore, the part with the highest reliability requirements in the circuit is implemented using it. The second layer is the in-memory computing module implemented using resistive random-access memory (RRAM). It is characterized by performing multi-bit analog in-memory computing, having a high device density, and being able to stack multiple layers (repeating the process). It can only efficiently implement the matrix-vector multiplication operation in neural networks. The third layer is the SRAM layer implemented using the back-end CMOS process. It is characterized by performing high-precision digital in-memory computing. In addition to the computing module, the back-end SRAM can also serve as a cache module, can stack multiple layers, and reduces the chip area. It is used to implement the matrix multiplication operation that requires frequent weight refreshing in neural networks. By designing a three-dimensional structure, the SRAM layer can be exactly stacked directly above or below the RRAM layer to obtain extremely high communication bandwidth through the interlayer dielectric vias. The order and number of layers of the second and third layers can be adjusted according to requirements. The hybrid architecture is as Figure 3 shown, Figure 3 In Figure 3 , taking a single layer of RRAM, and the in-memory computing and caching of SRAM being on the same layer as an example, in fact, the in-memory computing of SRAM and the SRAM cache can be split into two layers.

[0087] (2) Monolithic three-dimensional integration implementation process

[0088] Taking the stacking of one layer of RRAM and one layer of SRAM as an example, and the SRAM layer being directly above the RRAM layer, the first layer is fabricated using the foundry-standard CMOS logic process, and the second layer is the in-memory computing module implemented using HfO 2 -based resistive random-access memory. The third layer is the near-memory computing layer implemented using the back-end CFET process.

[0089] The first layer is fabricated using the foundry-standard CMOS logic process, which will not be elaborated further. The second and third layers are fabricated using a back-end integration process at a low temperature (<= 300°C), including the following steps: (a) Depositing a 30-nm TiN (physical vapor deposition, bottom electrode) / 8-nm HfO 2 (atomic layer deposition, resistive change layer) / 45-nm TaO x (physical vapor deposition, thermal enhancement layer) / 30-nm TiN (physical vapor deposition, top electrode) stack. (b) Using photolithography and dry etching processes, selectively etching the TiN / HfO 2 / TaO x / TiN stack to achieve the patterning of the resistive random-access memory. (c) Using plasma-enhanced chemical vapor deposition to deposit 400-nm SiO 2Thin film (passivation layer). (d) Using photolithography and dry etching processes, etch the SiO 2 thin film to form openings for interconnect contact points. (e) Electroplate a layer of W, and then use chemical mechanical polishing to grind away the W except in the SiO 2 holes (forming metal vias). (f) Use physical vapor deposition to deposit 400 nm of metal Al (metal interconnection). (g) Use photolithography and dry etching processes to selectively etch Al to form Al metal interconnect lines. (h) Use plasma enhanced chemical vapor deposition to deposit a 1000 nm SiO 2 thin film (passivation layer). (i) Use photolithography and dry etching processes to selectively etch the SiO 2 thin film to form openings. (j) Electroplate a layer of W, and then use chemical mechanical polishing to grind away the W except in the SiO 2 holes (forming metal vias). (k) Use a wet transfer method to deposit a layer of carbon nanotubes. (l) Use photolithography, electron beam evaporation to deposit 30 nm of Pd, and then lift-off to form a pattern as the source and drain of the carbon nanotube transistor. (m) Use photolithography and oxygen plasma etching processes to selectively etch the carbon nanotubes to isolate different devices. (n) Electron beam evaporation deposits 10 nm of Y 2 O 3 , as a buffer layer. (o) Use atomic layer deposition to grow 10 nm of HfO 2 as the gate oxide of the carbon nanotube transistor. (p) Use photolithography and dry etching processes to selectively etch the HfO 2 to achieve an opening in the gate oxide. (q) Use photolithography, electron beam evaporation to deposit 45 nm of Pd, and then lift-off to form a pattern as the common gate of the subsequent CFET. (r) Use atomic layer deposition to grow 15 nm of HfO 2 as the gate oxide of the IGZO transistor. (s) Use photolithography and dry etching processes to selectively etch the HfO 2 to achieve an opening in the gate oxide. (t) Use atomic layer deposition to grow 10 nm of IGZO. (u) Use photolithography, electron beam evaporation to deposit 20 nm of Ti and 65 nm of Pd, and then lift-off to form a pattern as the source and drain of the IGZO transistor. (v) Use photolithography and wet etching processes to selectively etch IGZO to isolate different devices. (w) Use subsequent passivation and metal interconnection processes to form a metal interconnection pattern. The transmission electron microscope (TEM) photograph after the process is as shown in Figure 4 Figure.

[0090] (3) Algorithm implementation mapping

[0091] Taking the in-memory computing chip with mixed precision for the ViT algorithm as an example, its network structure is as shown in Figure 1 Figure, where the linear transformation (WQ ,W K ,W V ,W O ) The matrix-vector multiplication is implemented by the RRAM analog in-memory computing layer. The matrix multiplication (Q*K T , ) is implemented by the SRAM digital in-memory computing module. The input cache and the intermediate value cache of the matrix multiplication are implemented by the SRAM cache module. Other logics and operations such as data interfaces and softmax operations are implemented by silicon-based CMOS circuits, as Figure 5 shown. By simulating the architecture proposed in this embodiment, an inference accuracy rate similar to that of the GPU is achieved. As Figure 6 shown, it shows that the computing function of this architecture is correct, and its energy efficiency is improved by 32.39 times compared with the GPU, as Figure 7 shown. By evaluating the chip area and throughput, it is found that by using this monolithic three-dimensional integration method, when the number of transistors in the upper SRAM layer and the RRAM layer is the same, compared with the traditional planar technology, it has an advantage of 2.22 times in computing throughput. If the in-memory computing module and the cache module of the third-layer SRAM are split into two layers, the throughput can be increased by 12.78 times compared with the two-dimensional architecture, as Figure 8 shown.

[0092] Next, a monolithic three-dimensional integration implementation device of the hybrid-precision in-memory computing architecture proposed according to the embodiments of the present application will be described with reference to the accompanying drawings.

[0093] Figure 9 FIG. is a block diagram of a monolithic three-dimensional integration implementation device of the hybrid-precision in-memory computing architecture of the embodiments of the present application.

[0094] As Figure 9 shown, the monolithic three-dimensional integration implementation device 10 of the hybrid-precision in-memory computing architecture includes: an acquisition module 201, a manufacturing module 202, and a stacking module 203.

[0095] Among them, the acquisition module 201 is used to acquire the silicon-based CMOS process and the back-end integration process in the monolithic three-dimensional integration technology; the manufacturing module 202 is used to manufacture control logic circuits on the hybrid-precision in-memory computing architecture using the silicon-based CMOS process; the stacking module 203 is used to vertically stack multiple functional layers on the manufactured logic circuits using the back-end integration process, and realize communication between the multiple functional layers through the interlayer dielectric vias in the vertical direction. Among them, an analog in-memory array, a digital in-memory array, and a cache are provided on the multiple functional layers. The analog in-memory array realizes the matrix-vector multiplication operation of fixed weights, and the digital in-memory array realizes the matrix multiplication operation of refreshed weights.

[0096] In an embodiment of the present application, the stacking module 203 is further configured to: obtain the target process step sequence of the back-end integration process; determine the target stacking sequence of the multi-layer functional layers according to the target process step sequence; and vertically stack the multi-layer functional layers on the control logic circuit through the back-end integration process and the target stacking sequence.

[0097] In an embodiment of the present application, the multi-layer functional layers include a functional layer for analog in-memory, a functional layer for digital in-memory, and a functional layer for cache. Among them, an analog in-memory array is provided on the functional layer for analog in-memory, a digital in-memory array is provided on the functional layer for digital in-memory, and a cache is provided on the functional layer for cache.

[0098] In an embodiment of the present application, the functional layer for digital in-memory and the functional layer for cache are integrated on the same functional layer, or the functional layer for digital in-memory and the functional layer for cache are independent functional layers.

[0099] In an embodiment of the present application, the analog in-memory array and the digital in-memory array adopt one or more of hafnium oxide-based resistive random access memory (RRAM), phase change memory (PCM), magnetic memory, ferroelectric memory, and conductive bridge memory.

[0100] In an embodiment of the present application, the control logic circuit adopts any one of a back-end complementary FET (CFET) structure of carbon nanotubes and indium gallium zinc oxide (IGZO), a back-end CFET structure based on tungsten diselenide and molybdenum disulfide, and a back-end CFET structure based on P-type silicon and N-type silicon.

[0101] In an embodiment of the present application, the stacking module 203 is further configured to: configure the number of layers of the respective corresponding functional layers of the analog in-memory array and the digital in-memory array according to the algorithm requirements of the mixed-precision in-memory computing architecture.

[0102] It should be noted that the foregoing explanation of the embodiment of the method for realizing the monolithic three-dimensional integration of the mixed-precision in-memory computing architecture also applies to the device for realizing the monolithic three-dimensional integration of the mixed-precision in-memory computing architecture of this embodiment, and will not be elaborated here.

[0103] According to the device for realizing the monolithic three-dimensional integration of the mixed-precision in-memory computing architecture proposed in an embodiment of the present application, a control logic circuit is fabricated on the mixed-precision in-memory computing architecture using a silicon-based CMOS process, multi-layer functional layers are vertically stacked on the fabricated logic circuit using a back-end integration process, and communication between the multi-layer functional layers is realized through interlayer dielectric vias in the vertical direction, so that data communication between RRAM, SRAM, and the silicon-based control module can be carried out through efficient on-chip interconnection, having a very high communication bandwidth, improving the computing efficiency of the chip, and at the same time reducing the chip area and improving the computing parallelism.

[0104] An embodiment of the present application also proposes a mixed-precision in-memory computing architecture 20, as Figure 10As shown in the figure, the hybrid-precision in-memory computing architecture 20 includes: a control logic circuit 301 and multiple functional layers 302.

[0105] Among them, the multiple functional layers 302 are vertically stacked on the control logic circuit 301. The multiple functional layers 302 communicate through interlayer dielectric vias in the vertical direction. An analog in-memory array, a digital in-memory array, and a cache are provided on the multiple functional layers 302. The analog in-memory array implements the matrix-vector multiplication operation with fixed weights, and the digital in-memory array implements the matrix multiplication operation for refreshing weights.

[0106] It can be understood that the embodiment of the present application is composed of the control logic circuit 301 and the multiple functional layers 302 vertically stacked thereon. In the multiple functional layers 302, efficient communication is achieved between the functional layers through dielectric vias in the vertical direction, and an analog and digital in-memory array and a cache structure are integrated. Specifically, the analog in-memory array is used to perform the matrix-vector multiplication operation based on fixed weights, and the digital in-memory array is responsible for processing the matrix multiplication tasks that require dynamic weight refreshing, improving the data processing speed and efficiency, and providing an optimized solution for complex computing tasks.

[0107] According to the hybrid-precision in-memory computing architecture proposed in the embodiment of the present application, efficient communication is achieved between the functional layers through dielectric vias in the vertical direction, and an analog and digital in-memory array and a cache structure are integrated, improving the computing efficiency of the chip, reducing the chip area, and increasing the computing parallelism at the same time.

[0108] Figure 11 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application. The electronic device may include:

[0109] A memory 401, a processor 402, and a computer program stored on the memory 401 and executable on the processor 402.

[0110] When the processor 402 executes the program, it implements the monolithic three-dimensional integration implementation method of the hybrid-precision in-memory computing architecture provided in the above embodiment.

[0111] Furthermore, the electronic device further includes:

[0112] A communication interface 403 for communication between the memory 401 and the processor 402.

[0113] The memory 401 is used to store a computer program executable on the processor 402.

[0114] The memory 401 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.

[0115] If the memory 401, the processor 402, and the communication interface 403 are implemented independently, the communication interface 403, the memory 401, and the processor 402 can be interconnected through a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0116] Optionally, in a specific implementation, if the memory 401, the processor 402, and the communication interface 403 are integrated on a single chip, the memory 401, the processor 402, and the communication interface 403 can communicate with each other through an internal interface.

[0117] The processor 402 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.

[0118] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.

[0119] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0120] Any process or method description shown in the flowchart or described otherwise herein can be understood to represent a module, segment, or portion of code including one or N executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0121] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays, field programmable gate arrays, etc.

[0122] Those of ordinary skill in the art of the present technology can understand that all or part of the steps carried by the methods of implementing the above embodiments can be completed by instructing relevant hardware through a program. The above program can be stored in a computer-readable storage medium, and when executed, includes one or a combination of the steps of the method embodiments.

[0123] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limitations on the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A method for implementing a monolithic three-dimensional integrated mixed-precision in-memory computing architecture, characterized in that: The following steps are involved: Acquire silicon-based CMOS process and back-end integration process in monolithic 3D integration technology; The control logic circuit is manufactured using silicon-based CMOS technology on a mixed-precision in-memory computing architecture; A back-end integration process is used to vertically stack multiple functional layers on the logic circuit, and communication between the multiple functional layers is achieved through vertical interlayer dielectric vias, wherein the multiple functional layers are provided with an analog memory array, a digital memory array and a cache, the analog memory array realizes matrix and vector multiplication operations with fixed weights, and the digital memory array realizes matrix multiplication operations with refreshed weights.

2. The method for implementing a monolithic three-dimensional integrated mixed-precision in-memory computing architecture according to claim 1, characterized in that: The method of vertically stacking multiple functional layers on a control logic circuit of a mixed precision in-memory computing architecture using a back-end integration process includes: Obtaining a target process step sequence of the back-end integration process; Determining a target stacking order of the multiple functional layers according to the target process step order; Multiple functional layers are vertically stacked on the control logic circuit through the back-end integration process and the target stacking sequence.

3. The method for implementing a monolithic three-dimensional integrated mixed-precision in-memory computing architecture according to claim 1, characterized in that: The multi-layer functional layer includes a functional layer in analog memory, a functional layer in digital memory and a functional layer in cache, wherein the functional layer in analog memory is provided with an analog memory array, the functional layer in digital memory is provided with a digital memory array, and the functional layer in cache is provided with a cache.

4. The method for implementing a monolithic three-dimensional integrated mixed-precision in-memory computing architecture according to claim 3, characterized in that: The functional layer in the digital storage and the functional layer in the cache are integrated in the same functional layer, or the functional layer in the digital storage and the functional layer in the cache are independent functional layers.

5. The method for implementing a monolithic three-dimensional integrated mixed-precision in-memory computing architecture according to claim 1, characterized in that: The analog memory array and the digital memory array use one or more of hafnium oxide-based resistive memory, phase change memory, magnetic memory, ferroelectric memory and conductive bridge memory.

6. The method for implementing a monolithic three-dimensional integrated mixed-precision in-memory computing architecture according to claim 1, characterized in that: The control logic circuit adopts any one of a back-end CFET structure of carbon nanotubes and IGZO, a back-end CFET structure based on tungsten diselenide and molybdenum disulfide, and a back-end CFET structure based on P-type silicon and N-type silicon.

7. The method for implementing a monolithic three-dimensional integrated mixed-precision in-memory computing architecture according to claim 1, characterized in that: The number of functional layers corresponding to the analog in-memory array and the digital in-memory array is configured according to the algorithm requirements of the mixed-precision in-memory computing architecture.

8. A single-chip three-dimensional integrated implementation device of a mixed-precision in-memory computing architecture, characterized in that: include: Acquisition module, used to acquire silicon-based CMOS process and back-end integration process in monolithic 3D integration technology; A manufacturing module for manufacturing control logic circuits on a mixed-precision in-memory computing architecture using silicon-based CMOS technology; A stacking module is used to vertically stack multiple functional layers on the logic circuit using a back-end integration process, and to achieve communication between the multiple functional layers through vertical interlayer dielectric vias, wherein the multiple functional layers are provided with an analog memory array, a digital memory array and a cache, the analog memory array achieves matrix and vector multiplication operations with fixed weights, and the digital memory array achieves matrix multiplication operations with refreshed weights.

9. A mixed precision in-memory computing architecture, characterized in that: include: Control logic circuit; The multi-layer functional layers are vertically stacked on the logic circuit, wherein the multi-layer functional layers communicate through interlayer dielectric vias in the vertical direction, and the multi-layer functional layers are provided with analog memory arrays, digital memory arrays and caches, the analog memory arrays realize matrix and vector multiplication operations with fixed weights, and the digital memory arrays realize matrix multiplication operations with refresh weights.

10. An electronic device, characterized in that: Including the mixed precision in-memory computing architecture as described in claim 9.

Citation Information

Cited By

  • Software and hardware collaborative optimization method of hybrid in-memory architecture

    CN120973728A

  • Storage and calculation integrated neural network processor with three-dimensional integration and substrate interconnection

    CN121998007A

  • Design method and layout structure of three-dimensional integrated circuit

    CN122334168A