Integrated circuit device and method for preparing integrated circuit device
By stacking the configuration of the logic control circuit layer, the in-memory computing circuit layer and the cache circuit layer in the integrated circuit device, and using capacitively uncapacitorized dynamic random access memory, the bottlenecks in the improvement of integration and performance of the two-dimensional chip are solved, and efficient data transmission and computing performance are achieved.
Patent Information
- Application Number
- PCT/CN2024/134534
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-11-26
- Publication Date
- 2025-06-05
AI Technical Summary
Existing two-dimensional chips have encountered bottlenecks in terms of integration and performance improvement, and cannot meet the needs of high computing power.
An integrated circuit device is designed, including a logic control circuit layer, an in-memory computing circuit layer and a buffer circuit layer. It realizes communication connections between each other through a stacked structure and conductive vias, and uses capacitive dynamic random access memory for data storage.
By setting the memory unit and the processing unit in the same device, the communication distance between the memory unit and the processing unit is reduced, the data transmission speed between memory and computing is improved, and the integration and performance of the chip are improved.
Smart Images

Figure CN2024134534_05062025_PF_FP_ABST
Abstract
Description
Integrated circuit device and method for manufacturing an integrated circuit device
[0001] This application claims priority to Chinese Patent Application No. 202311616388.3 filed on November 29, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby cited in their entirety as part of this application. Technical Field
[0002] Embodiments of the present disclosure relate to an integrated circuit device and a method of manufacturing the integrated circuit device. Background Art
[0003] With the rapid advancement of technology and the rapid development of the digital age, the demand for computing power has exploded. Two-dimensional chips, also known as planar chips, are integrated circuits designed and manufactured on a single plane. In a two-dimensional chip, all circuits and transistors are distributed on the same plane and connected by planar metal wires. However, due to the limitations of this planar structure, their integration and performance have reached bottlenecks, making them unable to meet the current demand for high computing power. Summary of the Invention
[0004] Some embodiments of the present disclosure provide an integrated circuit device, the integrated circuit device comprising:
[0005] a logic control circuit layer configured to control the data processing flow;
[0006] The in-memory computing circuit layer is configured to perform matrix-vector calculation tasks;
[0007] A cache circuit layer includes a dynamic random access memory to store data in the matrix vector calculation task, wherein the dynamic random access memory is a capacitor-less dynamic random access memory, wherein the cache circuit layer, the in-memory calculation circuit layer and the logic control circuit layer are at least partially stacked with each other in a first direction and are communicatively connected to each other.
[0008] For example, an integrated circuit device provided in some embodiments of the present disclosure also includes an interlayer passivation layer that separates the cache circuit layer, the in-memory computing circuit layer, and the logic control circuit layer from each other, wherein the interlayer passivation layer includes a plurality of conductive vias, and the cache circuit layer, the in-memory computing circuit layer, and the logic control circuit layer are communicatively connected through the conductive vias.
[0009] For example, in an integrated circuit device provided in some embodiments of the present disclosure, the interlayer passivation layer includes: a first passivation layer between the in-memory computing circuit layer and the cache circuit layer and a second passivation layer between the logic control circuit layer and the in-memory computing circuit layer, or a second passivation layer between the logic control circuit layer and the cache circuit layer. The plurality of conductive vias includes: a plurality of first metal vias in the first passivation layer and a plurality of second metal vias in the second passivation layer.
[0010] For example, in an integrated circuit device provided in some embodiments of the present disclosure, the capacitor-less dynamic random access memory includes an N-type first transistor and a P-type second transistor, wherein the material of the active layer of the first transistor is different from the material of the active layer of the second transistor.
[0011] For example, in an integrated circuit device provided in some embodiments of the present disclosure, the active layer of the first transistor includes at least one of the following semiconductor oxides: IGZO, ITO, IWO and IAZO; the active layer of the second transistor includes at least one of the following materials: CNT, WSe2 and MoS2.
[0012] For example, in an integrated circuit device provided in some embodiments of the present disclosure, the in-memory computing circuit layer includes at least one in-memory computing circuit sub-layer and multiple in-memory computing arrays, and the multiple in-memory computing arrays are located in the at least one in-memory computing sub-layer.
[0013] For example, in an integrated circuit device provided in some embodiments of the present disclosure, each of the multiple in-memory computing arrays includes at least one memristor array, and the memristor array includes a plurality of memristors arranged in multiple rows and columns.
[0014] For example, in an integrated circuit device provided in some embodiments of the present disclosure, the logic control circuit layer includes a logic control circuit and a peripheral circuit, wherein the logic control circuit is configured to control the overall function, and the peripheral circuit is configured to perform at least one of the following functions: routing, layer skipping, rearrangement, and activation function.
[0015] For example, an integrated circuit device provided in some embodiments of the present disclosure also includes a silicon substrate, wherein the logic control circuit layer is located on the silicon substrate, the in-memory computing circuit layer is located on a side of the logic control circuit layer away from the silicon substrate, and the cache circuit layer is located on a side of the in-memory computing circuit layer away from the silicon substrate.
[0016] Some embodiments of the present disclosure provide a method for preparing an integrated circuit device, for preparing the integrated circuit device described in any of the above embodiments, the method comprising: preparing a logic control circuit layer, wherein the logic control circuit layer is configured to control a data processing flow; preparing an in-memory computing circuit layer, wherein the in-memory computing circuit layer is configured to perform matrix-vector calculation tasks; and preparing a cache circuit layer, wherein the cache circuit layer includes a dynamic random access memory (DRAM) to store data in the matrix-vector calculation tasks, the DRAM being a capacitor-less dynamic random access memory. The cache circuit layer, the in-memory computing circuit layer, and the logic control circuit layer are at least partially stacked on each other in a first direction and are communicatively connected to each other.
[0017] For example, a method for preparing an integrated circuit device provided in some embodiments of the present disclosure also includes: preparing an interlayer passivation layer and preparing a plurality of conductive vias in the interlayer passivation layer, wherein the interlayer passivation layer is configured to separate the cache circuit layer, the in-memory computing circuit layer, and the logic control circuit layer from each other, and the plurality of conductive vias are configured to communicatively connect the cache circuit layer, the in-memory computing circuit layer, and the logic control circuit layer.
[0018] For example, a method for preparing an integrated circuit device provided in some embodiments of the present disclosure also includes: providing a silicon substrate to prepare the cache circuit layer, the in-memory computing circuit layer and the logic control circuit layer on the silicon substrate, wherein the logic control circuit layer is located on the silicon substrate, the in-memory computing circuit layer is located on a side of the logic control circuit layer away from the silicon substrate, and the cache circuit layer is located on a side of the in-memory computing circuit layer away from the silicon substrate.
[0019] For example, in a method for preparing an integrated circuit device provided in some embodiments of the present disclosure, the logic control circuit layer is prepared on the silicon substrate by a silicon-based CMOS process, and the cache circuit layer and the in-memory computing circuit layer are prepared by a back-end integration process, wherein the operating temperature of the back-end integration process is lower than that of the silicon-based CMOS process. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, rather than limiting the present disclosure.
[0021] FIG1 shows a schematic structural diagram of an exemplary memristor;
[0022] FIG2 shows a schematic diagram of an exemplary memristor array structure;
[0023] FIG3 shows a schematic diagram of the working principle of an exemplary in-memory computing array;
[0024] FIG4 shows a schematic structural diagram of an integrated circuit device provided by at least one embodiment of the present disclosure;
[0025] FIG5 shows a transmission electron microscope image of an integrated circuit device provided by at least one embodiment of the present disclosure;
[0026] FIG6 shows a schematic diagram of a 2TOC structure provided by at least one embodiment of the present disclosure;
[0027] FIG7 shows a schematic diagram of a working process of an integrated circuit device provided by at least one embodiment of the present disclosure; and
[0028] FIG8 is a schematic flow chart showing a method for fabricating an integrated circuit device according to at least one embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the embodiments of the present disclosure will be further described in detail below in conjunction with the accompanying drawings. The specific embodiments and drawings described herein are only used to explain the present disclosure, rather than to limit the disclosed embodiments. Unless there is a conflict, the various embodiments of the present disclosure and the various features therein may be combined with each other.
[0030] For ease of description, the drawings of the embodiments of the present disclosure only show the parts related to the embodiments of the present disclosure, and the parts not related to the embodiments of the present disclosure are not shown in the drawings. Each unit and module involved in the embodiments of the present disclosure may correspond to only one physical structure, or may be composed of multiple physical structures, or multiple units and modules may be integrated into one physical structure. In the absence of conflict, the functions and steps marked in the flowcharts and block diagrams of the embodiments of the present disclosure may occur in an order different from that marked in the drawings.
[0031] The flowcharts and block diagrams of the embodiments of the present disclosure illustrate the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to the embodiments of the present disclosure. Each box in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified functions. Moreover, each box or combination of boxes in the block diagram and flowchart may be implemented by a hardware-based system that implements the specified functions, or by a combination of hardware and computer instructions.
[0032] In order to keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and known components may be omitted. When any component of the embodiments of the present disclosure appears in more than one drawing, the component is represented by the same or similar reference numerals in each drawing.
[0033] Memristors (e.g., resistive random access memory, phase-change memory, conductive bridge memory, etc.) are non-volatile devices whose conductance state can be adjusted by applying an external stimulus. As two-terminal devices, memristors offer adjustable resistance and are non-volatile, making them widely used in integrated storage and computing applications. According to Kirchhoff's current law and Ohm's law, an array of memristors can perform multiplication and accumulation calculations in parallel, with both storage and computation occurring within each device in the array. This computing architecture enables integrated storage and computing without requiring large amounts of data movement.
[0034] For example, as an example of a memristor, resistive random access memory (RRAM) is a new type of non-volatile memory that uses the resistive properties of materials to store information. It consists of a phase change material, a lower electrode, and an upper electrode. During operation, external stimuli (such as voltage) drive ions and change the local structure of the storage medium (phase change material), resulting in a change in the resistance subsequently used to store data.
[0035] FIG1 shows a structural diagram of an exemplary memristor. As shown in FIG1 , the memristor RRAM includes a resistive layer 111 and an upper electrode 113 and a lower electrode 114 located on both sides of the resistive layer 111. In at least one example, the memristor may further include a functional layer 112. The functional layer 112 is an optional layer, and whether to add it or not can be determined based on the optimization direction of the memristor performance, and it can be designed accordingly. The resistive layer 111 can be, for example, a single layer, including a single type of binary metal oxide (such as NiO, AlOx, etc.), graphene oxide, multi-element perovskite oxide (such as STO, SZO, PCMO, etc.), or a multilayer, such as any optional stack of the above materials, for example, a stack of TixN and AlOx.
[0036] Figure 2 shows a schematic diagram of a memristor array structure. As shown in Figure 2, the memristor array is composed of multiple memristor units, which form an array of M rows and N columns, where M and N are both positive integers. Each memristor unit includes a switch element and one or more memristors. In Figure 2, WL <1> 、WL <2> ...WL <m>BL represents the word lines of the first row, the second row, ... the Mth row, respectively. The control electrodes (e.g., gates of transistors) of the switch elements in the memristor unit circuit of each row are connected to the word lines corresponding to the row; <1> BL <2> ...BL <n>They represent the bit lines of the first column, the second column, ... the Nth column, respectively. The memristors in the memristor unit circuit of each column are connected to the bit lines corresponding to the column; SL <1> , SL <2> ...SL <m>They represent the source lines of the first row, the second row, ... the Mth row respectively, and the sources of the transistors in the memristor unit circuit of each row are connected to the source line corresponding to the row.
[0037] The memristor unit in the memristor array of Figure 2 can be, for example, a 1T1R structure or a 2T2R structure, wherein the memristor unit of the 1T1R structure includes a switching transistor and a memristor, and the memristor unit of the 2T2R structure includes two switching transistors and two memristors. The present disclosure has no restrictions on the type, structure, etc. of the memristor device. It should be noted that the transistors used in the embodiments of the present disclosure can all be thin film transistors or field effect transistors (such as MOS field effect transistors) or other switching devices with the same characteristics. The source and drain of the transistor used here can be symmetrical in structure, so the source and drain can be indistinguishable in structure. The embodiments of the present disclosure do not limit the type of transistor used.
[0038] In-memory computing technology can be used to perform matrix operations, such as calculating the product of a vector and a matrix. The matrix element values are stored in the memory. The vector is input into the memory, and the memory itself can perform the operation and output the product result. For example, in-memory computing technology can be implemented using resistive random access memory (RRAM). RRAM can be thought of as a resistor with variable conductance. By mapping the first value to the RRAM's conductance and the second value to the RRAM's input voltage, Ohm's law can be used to determine that the output current is the product of the input voltage and the RRAM's conductance, that is, the product of the first value and the second value. By performing analog-to-digital conversion on the output current, a digital value that can be subsequently processed can be obtained.
[0039] Figure 3 shows a schematic diagram of the operating principle of an in-memory computing array. As shown in Figure 3, according to Kirchhoff's law, a memory array composed of non-volatile memory cells can perform matrix-vector multiplication. The data to be stored can be mapped to the resistance value of each memory cell and written into the memory array. The input vector is mapped to the read voltage of each row of the memory array. The output current of each column is the product of the input voltage and the conductance of the resistive memory, that is, ∑I = ∑V*G, where I represents the output current of each column in the memory array, V represents the input voltage, and G represents the conductance of the resistive memory.
[0040] Compared to the von Neumann architecture, the in-memory computing architecture offers advantages such as faster computing speed, lower power consumption, and higher integration density. By integrating computing functions into the memory unit, in-memory computing reduces the frequent movement of data between the data storage module and the computing module, and also reduces data transmission latency. Furthermore, in-memory computing integrates computing and storage functions on the same chip, reducing the need for external connections and wiring, resulting in a higher chip integration density and enabling its application in smaller, thinner electronic devices. Neural networks are key technologies in fields such as deep learning, requiring enormous computational effort and extremely high computational efficiency. In-memory computing not only meets the computational demands of neural networks, but also achieves high-performance computing with low power consumption.
[0041] In-memory computing technology based on resistive random access memory (RRAM) can quickly implement matrix-vector multiplication operations using basic physical laws, effectively implementing neural network calculations and improving computing power and energy efficiency. However, when faced with large-data throughput neural networks (such as those requiring high-speed, high-resolution video processing), a large amount of data is generated. The data in the in-memory computing array is cached through peripheral CMOS (Complementary Metal Oxide Semiconductor) circuit modules for temporary storage and retrieval. However, these cache modules often dominate the system area and power consumption.
[0042] Traditional storage and computing integrated chips based on two-dimensional planes are often limited by the cache modules on the silicon substrate because they have no area advantage and are difficult to achieve high-density integration. Moreover, within the two-dimensional plane, the data transmission efficiency between the in-memory computing array and the cache module is restricted by limited bandwidth, making it impossible to quickly process large amounts of data. As a result, the speed and energy consumption of the overall chip will perform poorly.
[0043] In addition, the preparation of traditional silicon-based integrated circuit semiconductor processes requires many process steps, including high-temperature processes such as active layer growth, ion implantation, and annealing, as well as temperature-sensitive processes such as metal interconnects. In the transistor device preparation process, these temperature-sensitive processes are often placed after the high-temperature process to avoid the impact of high temperature on device performance and process preparation. Therefore, for traditional silicon-based processes, after completing the preparation of a layer of transistors, it is difficult to prepare a second layer of devices on top of it, otherwise it will have a huge impact on the overall performance of the chip. Traditional silicon-based semiconductor integrated circuits are basically based on planes, with long data paths, and more speed and power consumption are lost in the data paths.
[0044] Monolithic 3D integration is an advanced semiconductor manufacturing technology that allows multiple layers of new logic, storage, and computing devices to be integrated in later processes to be stacked vertically on a single chip. This technology has two main advantages:
[0045] The first advantage is temperature compatibility with back-end process steps. Semiconductor processing is divided into front-end and back-end. The front-end generally refers to transistor fabrication, which operates at higher temperatures; the back-end refers to the lower-temperature process steps after transistor fabrication, such as metal interconnects. Some new storage and computing devices today have process temperatures that are back-end compatible, allowing them to be stacked vertically on a single chip.
[0046] The second advantage is that it significantly reduces chip area. Because it uses vertical stacking of devices and arrays, the data paths between layers are shorter, which can greatly increase data transmission bandwidth, improve chip speed and integration.
[0047] Of course, unlike the millimeter-scale diameter through-silicon via (TSV) process in the "three-dimensional integration technology" at the packaging level, the monolithic three-dimensional integration process does not require multi-layer silicon wafer drilling and interconnection. Instead, it uses inter-layer vias (ILV) at the hundred-nanometer scale to achieve ultra-high bandwidth interconnection of multi-layer device functional layers on a single-layer silicon wafer, greatly improving the efficiency of data exchange.
[0048] The inventors also noted that memory devices play an increasingly important role in storage-computing architectures. The current manufacturing process for large-capacity 1T1C DRAM (Dynamic Random Access Memory) requires specialized capacitors and transistors, making it difficult to match the advanced manufacturing processes for logic circuits. Furthermore, von Neumann storage-computing architectures often separate the DRAM memory from the CPU (Central Processing Unit), requiring long-distance communication for data exchange. Consequently, the data transmission speeds between storage and computing are gradually mismatched.
[0049] At least one embodiment of the present disclosure provides an integrated circuit device, comprising a logic control circuit layer, an in-memory computing circuit layer, and a cache circuit layer. The logic control circuit layer is configured to control a data processing flow; the in-memory computing circuit layer is configured to perform matrix-vector computing tasks; the cache circuit layer includes a dynamic random access memory (DRAM) to store data in matrix-vector computing tasks, wherein the DRAM is a capacitor-free dynamic random access memory, and the cache circuit layer, the in-memory computing circuit layer, and the logic control circuit layer are at least partially stacked with each other in a first direction and are communicatively connected to each other. The "first direction" here refers to the stacking direction of the three.
[0050] The embodiment of the present disclosure arranges the storage unit and the processing unit in the same device by stacking a cache circuit layer, an in-memory computing circuit layer, and a logic control circuit layer, thereby reducing the communication distance between the storage unit and the processing unit and improving the data transmission speed between storage and computing.
[0051] In at least one disclosed embodiment, the integrated circuit device further includes an interlayer passivation layer that separates the cache circuit layer, the in-memory computing circuit layer, and the logic control circuit layer from each other, wherein the interlayer passivation layer includes a plurality of conductive vias, and the cache circuit layer, the in-memory computing circuit layer, and the logic control circuit layer are communicatively connected through the conductive vias.
[0052] The interlayer passivation layer of the embodiment of the present disclosure is used for insulation and isolation, protecting the electrical devices in the cache circuit layer, the in-memory computing circuit layer and the logic control circuit layer, and realizing communication and interconnection among the cache circuit layer, the in-memory computing circuit layer and the logic control circuit layer through conductive vias.
[0053] For example, in at least one embodiment, the interlayer passivation layer may include: a first passivation layer between the in-memory computing circuit layer and the cache circuit layer and a second passivation layer between the logic control circuit layer and the in-memory computing circuit layer, or a second passivation layer between the logic control circuit layer and the cache circuit layer, and the multiple conductive vias include: a plurality of first metal vias in the first passivation layer and a plurality of second metal vias in the second passivation layer.
[0054] Figure 4 shows a schematic diagram of the structure of an integrated circuit device 100 provided in at least one embodiment of the present disclosure. As shown in Figure 4, integrated circuit device 100 includes a logic control circuit layer 101, an in-memory computing circuit layer 102, a cache circuit layer 103, a first passivation layer 104, a second passivation layer 105, a first metal via 106, a second metal via 107, and a silicon substrate 108.
[0055] In this embodiment, the logic control circuit layer 101, the in-memory computing circuit layer 102, and the cache circuit layer 103 are at least partially stacked and communicate with each other through multiple metal vias. For example, the logic control circuit layer 101, the in-memory computing circuit layer 102, and the cache circuit layer 103 can be completely stacked together in the first direction; for another example, two of them can be completely stacked together, while another can partially overlap with the other two; for example, any two of the three can partially stack with each other, and so on.
[0056] As shown in FIG4 , for example, the logic control circuit layer 101 is located on a silicon substrate 108, the in-memory computing circuit layer 102 is located on a side of the logic control circuit layer 101 away from the silicon substrate 108, and the cache circuit layer 103 is located on a side of the in-memory computing circuit layer 102 away from the silicon substrate 108. Here, the stacking direction is perpendicular to the surface of the silicon substrate 108.
[0057] For example, in the integrated circuit device shown in Figure 4, the logic control circuit layer 101 is connected to the in-memory computing circuit layer 102 through the second passivation layer 105, and the electrical interconnection between the logic control circuit layer 101 and the in-memory computing circuit layer 102 is achieved through the first metal via 106; the in-memory computing circuit layer 102 is connected to the cache circuit layer 103 through the first passivation layer 104, and the electrical interconnection between the cache circuit layer 103 and the in-memory computing circuit layer 102 is achieved through the second metal via 107.
[0058] For example, conductive vias can be vias directly derived from CMOS devices, which are used to realize metal interconnection. Usually, metal interconnection is prepared through the back-end process in integrated circuit manufacturing. The process temperature is relatively low, and it is a necessary step after the front-end transistor is completed, and is used to form electrical interconnection. In the preparation process, SiO2 is often used for passivation, isolation, and insulation; metal interconnection is divided into metal vias and metal connections. Metal vias are generally made of metal tungsten (W) material, and the process is to open the hole first and then fill it with W; W has low resistance and good contact performance, and is usually used to connect holes between different metal connection layers; metal connections can generally have many layers to facilitate layout and wiring, and are usually made of metal aluminum (Al) material or metal copper (Cu) material for large-scale metal interconnection. It should be noted that the present disclosure does not limit the materials used for metal vias.
[0059] For example, the interlayer passivation layer can be made of SiO2 material, which uses the insulating properties of SiO2 to achieve device passivation and protect the CMOS device from electrical interference from other non-insulating materials. Based on the embodiment shown in Figure 4, the positions of the in-memory computing circuit layer 102 and the cache circuit layer 103 can be interchanged.
[0060] In the above embodiment, by stacking the logic control circuit layer, the in-memory computing circuit layer, and the cache circuit layer in one direction, the separation of the memory and the processing unit is avoided, and the communication distance of data interaction is reduced.
[0061] In the above embodiment, for example, the logic control circuit layer 101 shown in Figure 4 includes a logic control circuit and a peripheral circuit, wherein the logic control circuit is configured to control the overall function, and the peripheral circuit is configured to perform at least one of the following functions: routing, layer skipping, rearrangement, and activation function.
[0062] Among them, routing refers to the process of directing and forwarding data received from the previous data interface to another data interface according to the destination address of the data; layer skipping is to skip certain layers in the integrated circuit during data transmission to reduce the data transmission time and improve the operating speed of the circuit; rearrangement refers to rearranging the data transmission order of devices in the circuit to shorten the data transmission path length in the circuit; activation function is a function running on the neurons of the artificial neural network, which is responsible for mapping the input of the neuron to the output.
[0063] For example, the in-memory computing circuit layer 102 shown in FIG4 includes one or more in-memory computing arrays. For example, each in-memory computing array includes at least one memristor array. When the in-memory computing array includes multiple memristor arrays, these memristor arrays are arranged in parallel. For example, they may have their own input circuits, word line driver circuits, bit line driver circuits, source line driver circuits, readout circuits, etc., or they may share some peripheral circuits, such as a shared word line driver circuit. Multiple in-memory computing arrays may be interconnected, for example, via a data bus.
[0064] In at least one embodiment of the present disclosure, the in-memory computing circuit layer includes at least one in-memory computing circuit sublayer and multiple in-memory computing arrays, and the multiple in-memory computing arrays are located in at least one in-memory computing sublayer. For example, the in-memory computing circuit layer can be stacked by multiple in-memory computing sublayers, each in-memory computing sublayer includes one or more in-memory computing arrays, and the in-memory computing arrays in different sublayers can be interconnected with each other through metal vias in the interlayer passivation layer.
[0065] Figure 5 shows a transmission electron microscope image of an integrated circuit device provided by at least one embodiment of the present disclosure, wherein the in-memory computing circuit layer is an in-memory computing circuit layer based on resistive random access memory, which can be used to implement matrix-vector multiplication in neural networks, can be stacked in multiple layers, and is compatible with back-end processes.
[0066] For example, each of the multiple in-memory computing arrays includes at least one memristor array, and the memristor array includes a plurality of memristors arranged in multiple rows and columns.
[0067] For example, in addition to using hafnium oxide-based resistive memory, the in-memory computing circuit layer can also use other types of applicable memristors, such as phase change memory, magnetic memory, ferroelectric memory, conductive bridge memory, etc. They can be implemented with different materials. For example, the most commonly used resistive memory can use a single layer of hafnium oxide, tantalum oxide, titanium oxide, silicon oxide, silicon nitride and other materials as the resistive layer, or a multi-layer structure, such as adding a layer of material to improve performance on the basis of the resistive layer, such as tantalum oxide, titanium oxide, yttrium oxide, aluminum oxide, etc. The chemical ratio of different materials can be adjusted according to actual needs.
[0068] For example, the integrated circuit device provided by at least one embodiment of the present disclosure can be used for neural network calculations. For example, the neural network weight parameters can be directly programmed as the conductance of the memristor array, or the weight parameters can be mapped to the conductance of the memristor array according to a certain rule.
[0069] As mentioned above, the current large-capacity memory 1T1C DRAM manufacturing process has special requirements for capacitors and transistors, which are difficult to match with the current advanced manufacturing process of logic circuits.
[0070] The DRAM used in the cache circuit layer in the embodiments of the present disclosure is exemplified below through some specific embodiments.
[0071] 2T0C memory devices, also known as capacitor-less DRAM, are highly scalable because they require no capacitors and meet the speed, density, and energy requirements of hardware systems. Compared to SRAM (Static Random-Access Memory), 2T0C devices offer higher density, larger capacity, and lower leakage current. Compared to DRAM, they offer non-destructive read and write speeds.
[0072] In at least one embodiment of the present disclosure, the DRAM in the cache circuit layer is a capacitor-less dynamic random access memory. For example, the capacitor-less dynamic random access memory includes a first N-type transistor and a second P-type transistor. For example, the material of the active layer of the first transistor is different from the material of the active layer of the second transistor. For example, the first transistor is a write transistor and the second transistor is a read transistor, where the write transistor is used to write received data to the read transistor. The cache circuit layer of the embodiment of the present disclosure uses a capacitor-less dynamic random access memory, which can improve the storage density of the memory.
[0073] FIG6 is a schematic diagram showing a DRAM having a 2T0C structure provided in a cache circuit layer according to at least one embodiment of the present disclosure.
[0074] As shown in Figure 6, this 2TOC DRAM structure includes a write transistor Wtr and a read transistor Rtr. The drain of the write transistor Wtr and the gate of the read transistor Rtr are connected to form a storage node SN. Data is stored by utilizing the capacitance characteristics of the transistor gates. In the embodiments of the present disclosure, in addition to NP-type 2TOC memory devices, PN-type 2TOC memory devices and other 2TOC memory devices can also be used.
[0075] Since the off-state current of silicon transistors is large and the leakage is severe, the data retention time is very short. If a 2T0C structure is used, the effective data retention time is usually no more than 1ms. Therefore, in at least one embodiment of the present disclosure, the active layer of the transistor of the 2T0C structure DRAM uses a non-silicon semiconductor material, and for example, the two transistors use semiconductor materials with different electrical properties for the active layer. Therefore, in at least one embodiment of the present disclosure, the active layer of the first transistor includes at least one of the following semiconductor oxides: IGZO, ITO, IWO and IAZO; the active layer of the second transistor includes at least one of the following materials: carbon nanotubes (CNT), WSe2 and MoS2.
[0076] The use of IGZO material to make write transistors increases the data retention time due to the low leakage of this material. In addition, the use of oxide semiconductors represented by indium gallium zinc oxide (IGZO) to prepare 2T0C devices makes it possible to achieve effective storage under small-size conditions. Compared with traditional silicon-based devices, TFTs (thin film transistors) made based on such oxide semiconductors have extremely low leakage, and the data retention time is increased by an order of magnitude compared to silicon devices, which is conducive to meeting the time requirements for data caching. Materials such as carbon nanotubes (CNTs), WSe2 and MoS2 have high mobility, which makes data operation time faster and helps to achieve data caching.
[0077] For example, a 2T0C memory device is composed of dual N-type transistors based on oxide semiconductors. However, the oxide semiconductor mobility of NN-type 2T0C DRAM is relatively low, which results in a low read current and, consequently, a slow read speed. Furthermore, under the same programming scheme, the read and write operations of the dual N-type 2T0C will affect the integrity of the stored data due to the coupling between the two N-type transistors. However, the read and write operations of the NP-type 2T0C are complementary to the stored data, extending the data retention time of the 2T0C DRAM.
[0078] Therefore, in the above embodiment, a low-leakage oxide semiconductor (such as IGZO) is used to prepare the write transistor of the 2T0C DRAM, and a high-mobility P-type material (such as CNT) is used to prepare the read transistor of the 2T0C DRAM, achieving high density, low leakage, long data retention time, and high-speed reading and writing.
[0079] At least one embodiment of the present disclosure further provides an electronic device, which includes the integrated circuit device of any of the above embodiments. The electronic device can be a server device, such as a server, or a client device, such as a mobile phone, laptop computer, or other terminal.
[0080] Figure 7 shows a schematic diagram of the workflow of an integrated circuit device provided by at least one embodiment of the present disclosure. For example, neural network data is input into the in-memory computing circuit layer, and the in-memory computing processing unit in the in-memory computing circuit layer performs neural network data processing. The in-memory computing processing unit can be a combination of multiple memristor units in the memristor array. The 2T0C cache array in the cache circuit layer stores the data used in the neural network calculation. After processing, the data passes through the data interface and routing circuit and can enter the cache circuit layer based on the 2T0C cache array for caching or enter other in-memory computing processing units, where the routing circuit is located in the logic control circuit layer. Finally, the data is efficiently transmitted through the routing circuit in the logic control circuit layer.
[0081] For example, the integrated circuit device provided in at least one embodiment of the present disclosure can be used to implement the YOLOv3 (You Only Look Once version.3) algorithm. YOLOv3 is an algorithm that uses a deep convolutional neural network to extract features and detect targets in images. It can identify and locate multiple objects in an image, or perform image segmentation and data analysis. It has applications in autonomous driving, video surveillance, robotic vision and other fields.
[0082] The YOLOv3 algorithm first extracts features from the input image using a deep convolutional neural network, divides the image into a grid, and predicts its bounding box. It then performs multi-scale predictions to identify objects. Finally, it uses a non-maximum suppression algorithm to filter out overlapping detections and output the final object detection result. In high-resolution video processing tasks, the quality of the recognition algorithm is measured by its high recognition speed. A higher number of frames per second indicates faster recognition, which offers a greater advantage in scenarios requiring real-time processing. The YOLOv3 algorithm boasts fast inference speed and, because its predictions are based on the global context of the image, offers good accuracy and wide applicability.
[0083] An in-memory computing circuit layer based on resistive random access memory (RRAM) is used to implement matrix-vector multiplication in neural networks. Data passes through a data interface and routing circuits and is cached in a 2TOC DRAM cache circuit layer. When running the YOLOv3 object detection algorithm, the architecture proposed in at least one embodiment of the present disclosure achieves a significantly higher frame per second (FPS) than conventional planar SRAM, achieving a 48.25x speed advantage while maintaining the same detection accuracy. Furthermore, the architecture's area is approximately 18.88% of that of conventional architectures.
[0084] At least one embodiment of the present disclosure further provides a method for manufacturing the integrated circuit device of at least one embodiment described above. FIG8 shows a flow chart of the method for manufacturing the integrated circuit device according to at least one embodiment of the present disclosure. As shown in FIG8 , the method for manufacturing the integrated circuit device includes:
[0085] Step S301: Prepare a logic control circuit layer, wherein the logic control circuit layer is configured to control a data processing flow.
[0086] Step S302: Prepare an in-memory computing circuit layer, wherein the in-memory computing circuit layer is configured to perform matrix-vector computing tasks.
[0087] Step S303: Prepare a cache circuit layer, wherein the cache circuit layer includes a dynamic random access memory (DRAM) to store data in the matrix vector calculation task, and the DRAM is a capacitor-less dynamic random access memory.
[0088] The cache circuit layer, the in-memory computing circuit layer and the logic control circuit layer are at least partially stacked on each other in a first direction and are communicatively connected to each other.
[0089] The embodiment of the present disclosure realizes setting the storage unit and the processing unit in the same device by stacking the cache circuit layer, the in-memory computing circuit layer and the logic control circuit layer, which can reduce the communication distance between the storage unit and the processing unit and improve the data transmission speed between storage and computing.
[0090] For example, the logic control circuit layer provided in the embodiment of the present disclosure may be prepared using a standard CMOS logic manufacturing process provided by a foundry. The CMOS logic manufacturing processes of different foundries may be different from each other, but this will not be described in detail here.
[0091] For example, in a specific example, preparing the in-memory computing circuit layer may include the following steps:
[0092] Step 201 : Deposit a stack of 30nm TiN (physical vapor deposition, lower electrode) / 8nm HfO2 (atomic layer deposition, resistive layer) / 45nm TaOx (physical vapor deposition, thermal enhancement layer) / 30nm TiN (physical vapor deposition, upper electrode).
[0093] Step 202: selectively etch the TiN / HfO2 / TaOx / TiN stack using photolithography and dry etching processes to achieve patterning of the resistive random access memory.
[0094] For example, in a specific example, preparing the cache circuit layer may include the following steps:
[0095] Step 211 : Define the gate region by photolithography, deposit 10 nm Ti and 15 nm Pd by electron beam evaporation, and then lift off to form a pattern to serve as the gate of the CNT and IGZO transistors.
[0096] Step 212: Atomic layer deposition is used to grow 10 nm HfO 2 as the gate oxide layer of the two top-gate transistors.
[0097] Step 213: Use photolithography and dry etching processes to selectively etch HfO2 to open holes in the gate oxide layer.
[0098] Step 214: Deposit a layer of carbon nanotubes using a wet transfer method.
[0099] Step 215 : Using photolithography and electron beam evaporation to deposit 45 nm of Pd, and then lift off to form a pattern to serve as the source and drain of the carbon nanotube transistor.
[0100] Step 216: Deposit 10 nm Y2O3 by electron beam evaporation as a passivation layer for the carbon nanotubes.
[0101] Step 217 : Using photolithography and oxygen plasma etching processes, selectively etch the carbon nanotubes to isolate different devices.
[0102] Step 218: Grow 10 nm IGZO using atomic layer deposition.
[0103] Step 219: Using photolithography and electron beam evaporation to deposit 30 nm Ti and 30 nm Pd, and then lift off to form patterns to serve as the source and drain of the IGZO transistor.
[0104] Step 220: Use photolithography and wet etching processes to selectively etch the IGZO to isolate different devices.
[0105] In at least one embodiment of the present disclosure, the method for preparing an integrated circuit device also includes preparing an interlayer passivation layer and preparing a plurality of conductive vias in the interlayer passivation layer, wherein the interlayer passivation layer is configured to separate the cache circuit layer, the in-memory computing circuit layer, and the logic control circuit layer from each other, and the plurality of conductive vias are configured to communicatively connect the cache circuit layer, the in-memory computing circuit layer, and the logic control circuit layer.
[0106] For example, in a specific example, preparing an interlayer passivation layer and preparing a plurality of conductive vias in the interlayer passivation layer may include the following steps:
[0107] Step 221: Deposit a 400nm SiO2 film (passivation layer) by plasma enhanced chemical vapor deposition.
[0108] Step 222: Use photolithography and dry etching processes to etch the SiO2 film to form openings for interconnection line connection points.
[0109] Step 223: deposit a layer of W by electroplating, and then use chemical mechanical polishing to grind away the W except for the SiO2 holes (forming conductive vias).
[0110] Step 224 : Using physical vapor deposition, deposit 400 nm of metal Al to form metal interconnects.
[0111] Step 225 : Using photolithography and dry etching processes, selectively etch Al to form Al metal interconnects.
[0112] Step 226: Plasma enhanced chemical vapor deposition is used to deposit a 1000 nm SiO2 film to form a passivation layer.
[0113] Step 227: Use photolithography and dry etching processes to selectively etch the SiO2 film to form openings.
[0114] Step 228: deposit a layer of W by electroplating, and then use chemical mechanical polishing to grind away the W except for the SiO2 holes to form conductive vias.
[0115] In at least one embodiment of the present disclosure, the method for preparing an integrated circuit device further includes: providing a silicon substrate to prepare a cache circuit layer, an in-memory computing circuit layer and a logic control circuit layer on the silicon substrate, wherein the logic control circuit layer is located on the silicon substrate, the in-memory computing circuit layer is located on a side of the logic control circuit layer away from the silicon substrate, and the cache circuit layer is located on a side of the in-memory computing circuit layer away from the silicon substrate.
[0116] For example, after the above steps of preparing the logic control circuit layer, preparing the cache circuit layer, preparing the in-memory computing circuit layer, preparing the interlayer passivation layer and preparing multiple conductive vias in the interlayer passivation layer are completed, subsequent passivation and metal interconnection processes can be used to form a metal interconnection pattern. The subsequent passivation process is used to protect the device, and the metal interconnection process can connect metal areas that can be used for testing.
[0117] It should be noted that in the methods for preparing integrated circuit devices provided in different embodiments of the present disclosure, the positions of the in-memory computing circuit layer and the cache circuit layer can be interchanged. For example, if the in-memory computing circuit layer is on the outermost side of the integrated circuit device, then when preparing the cache circuit layer, it is necessary to deposit another layer of Al2O3 after step S220 to protect the IGZO channel.
[0118] In at least one embodiment of the present disclosure, a logic control circuit layer is prepared on a silicon substrate by a silicon-based CMOS process, and a cache circuit layer and an in-memory computing circuit layer are prepared by a back-end integration process, wherein the operating temperature of the back-end integration process is lower than that of the silicon-based CMOS process.
[0119] For example, at least one embodiment of the present disclosure provides a method for fabricating an integrated circuit device, wherein the fabricated integrated circuit device mainly comprises three functional layers:
[0120] The first layer is the logic control circuit layer, manufactured using a standard silicon-based CMOS process. It features high performance, high reliability, and relatively mature technology. Its primary function is to control the overall chip functionality and provide neural network calculations other than matrix-vector multiplication (such as routing, layer skipping, rearrangement, activation functions, etc.). However, due to temperature limitations, a single chip can only implement one layer of silicon-based CMOS, so other back-end processes are stacked on top.
[0121] The second layer is an in-memory computing circuit layer based on resistive random access memory, which can be used to implement matrix-vector multiplication in neural networks. It can be stacked in multiple layers and is compatible with back-end processes.
[0122] The third layer is a cache circuit layer based on 2T0C DRAM, which is implemented using a back-end process and is used to cache large amounts of data in neural network calculations. The above three layers are vertically stacked through monolithic three-dimensional integration technology and achieve high-bandwidth communication through conductive vias.
[0123] At least one embodiment of the present disclosure provides an integrated circuit device that can implement matrix-vector multiplication operations using in-memory computing. For example, a cache circuit layer based on 2TOC DRAM can be used to implement data caching, and the aforementioned architecture is implemented using a monolithic three-dimensional integration approach. In the embodiments of the present disclosure, using a monolithic three-dimensional integration approach, data communication between the in-memory computing and cache modules can be carried out via efficient on-chip interconnects, resulting in high communication bandwidth. Furthermore, the presence of the cache circuit layer allows for high-capacity data throughput, improving the computational efficiency of the chip in processing neural networks and enabling high-speed, large-data neural network task computations.
[0124] It is understood that the above embodiments are merely exemplary embodiments for illustrating the principles of the present disclosure, and the present disclosure is not limited thereto. Those skilled in the art may make various modifications and improvements without departing from the spirit and substance of the present disclosure, and such modifications and improvements are also considered to be within the scope of protection of the present disclosure.
[0125] Regarding this disclosure, the following points need to be explained:
[0126] (1) The drawings of the embodiments of the present disclosure only relate to the structures related to the embodiments of the present disclosure. Other structures may refer to conventional designs.
[0127] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present disclosure, the thickness of layers or regions is exaggerated or reduced, that is, these drawings are not drawn according to the actual scale.
[0128] (3) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to form new embodiments.
[0129] The above description is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be based on the protection scope of the claims.< / m> < / n> < / m>
Claims
1. An integrated circuit device comprising: The logic control circuit layer is configured to control the data processing flow. The in-memory computing circuit layer is configured to perform matrix-vector computing tasks. A cache circuit layer, comprising a dynamic random access memory to store data in the matrix vector calculation task, wherein the dynamic random access memory is a capacitor-free dynamic random access memory, The cache circuit layer, the in-memory computing circuit layer and the logic control circuit layer are at least partially stacked on each other in a first direction and are communicatively connected with each other.
2. The integrated circuit device according to claim 1, further comprising an interlayer passivation layer for separating the cache circuit layer, the in-memory computing circuit layer, and the logic control circuit layer from each other, in, The interlayer passivation layer includes a plurality of conductive vias, and the cache circuit layer, the in-memory computing circuit layer and the logic control circuit layer are communicatively connected through the conductive vias.
3. The integrated circuit device according to claim 2, wherein: The interlayer passivation layer includes: a first passivation layer between the in-memory computing circuit layer and the cache circuit layer and a second passivation layer between the logic control circuit layer and the in-memory computing circuit layer, or a second passivation layer between the logic control circuit layer and the cache circuit layer, The plurality of conductive vias include: a plurality of first metal vias in the first passivation layer and a plurality of second metal vias in the second passivation layer.
4. The integrated circuit device according to any one of claims 1 to 3, wherein: The capacitor-less dynamic random access memory includes an N-type first transistor and a P-type second transistor, wherein the material of the active layer of the first transistor is different from the material of the active layer of the second transistor.
5. The integrated circuit device according to claim 4, wherein: The active layer of the first transistor comprises at least one of the following semiconductor oxides: IGZO, ITO, IWO and IAZO; The active layer of the second transistor includes at least one of the following materials: CNT, WSe2 and MoS2.
6. The integrated circuit device according to any one of claims 1 to 5, wherein: The in-memory computing circuit layer includes at least one in-memory computing circuit sub-layer and a plurality of in-memory computing arrays, and the plurality of in-memory computing arrays are located in the at least one in-memory computing sub-layer.
7. The integrated circuit device according to claim 6, wherein: Each of the multiple in-memory computing arrays includes at least one memristor array, and the memristor array includes a plurality of memristors arranged in multiple rows and columns.
8. The integrated circuit device according to any one of claims 1 to 7, wherein: The logic control circuit layer includes a logic control circuit and a peripheral circuit, wherein the logic control circuit is configured to control the overall function, and the peripheral circuit is configured to perform at least one of the following functions: routing, layer skipping, rearrangement, and activation function.
9. The integrated circuit device according to any one of claims 1 to 8, further comprising a silicon substrate, in, The logic control circuit layer is located on the silicon substrate, the in-memory computing circuit layer is located on a side of the logic control circuit layer away from the silicon substrate, and the cache circuit layer is located on a side of the in-memory computing circuit layer away from the silicon substrate.
10. A method for preparing an integrated circuit device according to any one of claims 1 to 9, comprising: preparing a logic control circuit layer, wherein the logic control circuit layer is configured to control a data processing flow, preparing an in-memory computing circuit layer, wherein the in-memory computing circuit layer is configured to perform matrix vector computing tasks, Prepare a cache circuit layer, wherein the cache circuit layer includes a dynamic random access memory DRAM to store data in the matrix vector calculation task, and the dynamic random access memory is a capacitor-free dynamic random access memory; The cache circuit layer, the in-memory computing circuit layer and the logic control circuit layer are at least partially stacked on each other in a first direction and are communicatively connected with each other.
11. The method for preparing an integrated circuit device according to claim 10, further comprising: preparing an interlayer passivation layer and preparing a plurality of conductive vias in the interlayer passivation layer, Among them, the interlayer passivation layer is configured to separate the cache circuit layer, the in-memory computing circuit layer and the logic control circuit layer from each other, and the multiple conductive vias are configured to communicatively connect the cache circuit layer, the in-memory computing circuit layer and the logic control circuit layer.
12. The method for preparing an integrated circuit device according to claim 10 or 11, further comprising: Providing a silicon substrate to prepare the cache circuit layer, the in-memory computing circuit layer and the logic control circuit layer on the silicon substrate, The logic control circuit layer is located on the silicon substrate, the in-memory computing circuit layer is located on a side of the logic control circuit layer away from the silicon substrate, and the cache circuit layer is located on a side of the in-memory computing circuit layer away from the silicon substrate.
13. The method for preparing an integrated circuit device according to claim 12, wherein: The logic control circuit layer is prepared on the silicon substrate by a silicon-based CMOS process, and the cache circuit layer and the in-memory computing circuit layer are prepared by a back-end integration process. Wherein, the operating temperature of the back-end integration process is lower than that of the silicon-based CMOS process.
Citation Information
Patent Citations
In-memory computing module and method, in-memory computing network and construction method
CN113704137A
Data processing apparatus and method, and method for manufacturing data processing apparatus
CN114139644A
Storage and calculation integrated chip, operation method, manufacturing method and electronic equipment
CN115831185A
Integrated circuit device and method of manufacturing the same
CN117641908A
Integrated circuit with vertically integrated passive variable resistance memory and method for making the same
US20130051117A1