Integrated circuit chip and preparation method of integrated circuit chip
Through the three-dimensional vertically integrated integrated circuit chip architecture, combined with the high-density storage characteristics of NAND flash memory and hybrid bonding technology, the problem of separation of storage and computing in traditional hardware architecture is solved, and the integrated solution of end-side computing and storage with high bandwidth, low latency and low cost is realized, which is suitable for end-side intelligent devices.
Patent Information
- Application Number
- CN202510380485.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional hardware architectures separated from computing and storage are difficult to meet the real-time and privacy protection needs of end-side smart devices. SRAM storage density is insufficient, and HBM solutions face problems such as large packaging area, high power consumption and high cost.
Adopting a three-dimensional vertically integrated integrated circuit chip architecture, including a storage array layer, a control circuit layer and a computing circuit layer, the NAND storage unit, control circuit and computing circuit are superimposed through hybrid bonding technology, and the high-density storage characteristics of NAND flash memory are utilized, combined with the packaging to shorten the data interaction path, and achieve high bandwidth and low latency data access.
While meeting the storage needs of the end-side artificial intelligence model, it significantly reduces packaging area, power consumption and manufacturing costs, provides an efficient hardware foundation, and supports real-time computing and low-latency data access.
Smart Images

Figure CN120264774A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to an integrated circuit chip and a method for manufacturing an integrated circuit chip. Background Art
[0002] With the development of artificial intelligence technology, the demand for local data processing in edge intelligent devices has grown. However, the traditional hardware architecture with separated computing and storage is difficult to meet the requirements of real-time performance and privacy protection. Although 3D V-Cache technology shortens the distance between storage and computing by three-dimensionally stacking SRAM (Static Random Access Memory) and computing chips, thereby improving the bandwidth, the capacity of SRAM is limited. While using high-bandwidth memory (HBM) increases the capacity and bandwidth, it is not suitable for edge devices due to issues such as power consumption, cost, and packaging volume. Summary of the Invention
[0003] At least one embodiment of the present disclosure provides an integrated circuit chip, including: a storage array layer including at least one storage array, where each of the at least one storage array includes a plurality of NAND (NAND Flash Memory) storage units arranged in multiple rows and columns; a control circuit layer configured to perform data read and write operations on the at least one storage array; and a computing circuit layer configured to send data read and write instructions to the control circuit layer to receive data in the at least one storage array and perform computing tasks; wherein the storage array, the control circuit layer, and the computing circuit layer are stacked relative to a first substrate, the storage array layer is disposed on the first substrate, the control circuit layer is disposed on a side of the storage array layer away from the first substrate, and the computing circuit layer is disposed on a side of the control circuit layer away from the first substrate.
[0004] For example, in the integrated circuit chip provided by an embodiment of the present disclosure, the control circuit layer includes a page buffer device to cache data of a target storage page in the at least one storage array; the page buffer device includes a sensing circuit.
[0005] For example, in the integrated circuit chip provided by an embodiment of the present disclosure, the control circuit layer further includes a bit line selector and a scan selector; the control circuit layer is further configured to dynamically activate the target storage page through the bit line selector and the scan selector.
[0006] For example, in an integrated circuit chip provided in an embodiment of the present disclosure, the integrated circuit chip further includes a second substrate; wherein, the control circuit layer is disposed on the second substrate, and on a side of the second substrate facing the first substrate, the computing circuit layer is disposed on a side of the second substrate facing away from the first substrate, the second substrate includes a first through-silicon via, and the control circuit layer and the computing circuit layer are interconnected through the first through-silicon via; the computing circuit layer is further configured to send a data read / write instruction to the control circuit layer through the first through-silicon via.
[0007] For example, in an integrated circuit chip provided in an embodiment of the present disclosure, the integrated circuit chip further includes a third substrate; wherein, the computing circuit layer is disposed on a side of the third substrate facing the second substrate; the third substrate includes a second through-silicon via and a conductive pad electrically connected to the second through-silicon via is disposed on a side facing away from the computing circuit layer.
[0008] For example, in an integrated circuit chip provided in an embodiment of the present disclosure, the computing circuit layer further includes a cache unit; the cache unit is configured to form a data path with the sensing circuit through the first through-silicon via.
[0009] For example, in an integrated circuit chip provided in an embodiment of the present disclosure, the storage array layer and the control circuit layer are actively aligned through a bonding technique to form a first bonding interface.
[0010] For example, in an integrated circuit chip provided in an embodiment of the present disclosure, the control circuit layer and the computing circuit layer are non-actively aligned through a bonding technique to form a second bonding interface.
[0011] At least one embodiment of the present disclosure provides a method for manufacturing an integrated circuit chip, including: preparing a storage array layer on a first substrate; preparing a control circuit layer on a side of the storage array layer away from the first substrate; and preparing a computing circuit layer on a side of the control circuit layer away from the first substrate; the storage array layer includes at least one storage array, and each of the at least one storage array includes a plurality of NAND storage units arranged in multiple rows and multiple columns; the control circuit layer is configured to perform data read / write operations on the at least one storage array; and the computing circuit layer is configured to send a data read / write instruction to the control circuit layer to receive data in the at least one storage array and perform a computing task.
[0012] For example, in the method for manufacturing an integrated circuit chip provided in an embodiment of the present disclosure, the manufacturing method further includes: providing a second substrate; wherein, the control circuit layer is disposed on the second substrate, and on a side of the second substrate facing the first substrate, the computing circuit layer is disposed on a side of the second substrate facing away from the first substrate, the second substrate includes a first through-silicon via, and the control circuit layer and the computing circuit layer are interconnected through the first through-silicon via; the computing circuit layer is further configured to send data read / write instructions to the control circuit layer through the first through-silicon via.
[0013] For example, in the method for manufacturing an integrated circuit chip provided in an embodiment of the present disclosure, the manufacturing method further includes: providing a third substrate; wherein, the computing circuit layer is disposed on a side of the third substrate facing the second substrate; the third substrate includes a second through-silicon via and a conductive pad electrically connected to the second through-silicon via is disposed on a side facing away from the computing circuit layer.
[0014] At least one embodiment of the present disclosure provides an electronic device, including the integrated circuit chip provided in any embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0016] Figure 1 FIG. shows a schematic block diagram of an integrated circuit chip provided in at least one embodiment of the present disclosure;
[0017] Figure 2 FIG. shows a schematic block diagram of a memory array layer in an integrated circuit chip provided in at least one embodiment of the present disclosure;
[0018] Figure 3 FIG. shows a schematic block diagram of a control circuit layer in an integrated circuit chip provided in at least one embodiment of the present disclosure;
[0019] Figure 4 FIG. shows a schematic block diagram of a computing circuit layer in an integrated circuit chip provided in at least one embodiment of the present disclosure;
[0020] Figure 5 FIG. shows a schematic diagram of a method for manufacturing an integrated circuit chip provided in at least one embodiment of the present disclosure;
[0021] Figure 6 FIG. shows an application schematic diagram of a method for manufacturing an integrated circuit chip provided in at least one embodiment of the present disclosure;
[0022] Figure 7The schematic diagram of the architecture of an integrated circuit chip provided by at least one embodiment of the present disclosure is shown; and
[0023] Figure 8 The schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure is shown. Detailed implementation manners
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0025] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second", and the like used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, the terms such as "a", "an", or "the" do not denote a quantity limitation, but mean that there is at least one. The terms such as "include" or "comprise" mean that the elements or items appearing before the term cover the elements or items listed after the term and their equivalents, without excluding other elements or items. The terms such as "connect" or "couple" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0026] With the rapid development of artificial intelligence technology, the demand for data local processing by edge intelligent devices (such as mobile phones, tablets, etc.) is becoming increasingly urgent. Edge artificial intelligence can directly complete data processing at the device end, reducing the latency and privacy leakage risks brought by cloud transmission, and becoming the main technology to promote the implementation of intelligent terminal applications. However, the data volume of artificial intelligence models increases exponentially, posing higher requirements on the storage capacity and bandwidth of the computing platform. In traditional hardware architectures, the design of physically separating the computing chip (such as a CPU or GPU) from the storage device (such as DRAM (Dynamic Random Access Memory), NAND flash (Flash), or disk) may result in a long data transmission path and limited bandwidth, thus making it difficult to meet the real-time computing requirements of large models at the edge.
[0027] To address the bottleneck of the separation between storage and computing, for example, an integration solution based on advanced packaging technology can be adopted. For instance, in the technical solution of three-dimensionally stacking SRAM and a computing chip using hybrid bonding technology (3D V-Cache technology), the bandwidth is increased by shortening the distance between the storage and computing circuit layers. However, the storage density of SRAM is relatively low (usually in the order of dozens of MB), which cannot meet the storage requirements of large amounts of data for end-side artificial intelligence models. Another example is the technical solution of 2.5D packaging of a high-bandwidth memory (HBM, High Bandwidth Memory) and a computing chip. Although the bandwidth and capacity can be significantly improved, the multi-layer DRAM stacking structure of HBM may result in a large packaging area, high power consumption, and high cost, making it difficult to adapt to end-side devices that are sensitive to size and energy efficiency.
[0028] The inventors of the present disclosure have noticed that in the design of the separation architecture between storage and computing, the storage bandwidth is insufficient to support the real-time inference of end-side large models; the on-chip cache capacity based on SRAM is limited and it is difficult to store large-scale model parameters; although DRAM integration solutions such as HBM improve the capacity and bandwidth, they face a significant increase in cost, power consumption, and packaging volume, which conflicts with the lightweight requirements of end-side devices. Therefore, there is an urgent need for an end-side computing and storage integration solution that combines high storage capacity, high bandwidth, low power consumption, small area, and controllable cost.
[0029] In view of this, an embodiment of the present disclosure provides an integrated circuit chip. The integrated circuit chip includes a storage array layer, a control circuit layer, and a computing circuit layer. The storage array layer includes at least one storage array, where each of the at least one storage arrays includes a plurality of NAND storage units arranged in multiple rows and columns; the control circuit layer is configured to perform data read and write operations on the at least one storage array; the computing circuit layer is configured to send data read and write instructions to the control circuit layer to receive data in the at least one storage array and execute computing tasks; the storage array, the control circuit layer, and the computing circuit layer are stacked relative to a first substrate, the storage array layer is disposed on the first substrate, the control circuit layer is disposed on a side of the storage array layer away from the first substrate, and the computing circuit layer is disposed on a side of the control circuit layer away from the first substrate. The integrated circuit chip can be based on a three-dimensional storage and computing chip architecture of NAND flash memory. By vertically integrating the storage array layer, the control circuit layer, and the computing circuit layer three-dimensionally, the high-density storage characteristics of NAND flash memory are fully utilized (for example, a single chip can reach hundreds of GB to TB level), and the data interaction path is shortened by combining packaging, so as to achieve high-bandwidth and low-latency data access. The integrated circuit chip can meet the storage requirements of, for example, end-side artificial intelligence models while significantly reducing the packaging area, power consumption, and manufacturing cost, providing an efficient hardware foundation for the deployment of end-side large models.
[0030] Moreover, embodiments of the present disclosure also provide a method for manufacturing the above-mentioned integrated circuit chip.
[0031] Figure 1 The schematic block diagram of an integrated circuit chip provided by at least one embodiment of the present disclosure is shown.
[0032] As Figure 1 shown, in some embodiments of the present disclosure, the integrated circuit chip 100 includes a memory array layer 101, a control circuit layer 102, and a computing circuit layer 103. The memory array layer 101 includes at least one memory array, and each of the at least one memory array includes a plurality of NAND memory cells arranged in multiple rows and columns. The control circuit layer 102 is configured to perform data read and write operations on the at least one memory array. The computing circuit layer 103 is configured to send data read and write instructions to the control circuit layer 102 to receive data in the at least one memory array and perform computing tasks. The memory array, the control circuit layer 102, and the computing circuit layer 103 are stacked relative to the first substrate, the memory array layer 101 is disposed on the first substrate, the control circuit layer 102 is disposed on the side of the memory array layer 101 away from the first substrate, and the computing circuit layer 103 is disposed on the side of the control circuit layer 102 away from the first substrate.
[0033] For example, the integrated circuit chip 100 may adopt a three-dimensional vertical stacking architecture, and the memory array layer 101, the control circuit layer 102, and the computing circuit layer 103 are integrated into a single package structure through hybrid bonding technology to achieve efficient cooperation between storage and computing.
[0034] For example, the first substrate may serve as the support substrate of the integrated circuit chip 100, and the first substrate may be a silicon substrate or other compatible semiconductor materials.
[0035] For example, an interconnect metal layer may be provided on the surface of the first substrate for electrical signal transmission and physical connection between the functional layers.
[0036] For example, the memory array layer 101 is disposed above the first substrate.
[0037] Figure 2 The schematic block diagram of the memory array layer in an integrated circuit chip provided by at least one embodiment of the present disclosure is shown.
[0038] As Figure 2 shown, the memory array layer 101 may include at least one memory array composed of a plurality of NAND memory cells.
[0039] For example, NAND memory cells in each memory array (i.e., NAND flash memory) can be arranged in a high density in multiple rows and columns. For example, a 3D NAND stacking structure (such as vertical channel or charge trap type cells, etc.) can be adopted, with single-layer or multi-layer stacking to achieve higher storage capacity (for example, the capacity of a single chip can reach hundreds of GB to the TB level).
[0040] For example, the layout of the NAND memory cell array is adapted to the miniaturization requirements of the end-side device. For example, the storage density can be improved by reducing the size of the NAND memory cells or optimizing the pitch of the NAND memory cells.
[0041] For example, the memory array layer 101 can be electrically connected to the underlying first substrate and the upper circuit on the side away from the first substrate through through-silicon vias or micro-bumps in the hybrid bonding interface.
[0042] For example, the control circuit layer 102 is disposed on the side of the memory array layer 101 away from the first substrate (i.e., above the memory array layer 101).
[0043] For example, the control circuit layer 102 can include a memory controller for the NAND memory array, an address decoder, a read / write drive circuit, or an error correction coding (ECC) module.
[0044] For example, the memory controller can be configured to receive read / write instructions from the computing circuit layer 103 and generate corresponding row select signals, column select signals, and voltage control signals according to the instructions to drive the read / write operations of specific memory cells in the memory array.
[0045] For example, the control circuit layer 102 can be directly interconnected with the memory array layer 101 through hybrid bonding. For example, copper-copper bonding or oxide-oxide bonding technology can be adopted to form high-density, low-impedance vertical interconnection channels, significantly shortening the data access path and improving the bandwidth (for example, up to hundreds of GB per second).
[0046] For example, the computing circuit layer 103 is disposed on the side of the control circuit layer 102 away from the first substrate (i.e., above the control circuit layer 102). The computing circuit layer 103 can include artificial intelligence computing units (such as NPU, GPGPU, or dedicated accelerator cores), instruction scheduling modules, or data cache units, etc.
[0047] For example, the computing circuit layer 103 can send data request instructions (such as reading model parameters or writing intermediate results) to the control circuit layer 102 through an internal bus or a network-on-chip (NoC). For example, it can directly obtain raw data from the memory array layer 101 or write the calculation results, bypassing the long-distance data transmission of off-chip memory in the traditional architecture.
[0048] For example, the computing circuit layer 103 can achieve high-bandwidth interconnection with the control circuit layer 102 through a hybrid bonding interface, enabling the computing circuit layer 103 to access data in the storage array with extremely low latency. For example, it supports the fast loading of weights and the dynamic update of activation values in real-time inference tasks.
[0049] For example, the storage array layer 101, the control circuit layer 102, and the computing circuit layer 103 are stacked vertically. The hybrid bonding technology is used to tightly integrate the storage, control, and computing units. The transmission distance of data from the storage array layer 101 to the computing circuit layer 103 is shortened to the micron level, significantly reducing the transmission latency and power consumption, and achieving vertical integration and signal path optimization.
[0050] For example, the control circuit layer 102 can be configured to drive the read and write operations of multiple storage arrays simultaneously. Combining with the multi-core parallel computing ability of the computing circuit layer 103, it realizes the matching of storage bandwidth and computing power, meets the hierarchical parameter loading and pipeline computing requirements of end-side large models, and achieves parallel data access.
[0051] For example, the high-density characteristic of NAND storage cells enables the storage array layer 101 to provide ultra-large capacity within a limited area, solving the problems of area expansion and cost explosion caused by the stacking of multiple layers of DRAM in the HBM solution. At the same time, the hybrid bonding technology reduces the dependence on the interposer, further reducing the packaging complexity and achieving physical area and cost control.
[0052] In a possible implementation manner, for example, when the end-side device runs a Transformer-class large model, the computing circuit layer 103 can be configured to load the weight matrix in blocks from the storage array layer 101 on demand. The control circuit layer 102 can be configured to pre-schedule adjacent data blocks in advance through an address mapping and prefetch mechanism. The computing circuit layer 103 can be configured to obtain the next data block immediately after completing the operation of the current block, realizing "near-storage computing" of model parameters. The overall energy efficiency ratio is more than 5 times higher than that of the traditional separated architecture.
[0053] Figure 3 The schematic block diagram of the control circuit layer in an integrated circuit chip provided by at least one embodiment of the present disclosure is shown.
[0054] As Figure 3 shown, in some embodiments of the present disclosure, the control circuit layer 102 includes a page buffer device to cache data of a target storage page in at least one storage array.
[0055] For example, a page buffer device is integrated in the control circuit layer 102 to efficiently cache data of the target storage page in the storage array layer 101 and achieve fast reading and temporary storage of the data.
[0056] For example, the page buffer device may include a sensing circuit, such as a sense amplifier (SA).
[0057] For example, the page buffer device may include the following sub-modules: a sensing circuit, a differential amplification module, a data latch module, a timing control module, a data cache unit, or a logic control module, etc.
[0058] For example, each sensing circuit may correspond to one or more bit lines in a NAND memory array, and may be configured to detect the charge state of a target memory cell during a read operation and convert it into a digital signal.
[0059] For example, the differential amplification module may be configured to amplify the logic state of the memory cell (e.g., distinguish between "0" or "1") by comparing the difference between the bit line and a reference voltage.
[0060] For example, the data latch module may be configured to temporarily store the converted data until the reading of an entire page is completed.
[0061] For example, the timing control module may be configured to work in cooperation with a row / column decoder, and trigger corresponding sensing actions according to the timing of the word line and the layer select signal.
[0062] For example, the data cache unit may be configured to temporarily store the complete data of the current page (e.g., 4KB or 16KB), and may also support a burst transfer mode to batch transfer the entire page of data to the computing circuit layer 103 through a high-bandwidth interface.
[0063] For example, the logic control module may be configured to receive read / write instructions from the computing circuit layer 103, coordinate the activation timing of the SA, manage the read / write pointers of the data cache unit, and handle operations such as error correction coding (ECC) verification.
[0064] For example, the number of columns in each memory array (Column) may be configured to be an integer multiple of the number of SAs. Exemplarily, if the memory array contains 1024 columns, the SAs can be configured as 128 (each SA processes 8 columns) or 1024 (each SA processes 1 column).
[0065] For example, by maximizing the number of SAs (e.g., equal to the number of columns in the memory array), it is possible to achieve parallel reading of all columns of an entire page within a single cycle, thereby boosting the memory bandwidth to the theoretical peak (e.g., when each SA supports a rate of 1 Gbps, 1024 SAs can achieve a bandwidth of 1 Tbps).
[0066] For example, Bitline Selectors can work in conjunction with Control Gate Selectors to activate corresponding bitline groups and wordlines according to the physical address of the target storage page (such as layer number, wordline number), and transfer the data of the target page to the associated SA.
[0067] For example, in a 3D NAND array, a specific layer in the vertical stack can be determined by a layer selector, and then the target bitline in the horizontal direction can be located by a bitline selector, finally locking the storage cell page to be read.
[0068] For example, the control circuit layer 102 can be configured to be directly interconnected with the upper computing circuit layer 103 through a hybrid bonding interface, and the data cached in the page buffer device can be transmitted to the registers or L2 cache of the computing cores in the computing circuit layer 103 through vertical interconnection channels (such as microbumps or TSV (Through-Silicon Via)) with extremely low latency.
[0069] For example, during the execution of a computing task, the page buffer device can be configured to prefetch the model parameters or intermediate data required for the next computing stage to achieve pipelined processing of "computing - transmission - storage", so as to reduce the idling of the computing units in the computing circuit layer 103 due to data waiting.
[0070] For example, if the storage array layer 101 includes multiple independent NAND storage arrays, each storage array can be configured with a dedicated page buffer device, and the number of SAs matches the number of columns of the corresponding array. Exemplarily, for an architecture including 4 storage arrays, the control circuit layer 102 can integrate 4 groups of page buffer devices, each group containing 1024 SAs, and the total bandwidth is increased to 4 Tbps.
[0071] For example, the number of SAs and the page buffer capacity can also be dynamically adjusted according to the application scenario of the edge device, and the embodiments of the present disclosure do not limit this. For example, for a lightweight model, some SAs can be turned off to reduce power consumption; for large model inference, all SAs can be enabled to maximize the bandwidth.
[0072] For example, the input end of each SA can be directly connected to the bitline of the storage array through metal interconnections. If one SA is responsible for multiple columns of bitlines (for example, 1 SA processes 8 columns), the target bitline group can be dynamically gated through bitline selectors to form a multiplexing structure.
[0073] For example, in a 3D NAND structure, the SA can receive a hierarchical selection signal through layer selectors (Control Gate Selectors) to distinguish memory cells in different layers in the vertical stack.
[0074] For example, the output terminal of each SA can be connected to the corresponding memory bit of the data cache unit through a parallel bus. For example, 1024 SAs correspond to a 1024-bit wide bus, and a whole page of data (such as 4KB) can be written into the cache in a single cycle.
[0075] For example, the output terminal of the data cache unit can be connected to the L2 cache or register of the computing circuit layer 103 through a vertical interconnection channel formed by bonding technology (such as copper-copper microbumps or TSVs), supporting high-bandwidth burst transmission (such as transmitting 512 bits of data per cycle).
[0076] For example, the logic control module can be configured to receive read and write commands of the computing circuit layer 103 through an instruction bus, and generate corresponding control signals after parsing. For example, the corresponding control signals can include row / column decoding signals, SA enable signals, or layer selection signals.
[0077] For example, the row / column decoding signal can be used to drive the row decoder (Wordline Driver) and column decoder (Bitline Driver) to select the target memory cell.
[0078] For example, the SA enable signal can be used to control the start timing and working mode of the SA (such as reading, writing, or erasing).
[0079] For example, the layer selection signal can be used to specify the target layer in the 3D stack.
[0080] For example, the bitline selector can adopt a tree-like multiplexing structure (such as 1:8 or 1:16). When the number of SAs is less than the number of columns of the memory array, the resource utilization rate can be improved through time-division multiplexing.
[0081] For example, low-impedance metal wiring (such as copper interconnection) can be used between the selector and the bitline to reduce signal attenuation to achieve a low-impedance design.
[0082] For example, the bus width can be aligned with the number of SAs to achieve bandwidth matching of the vertical interconnection. Exemplarily, if the memory array contains 1024 columns and the number of SAs is 1024, the vertical interconnection bus is designed to be 1024-bit wide, so that a whole page of data can be transmitted in a single cycle.
[0083] For example, the data cache unit and the computing circuit layer 103 can adopt a source synchronous clock (such as a DDR interface) to further reduce the transmission delay difference.
[0084] For example, the sensing circuit for the storage array can be divided into multiple power supply clusters, and the inactive clusters can enter the low-power mode to achieve clustered power supply for the sensing circuit.
[0085] In a possible implementation, take the control circuit layer 102 including 4 NAND storage arrays as an example. For example, each storage array includes 1024 column bitlines and is configured with 1024 SAs (1 SA / column). For example, the data cache unit includes 4 groups of independent caches, each with a capacity of 4 KB, and is connected to the computing layer through 4 1024-bit wide buses. For example, the logic control module is configured to schedule the read and write operations of the 4 storage arrays through the controller in a time-sharing manner to support parallel data access (total bandwidth up to 4×1024 bits / cycle).
[0086] In a possible implementation, the exemplary operation process of the integrated circuit chip 100 provided by the embodiments of the present disclosure may include: when the computing circuit layer 103 requests to load the neural network weights of a certain layer, the computing management unit sends a read instruction including the target storage address to the control circuit layer 102; the logic control module parses the instruction and drives the row decoder and the layer selector to locate the target storage page; the bitline selector activates the corresponding bitlines, and the sensing circuit senses all column data in parallel and latches it into the page buffer; after the data in the page buffer is verified by ECC, it is transmitted to the L2 cache of the computing core through the vertical interconnection channel; the computing circuit layer 103 reads the data from the cache and performs matrix operations, and at the same time, the page buffer prefetches the next batch of weight data. This enables the edge device to achieve a balance among energy efficiency, cost, and performance, providing a hardware foundation for deploying large models with tens of billions of parameters.
[0087] In at least one embodiment of the present disclosure, for example, through the parallel configuration of the sensing circuit, the data transfer rate between the storage array layer 101 and the computing circuit layer 103 breaks through the bottleneck of traditional off-chip memories, meeting the real-time computing requirements of edge large models. For example, through the bonding technology, the physical path of data from the storage unit to the computing circuit layer 103 is shortened, and combined with the fast response ability of the sensing circuit, the energy consumption and latency of data access are significantly reduced. For example, through vertical stacking, high-density storage and high computing power are coordinated within a limited area, reducing the additional area overhead of the interposer in 2.5D packaging.
[0088] In some embodiments of the present disclosure, the control circuit layer 102 further includes a bitline selector and a scan selector. The control circuit layer 102 is further configured to dynamically activate the target storage page through the bitline selector and the scan selector.
[0089] In some embodiments of the present disclosure, the control circuit layer 102 further integrates bitline selectors and scan selectors (also known as layer selectors / control gate selectors), which work together to achieve dynamic activation and precise access to target memory pages. This design significantly improves memory access efficiency and flexibility through signal gating and timing control in the spatial dimension, and an exemplary implementation is as follows:
[0090] For example, the bitline selectors can gate a specific group of bitlines in the target memory array in the horizontal direction (X-Y plane), controlling the transmission path of data from the memory cells to the bitline sensing circuit.
[0091] For example, the bitline selectors can support on-demand switching of the activated bitline range, such as only conducting the bitlines corresponding to the current page during a read operation, reducing ineffective power consumption.
[0092] For example, the scan selectors can gate a specific memory layer in the 3D NAND stack structure in the vertical direction (Z-axis), activating the wordlines in the target layer through the control gate voltage.
[0093] For example, the scan selectors can support cross-layer scanning. For example, during multi-page consecutive reads, the activated layer numbers are switched in sequence to improve throughput.
[0094] For example, the bitline selectors can adopt a multi-level tree-shaped switch network (such as a 1:N multiplexer), and each selector unit is connected to a group of bitlines (for example, 8 bitlines). For example, the selector switches can include low-impedance transistors, which are arranged at the edge of the memory array and are directly connected to the bitlines through metal wiring (such as copper interconnects).
[0095] For example, the bitline selectors can be configured to receive column address decoding signals from the logic control module and generate corresponding bitline gating enable signals (BL_EN).
[0096] For example, if the memory array contains 1024 columns of bitlines, the bitline selectors can divide them into 128 groups (8 columns per group) and dynamically gate the target group through a 7-bit address signal (2^7 = 128).
[0097] For example, the scan selectors can be integrated at the vertical stacking interface of the memory array. For example, the scan selectors can include a layer decoding circuit and a high-voltage driving module.
[0098] For example, each scan selector can correspond to a memory layer and is connected to the control gate (CG) of that layer through an independent control line.
[0099] For example, the scan selector can be configured to receive a layer address signal (LayerAddress) and a timing control signal (Timing Ctrl) from the logic control module, and generate a layer-specific high-voltage pulse (such as a 20V programming voltage or a 5V read voltage).
[0100] For example, in a 64-layer 3D NAND, the layer address is 6 bits (2^6 = 64), and the scan selector can activate the CG signal of the corresponding layer according to the address.
[0101] For example, the logic control module can be configured to receive the storage page address (including layer number, word line number, and block number) sent by the computing circuit layer 103, and decompose it into a layer address (Layer Address), a row address (Row Address), and a column address (Column Address).
[0102] For example, the layer address can be used to drive the scan selector to activate the target storage layer.
[0103] For example, the row address can be used to drive the wordline driver to select the target word line (WL).
[0104] For example, the column address can be used to drive the bit line selector to conduct the target bit line group.
[0105] For example, the scan selector can be configured to preferentially activate the target layer at the initial stage of the operation. After the control gate voltage is stabilized, the bit line selector and the wordline driver are triggered to reduce the cross-layer signal interference.
[0106] For example, if the number of groups of the bit line selector (such as 128 groups) is less than the total number of columns (such as 1024 columns), different bit line groups are sequentially selected through time-division multiplexing, and the full-page reading is realized in cooperation with the pipelined operation of the sensing circuit to achieve bit line time-division multiplexing.
[0107] For example, for the dynamic page switching process (taking reading as an example), for example, in the pre-charge stage, the scan selector can turn off the high-voltage signals of all layers, the bit line selector disconnects the bit line connection, and the storage array enters the pre-charge state.
[0108] For example, the scan selector can apply a read voltage (such as Vread) to the target layer according to the layer address, and the wordline driver selects the target word line to achieve layer and word line activation.
[0109] For example, the bit line selector can conduct the target bit line group according to the column address, and the sensing circuit senses the bit line signals in parallel and latches the data to achieve bit line selection and sensing.
[0110] For example, for continuous cross-page access, when adjacent memory pages need to be read, the scan selector can maintain the current layer activation state, and the row decoder increments the word line number; when cross-layer access is required, the layer address can be reconfigured and the above process can be repeated.
[0111] For example, through the coordinated control of the bit line selector and the scan selector, three-dimensional addressing (X-Y-Z coordinates) of 3D NAND memory cells is achieved, solving the address conflict problem of traditional planar architectures. For example, the dynamic activation mechanism only powers the bit line layer and the memory layer associated with the target page, and the inactive area enters the low-power state, reducing the overall power consumption by 30%-50%. For example, by increasing the number of groups of the bit line selector (such as increasing from 128 groups to 256 groups) or the number of parallel layers of the scan selector (such as activating 2 layers simultaneously), the memory bandwidth can be linearly increased to adapt to the needs of different-scale edge models. For example, the scan selector can adopt isolation drive technology (such as differential high-voltage drive) to suppress inter-layer crosstalk; the bit line selector can further integrate noise-shielding wiring to reduce signal coupling.
[0112] In one possible implementation, for example, when the edge device executes an image recognition task, the computing circuit layer 103 needs to continuously load convolution kernel weights and feature map data. For example, in the weight loading stage, the control circuit layer 102 can activate a specific layer (such as Layer 5) that stores weight parameters through the scan selector, the bit line selector can select the corresponding bit line group according to the convolution kernel size, and the sensing circuit can batch-transfer the weight data to the computing core. For example, in the feature map writing stage, the scan selector can be switched to the layer that caches intermediate results (such as Layer 12), and the bit line selector dynamically expands the selection range to match the feature map size to achieve high-speed writing. The dynamic page switching latency is less than 10 ns, supporting millions of page access requests per second, meeting the processing requirements of 30 fps real-time video streams.
[0113] In at least one embodiment of the present disclosure, through the dynamic cooperation of the bit line selector and the scan selector, efficient and precise access to the 3D NAND memory array is achieved. Combining the physical advantages of the vertical stacking architecture, a high-bandwidth, low-power, and scalable storage-computing integrated solution is provided for edge artificial intelligence.
[0114] In some embodiments of the present disclosure, the integrated circuit chip 100 further includes a second substrate.
[0115] The control circuit layer 102 is disposed on the second substrate and on the side of the second substrate facing the first substrate, the computing circuit layer 103 is disposed on the side of the second substrate facing away from the first substrate, the second substrate includes a first through-silicon via, and the control circuit layer 102 and the computing circuit layer 103 are interconnected through the first through-silicon via. The computing circuit layer 103 can be further configured to send data read / write instructions to the control circuit layer 102 through the first through-silicon via.
[0116] For example, in at least one embodiment of the present disclosure, Figure 1 As shown, the storage array layer 101 can be bonded face-to-face with the control circuit layer 102, that is, the active surfaces of the two chips (dies) of the storage array layer 101 and the control circuit layer 102 are opposite; the control circuit layer 102 can be bonded face-to-back with the computing circuit layer 103, that is, the control circuit layer 102 and the computing circuit layer 103 are non-actively bonded.
[0117] For example, the integrated circuit chip 100 may adopt a dual-substrate three-dimensional stacking architecture, and by introducing a second substrate and through silicon via (TSV) technology, further optimize the interconnection efficiency between the control circuit layer 102 and the computing circuit layer 103, while enhancing the scalability and signal integrity of the chip.
[0118] For example, the second substrate may be a silicon substrate or other high-density interconnect substrates.
[0119] For example, a first through silicon via (TSV) may be integrated inside the second substrate, and copper or other low-impedance conductive materials may be filled inside the first through silicon via to form a vertical electrical channel that penetrates the substrate. For example, the second substrate may serve as a carrier and interconnection medium between the control circuit layer 102 and the computing circuit layer 103, and has both mechanical support and signal relay functions.
[0120] For example, the control circuit layer 102 may be fabricated on a side of the second substrate facing the first substrate (ie, a direction close to the memory array layer 101 ), and may be connected to the memory array layer 101 via a hybrid bonding interface.
[0121] For example, the computing circuit layer 103 may be fabricated on a side of the second substrate facing away from the first substrate (ie, in a direction away from the storage array layer 101 ), and interconnected with the control circuit layer 102 through a first through silicon via.
[0122] Figure 4 A schematic block diagram of a computing circuit layer in an integrated circuit chip provided by at least one embodiment of the present disclosure is shown.
[0123] like Figure 4 As shown, for example, the GPU architecture (such as computing core, L2 cache, interface unit, etc.) of the computing circuit layer 103 can adopt a high-performance process (such as 5nm FinFET, etc.) and achieve heterogeneous integration with the mature process (such as 7nm, 14nm or 28nm, etc.) of the lower control circuit layer 102 through silicon vias.
[0124] For example, first through silicon vias (TSVs) may be distributed in the second substrate in an array form, and the through hole spacing is optimized according to the signal type (for example, a data bus uses high-density TSVs, and a control signal uses sparse distribution).
[0125] For example, the diameter of each TSV can be configured to be less than 5 μm, the depth can be configured to match the thickness of the second substrate (such as 50 μm), and the impedance can be configured to be below 10 mΩ to support the transmission of high-frequency signals at the GHz level.
[0126] For example, in a possible transmission path of data read / write instructions, the computing management unit of the computing circuit layer 103 can generate data read / write instructions according to task requirements (such as reading model weights or writing intermediate results). The instructions include the target storage address (layer number, word line number, column number) and the operation type (read / write). The instructions can be transmitted to the TSV interface controller of the second substrate through the on-chip network (NoC) or a dedicated instruction bus. The controller allocates the corresponding TSV channel according to the address mapping table. The instruction signal can penetrate the second substrate from the computing circuit layer 103 through the conductive path of the TSV and reach the logic control module of the lower control circuit layer 102. The logic control module parses the instructions and can drive the row decoder, bit line selector, and scan selector to activate the target storage page according to the received instructions. The read data can enter the page buffer device of the control circuit layer 102 from the storage array layer 101 through the hybrid bonding interface and then be transmitted back to the L2 cache or register of the computing circuit layer 103 through the TSV.
[0127] For example, the second substrate can be used as a thermal buffer layer, and the TSV array of the second substrate can be further integrated with cooling channels (such as microfluidic cooling channels, etc.) to export the high heat of the computing circuit layer 103, so as to reduce the influence of thermal coupling on the stability of the storage unit.
[0128] In a possible implementation, for example, when the edge device runs a natural language processing (NLP) model, the computing management unit can send a batch read instruction to the control circuit layer 102 through the TSV to preload the model parameters from the storage array into the L2 cache. For example, the tensor computing unit can read data from the cache to perform matrix multiplication inference operations, and the intermediate results are written back to the specified area of the storage array in real time through the TSV. For example, when the interface unit interacts with other chips through off-chip communication, the computing management unit can independently access the storage array through the TSV to achieve task parallelism of computing and communication.
[0129] In at least one embodiment of the present disclosure, by introducing the second substrate and the through-silicon via technology, the high-density three-dimensional interconnection between the control circuit layer 102 and the computing circuit layer 103 is realized. Combining the computing core of the GPU architecture and the vertical integration of the NAND storage array, an edge artificial intelligence chip with ultra-high computing power, ultra-large storage capacity, and ultra-low power consumption is constructed, providing a hardware foundation for deploying large models with tens of billions of parameters on lightweight devices, and at the same time supporting future higher performance requirements through modular expansion.
[0130] In some embodiments of the present disclosure, the integrated circuit chip 100 further includes a third substrate.
[0131] The computing circuit layer 103 is disposed on a side of the third substrate facing the second substrate. The third substrate includes second through-silicon vias and has conductive pads electrically connected to the second through-silicon vias disposed on a side facing away from the computing circuit layer 103.
[0132] For example, the third substrate is located on a side of the computing circuit layer 103 facing away from the second substrate, that is, the computing circuit layer 103 is disposed on a side of the third substrate facing the second substrate.
[0133] For example, conductive pads are disposed on the other surface of the third substrate facing away from the computing circuit layer 103, which can be used for electrical and mechanical connections with external circuits (such as a PCB board, other chip modules, or a heat dissipation structure).
[0134] For example, high-speed data interaction between the computing circuit layer 103 and an external system can be achieved through the second through-silicon vias.
[0135] For example, the conductive pads integrate power input pins and a ground plane, and can also serve as a heat conduction path to export the heat generated by the computing circuit layer 103.
[0136] For example, the second through-silicon vias (TSVs) can penetrate the third substrate in the form of a high-density array, and the via pitch is designed differently according to signal types (for example, the via pitch in the data bus area ≤ 10 μm, and the pitch in the power / ground area ≤ 20 μm).
[0137] For example, the diameter range of the second through-silicon vias can be configured to be 1 - 5 μm, the depth can be configured to match the thickness of the third substrate (for example, 50 - 100 μm), and the internal filling material can include highly conductive materials such as copper or tungsten.
[0138] For example, one end of the second through-silicon via can be connected to the interface unit (such as the PHY layer) of the computing circuit layer 103, and the other end can be connected to the conductive pad, forming a vertical transmission channel from the computing core to the external system.
[0139] For example, the conductive pads can be made of copper or aluminum materials, and the surface can be gold-plated or nickel-plated to prevent oxidation.
[0140] For example, the size of a single conductive pad can be configured to be 50 μm × 50 μm, and the pitch complies with industry standards (such as a 100-μm pitch BGA package).
[0141] For example, the array layout of the conductive pad is divided into functional blocks: data I / O area, power / ground area or control signal area, etc., and the embodiments of the present disclosure are not limited to this. For example, the data I / O area can be used for high-speed differential pair pads (such as PCIe, DDR interface). For example, the power / ground area can be used for large-area power grids and low-impedance ground pads. For example, the control signal area can be used for low-speed single-ended signal pads (such as I2C, SPI).
[0142] For example, the interface module of the computing circuit layer 103 may be a TSV-compatible architecture, and its output driver may be directly connected to the second through-silicon via.
[0143] For example, the GPU's off-chip communication interface (such as GDDR6 PHY) can be directly connected to the conductive pad through TSV, for example, to support a transmission rate of up to 16 Gbps / pin.
[0144] For example, data generated by the computing core in the computing circuit layer 103 can be transmitted to the interface unit through the on-chip network (NoC), and vertically transmitted to the conductive pad via TSV, and the path length can be shortened to the millimeter level (the path in traditional packaging is, for example, several centimeters).
[0145] In at least one embodiment of the present disclosure, by introducing the third substrate and the second through silicon via, a three-dimensional heterogeneous integrated system is constructed from the storage array layer 101, the control circuit layer 102 and the computing circuit layer 103. In addition, the design of the conductive pad not only opens up an efficient interface between the chip and the external system, but also breaks through the performance and energy efficiency bottleneck of traditional end-side devices through electrical-thermal coordinated optimization.
[0146] In some embodiments of the present disclosure, the computing circuit layer 103 further includes a cache unit. The cache unit is configured to form a data path with the sensing circuit through the first through silicon via.
[0147] For example, the computing circuit layer 103 further integrates a cache unit (such as an L2 cache or a dedicated data buffer) and establishes a high-bandwidth, low-latency vertical data path with the sensing circuit of the control circuit layer 102 through the first through silicon via (TSV) to achieve data interaction between the storage array and the computing core.
[0148] For example, the sensing circuit may be located in the page buffer device of the control circuit layer 102, and may be configured to convert the analog signal in the storage array into digital data. For example, the converted data may vertically penetrate the second substrate through the first through silicon via (TSV) and be transmitted to the cache unit of the computing circuit layer 103. For example, after receiving the data, the cache unit (such as L2 cache) may distribute the data to the arithmetic logic unit (ALU) or tensor computing unit according to the needs of the computing core, or temporarily store the intermediate computing results, etc.
[0149] For example, the number of TSVs can match the number of bitline sensing circuits. Exemplarily, if each memory array includes 1024 SAs, 1024 TSV channels can be correspondingly set to form a 1:1 parallel data bus, so as to improve the complete transmission of the entire page data (such as 4KB) within a single cycle.
[0150] For example, electrical connection can be achieved between the control circuit layer 102 and the computing circuit layer 103 through bonding. For example, the microbumps at the bonding interface can be aligned with the TSVs to reduce signal reflection and impedance mismatch.
[0151] In a possible implementation manner, for example, in edge-side large model inference, during the model initialization phase, the cache unit can batch load the weight parameters in the memory array to the L2 cache through the TSVs, and utilize the high bandwidth characteristic to achieve model loading within seconds. For example, for real-time pipelined computing, when the computing core processes the activation values of the current layer, the cache unit can prefetch the weights of the next layer and maintain the transmission through the TSVs, eliminating the waiting delay of the computing unit and increasing the throughput by more than 30%.
[0152] In a possible implementation manner, for example, in image processing acceleration, the intermediate feature maps of a convolutional neural network (CNN) can be written back to the memory array in real time through the TSVs, releasing the cache space and supporting multi-task sharing of feature data at the same time. For example, dynamic resolution adaptation can be performed. According to the input image resolution, the cache unit can dynamically adjust the number of TSV channels (for example, expand from 512 channels to 1024 channels) to match the transmission requirements of feature maps of different sizes.
[0153] In at least one embodiment of the present disclosure, through the direct connection design of the cache unit and the sensing circuit by TSVs, and in combination with three-dimensional stacking and hybrid bonding technologies, a memory-computation integrated solution with performance, energy efficiency and reliability is provided for edge-side artificial intelligence.
[0154] In some embodiments of the present disclosure, the memory array layer 101 and the control circuit layer 102 are actively aligned through bonding technology to form a first bonding interface.
[0155] For example, in the embodiments of the present disclosure, the bonding technology may include adopting hybrid bonding, fusion bonding, transfer bonding, eutectic bonding technology, etc., and the embodiments of the present disclosure are not limited thereto.
[0156] For example, through a bonding technique, the active surface of the storage array layer 101 (i.e., the surface including the NAND storage cell array) can be face-to-face bonded with the active surface of the control circuit layer 102 (i.e., the surface integrating circuits such as column decoders, row decoders, sense amplifiers, etc.).
[0157] For example, copper-copper bonding points and oxide-oxide bonding regions can be provided at the first bonding interface. For example, copper interconnect points (e.g., with a diameter ≤ 1μm) can be precisely aligned with the bit lines (BL), word lines (WL), DSG lines, SSG lines, and common source lines of the storage array and the corresponding drive circuit pads of the control circuit layer 102 to form a low-impedance electrical connection.
[0158] For example, active alignment can achieve a micron-level interconnect pitch (e.g., 0.5μm) through hybrid bonding, supporting tens of millions of interconnect points integrated on a single bonding interface, thereby meeting the precise docking of the bit lines (thousands to tens of thousands) of the 3D NAND storage array with other circuits.
[0159] For example, through the active alignment of the storage array layer 101 and the control circuit layer 102, the distance between the storage cells and the control circuit layer 102 is reduced to the sub-micron level, the bit line resistance is reduced by 50%, the word line drive delay is reduced to the nanosecond level, and the read / write speed is significantly improved (e.g., the read delay < 10ns).
[0160] For example, before bonding, the storage array layer 101 and the control circuit layer 102 can respectively adopt independent process nodes (e.g., the storage layer uses a 20nm 3D NAND process, and the control layer uses a 28nm CMOS process. The embodiments of the present disclosure do not limit the specific process), and work together after bonding to reduce the manufacturing cost.
[0161] For example, chemical mechanical polishing (CMP) can be performed on the first bonding interface of the storage array layer 101 and the control circuit layer 102 to make the surface roughness of the first bonding interface < 0.5nm.
[0162] For example, the temperature can be raised to above 400°C in an inert gas environment to make the copper interconnect points diffuse and fuse, and the oxide layer forms covalent bonds to achieve the active alignment of the storage array layer 101 and the control circuit layer 102.
[0163] In some embodiments of the present disclosure, the control circuit layer 102 and the computing circuit layer 103 achieve non-active alignment of the control circuit layer 102 through a bonding technique to form a second bonding interface.
[0164] For example, the control circuit layer 102 and the computing circuit layer 103 can form a second bonding interface through a non-active alignment bonding technique and combine through-silicon vias (TSV) to achieve cross-layer signal transmission.
[0165] For example, bonding techniques that are the same as or different from those used to form the first bonding interface described above can be employed to bond the non-active surface (i.e., the back surface of the substrate) of the control circuit layer 102 to the active surface (i.e., the surface including the computing core) of the computing circuit layer 103.
[0166] For example, through-silicon vias (TSVs) or interlayer vias (ILVs) (e.g., with a diameter ≤ 5 μm) can be prefabricated in the control circuit layer 102 to penetrate the thinned substrate (e.g., with a thickness ≤ 50 μm), and highly conductive materials such as copper or tungsten are filled in the vias.
[0167] For example, one end of a TSV can be connected to the relevant circuits (such as page buffer devices, logic control units, etc.) of the control circuit layer 102, and the other end can be interconnected with the interface unit (such as a bus controller) of the computing circuit layer 103 through a second bonding interface to form a vertical data channel.
[0168] For example, the control circuit layer 102 and the computing circuit layer 103 can respectively adopt different process nodes (e.g., the control circuit layer uses 28nm CMOS and the computing circuit layer uses 16nm FinFET, and the embodiments of the present disclosure do not limit the specific process), and cross-process interconnection is achieved through TSVs, taking into account both performance and cost.
[0169] For example, a thermal interface material (TIM) can be integrated at the second bonding interface to quickly conduct the heat of the computing circuit layer 103 through the TSVs to the substrate.
[0170] For example, a eutectic bond can also be used to provide a high-strength mechanical connection to prevent delamination between stacked layers.
[0171] For example, the second substrate of the control circuit layer 102 can also be polished and wet-etched. Exemplarily, the thickness of the second substrate is thinned from the original 750 μm to less than 50 μm to expose the bottom metal pads of the TSVs.
[0172] For example, the active surface of the computing circuit layer 103 can be aligned with the thinned back surface (including TSV pads) of the control circuit layer 102, and an electrical connection is formed through eutectic bonding (such as Sn-Ag alloy) or thermocompression bonding.
[0173] For example, after bonding, a polymer underfill can be filled to enhance the mechanical strength of the interface and prevent moisture erosion.
[0174] In a possible implementation, for example, when the edge device executes an image classification task, such as during the weight loading phase, the computing circuit layer 103 can send a read instruction to the control circuit layer 102 through the TSV of the second bonding interface; the control circuit layer 102 can read the model weights from the storage array layer 101 through the hybrid bonding interconnection of the first bonding interface and temporarily store them in the page buffer; the weight data can be transmitted to the L2 cache of the computing circuit layer 103 through the TSV. For example, during the inference calculation phase, the computing core can read the weights from the L2 cache, process the input image, and generate intermediate features; the intermediate features can be written back to the specified block of the storage array layer 101 through the TSV to release the cache space. For example, during the result output phase, the classification result can be output to the external display module through the interface unit of the computing circuit layer 103 and the conductive pads of the third substrate.
[0175] In at least one embodiment of the present disclosure, through the double bonding interface design of active alignment and non-active alignment, the three-dimensional efficient collaboration of the storage, control, and computing layers is achieved. Combining the hybrid bonding and TSV technologies, the bandwidth and latency limitations of the traditional architecture are broken through at the micron scale, providing a high-integration, low-power, and highly scalable memory-computation integrated chip solution for edge artificial intelligence.
[0176] In some embodiments of the present disclosure, the storage array layer can also be arranged to face-to-back bond with the control circuit layer, and the control circuit layer can also be arranged to face-to-back bond with the computing circuit layer. That is, in this implementation, the chips (dies) of each circuit layer are all oriented in the same direction. For example, the storage array layer is arranged on the first substrate, the second substrate is arranged on the side of the storage array layer away from the first substrate, the control circuit layer is arranged on the side of the second substrate away from the first substrate, the third substrate is arranged on the side of the control circuit layer away from the first substrate, and the computing circuit layer is arranged on the side of the third substrate away from the first substrate.
[0177] Figure 5 The schematic diagram of a method for manufacturing an integrated circuit chip provided by at least one embodiment of the present disclosure is shown.
[0178] In some embodiments of the present disclosure, a method for manufacturing an integrated circuit chip is also provided. As Figure 5 shown, the method for manufacturing the integrated circuit chip includes steps S200 to S240.
[0179] Step S200: Prepare the storage array layer on the first substrate.
[0180] Step S220: Prepare the control circuit layer on the side of the storage array layer away from the first substrate.
[0181] Step S240: Prepare the computing circuit layer on the side of the control circuit layer away from the first substrate.
[0182] For example, the storage array layer includes at least one storage array, and each of the at least one storage array includes a plurality of NAND storage cells arranged in multiple rows and columns. The control circuit layer is configured to perform data read and write operations on the at least one storage array. The computing circuit layer is configured to send data read and write instructions to the control circuit layer to receive data in the at least one storage array and perform computing tasks.
[0183] For step S200, for example, a 3D NAND storage cell array or a 2D NAND storage cell array can be fabricated on a silicon-based first substrate.
[0184] Taking 3D NAND as an example, for example, a vertical channel structure can be formed by alternately depositing oxides and polysilicon / metal layers, and each channel can include a plurality of storage cells. For example, for the bit line (BL), the metal lines in the horizontal direction can connect the storage cells in the same row, and for example, a copper or aluminum dual-damascene process can be used. For example, for the word line (WL), the metal lines in each layer in the vertical stack can connect the control gates at the same layer position, and the contact points are exposed through staircase etching. For example, for the select line, the top and bottom select transistors can be connected to the DSG line (top side) and the SSG line (bottom side) respectively, and the transistor channels are formed by ion implantation.
[0185] For example, substrate pretreatment can be performed. For example, the silicon substrate can be cleaned and a buffer layer (such as SiO2) can be deposited. For example, stack deposition can be performed. For example, oxides and sacrificial layers (such as SiN) can be alternately grown by chemical vapor deposition (CVD). For example, channel etching can be performed. For example, deep reactive ion etching (DRIE) can be used to form vertical holes, and polysilicon can be filled as the channel. For example, gate replacement can be performed. For example, the sacrificial layer can be removed by wet etching, a high-k dielectric and a metal gate (such as TiN / W) can be deposited to form the control gate of the storage cell. For example, interconnect metallization can be performed. For example, a bit line, a word line and a select line network can be formed by lithography and electroplating processes.
[0186] For step S220, for example, a 28nm or other CMOS process can be used to balance performance and cost, and the embodiments of the present disclosure are not limited thereto. For example, high-voltage devices (such as word line driving transistors) can use an LDMOS (lateral diffused MOS) design, and for example, the breakdown voltage can be set to ≥30V.
[0187] For example, for the storage array layer and the control circuit layer to achieve active alignment through a bonding technique to form a first bonding interface, the surface of the storage array layer can be processed.
[0188] For example, the surface of the storage layer can be polished by CMP to make the copper pads (diameter ≤1μm) coplanar with the oxide layer.
[0189] For example, through hybrid bonding, the copper pads (with a diameter of 1 μm) can be coplanarly aligned with the oxide layer (error < 50 nm).
[0190] For example, a low-impedance interconnection can be formed through low-temperature pre-bonding (200 °C / 10 kN) followed by high-temperature curing (400 °C).
[0191] For example, electrical verification can also be performed to test the bit-line resistance (< 10 Ω) and the word-line delay (< 5 ns).
[0192] In some embodiments of the present disclosure, after step S220 in the method for manufacturing the integrated circuit chip, step S230 is further included.
[0193] Step S230: Provide a second substrate.
[0194] For example, the control circuit layer is disposed on the second substrate, and on the side of the second substrate facing the first substrate, the computing circuit layer is disposed on the side of the second substrate facing away from the first substrate. The second substrate includes a first through-silicon via, and the control circuit layer and the computing circuit layer are interconnected through the first through-silicon via. The computing circuit layer can be further configured to send data read / write instructions to the control circuit layer through the first through-silicon via.
[0195] For step S230, for example, the second substrate can serve as a carrier for the control circuit layer, integrating the first through-silicon via (TSV) to achieve vertical interconnection between the control circuit layer and the computing circuit layer.
[0196] For example, for the preparation of the TSV, operations such as via etching, metal filling, and RDL wiring can be performed.
[0197] Exemplarily, for via etching, for example, a via with a diameter of 5 μm and a depth of 50 μm can be etched through DRIE. For metal filling, for example, the via can be filled with copper by electroplating and polished to a flat surface by CMP. For RDL wiring, for example, copper can be deposited first and then the redistribution layer (RDL) to expand the TSV pads to 10 μm × 10 μm.
[0198] For step S240, the computing circuit layer is fabricated on the side of the control circuit layer away from the first substrate. For example, the FinFET process can be used to integrate an arithmetic logic unit (ALU), a tensor processing unit (TPU), a register file, a cache unit, and an interface unit in the computing circuit layer.
[0199] For example, the computing circuit layer can adopt an advanced process (such as 7 nm or 5 nm process), and is heterogeneously integrated with the process of the control circuit layer (such as 28 nm process) through TSV.
[0200] For example, for the control circuit layer and the computing circuit layer, non-active alignment of the control circuit layer is achieved through a bonding technique to form a second bonding interface. For example, the control circuit layer can be thinned. Exemplarily, the thickness of the second substrate can be thinned from 750 μm to 50 μm by mechanical grinding. For example, for the formation of TSVs, through-holes can be etched by DRIE, filled with copper, and chemically mechanically polished (CMP) until the surface is flat.
[0201] For example, for the bonding process of the control circuit layer and the computing circuit layer, for example, for the preparation of the computing circuit layer, on the side of the second substrate facing the first substrate, the manufacturing of the computing circuit (such as a GPU architecture) is achieved, and pads (such as copper pads, size 10 μm × 10 μm) are deposited on the surface. For example, the active surface (including pads) of the computing circuit layer can be aligned with the thinned back surface (including TSVs) of the control circuit layer, and a bonding technique (such as eutectic bonding, Sn-Ag eutectic alloy (melting point 221 °C) bonding) is used to form a low-impedance interconnection. For example, reinforcement of the second bonding interface can also be performed, for example, by injecting an underfill polymer, and after curing, the mechanical stability is enhanced.
[0202] In some embodiments of the present disclosure, the method for manufacturing the integrated circuit chip further includes step S250: providing a third substrate.
[0203] For example, the computing circuit layer is disposed on the side of the third substrate facing the second substrate; the third substrate includes second through-silicon vias and conductive pads electrically connected to the second through-silicon vias are disposed on the side facing away from the computing circuit layer.
[0204] In step S250, the third substrate can serve as a carrier for the computing circuit layer, integrating the second through-silicon vias (TSVs) and the conductive pads to achieve high-speed interaction between the chip and the external system.
[0205] For example, the conductive pads can include copper-gold plated pads (size 50 μm × 50 μm, pitch 100 μm).
[0206] For example, integration of TSVs and conductive pads can be prepared. For example, through-holes can be etched and filled with copper, with a diameter of 2 μm and a depth of 50 μm to prepare TSVs. For example, RDL can be deposited to connect the TSVs to the conductive pads. For example, a high thermal conductivity material (such as graphene) can also be filled in the TSVs to conduct heat from the computing layer.
[0207] For example, for the specific operation methods, functions, or beneficial effects related to steps S200 to S250, for example, reference can also be made to the relevant descriptions of the embodiments of the integrated circuit chip provided in any embodiment of the present disclosure, which will not be elaborated here.
[0208] Figure 6The figure shows an application schematic diagram of a method for manufacturing an integrated circuit chip provided by at least one embodiment of the present disclosure.
[0209] Taking the method for manufacturing a three-dimensional storage computing chip of NAND flash as an example below, the method for manufacturing an integrated circuit chip provided by the present disclosure will be described exemplarily. This manufacturing method can be used, for example, to manufacture Figure 7 the integrated circuit chip shown.
[0210] As Figure 7 shown, the integrated circuit chip includes a storage array layer, a control circuit layer, and a computing circuit layer. Among them, the storage array layer further includes a semiconductor layer, a storage cell array, and an interconnect layer; the control circuit layer includes an interconnect layer, a device layer, and a semiconductor layer; the computing circuit layer includes an interconnect layer, a device layer, and a semiconductor layer. For example, the storage array layer and the control circuit layer are bonded through a bonding layer, and the control circuit layer and the computing circuit layer are bonded through a bonding layer. For example, each bonding layer includes a plurality of bonding contacts. For example, the control circuit layer and the computing circuit layer each further include a plurality of transistors to implement the corresponding chip circuit design.
[0211] For example, the manufacturing method includes the following steps (1) to (8).
[0212] (1) Preparation of the storage array layer. Build a 3D / 2D NAND storage cell array on the first substrate to achieve high-density data storage.
[0213] (1.1) Substrate pretreatment: Clean the silicon substrate and deposit a buffer layer (such as SiO2).
[0214] (1.2) 3D NAND stacking: Alternately deposit oxides and sacrificial layers (such as SiN) to form a multi-layer stacked structure (such as 64 layers); perform deep reactive ion etching (DRIE) to form vertical channel holes, and fill polysilicon as the storage cell channels; remove the sacrificial layer by wet etching, deposit a high-k dielectric (such as HfO2) and a metal gate (such as TiN / W) to complete gate replacement.
[0215] (1.3) Interconnect wiring: Use photolithography and electroplating processes to form bit lines (BL), word lines (WL), DSG lines (top side select lines), and SSG lines (bottom side select lines).
[0216] (1.4) Multi-array partitioning: Divide multiple storage planes / blocks through deep trench isolation (DTI) to support parallel access.
[0217] (2) Preparation of the control circuit layer. Integrate peripheral circuits on the second substrate to achieve read / write control of the storage array.
[0218] (2.1) Peripheral circuit design: Column decoder / bit line driver: Drive the bit line voltage (0.5V read / 20V program). Row decoder / word line driver: Select the word line and apply a stepped voltage. Sense amplifier (SA): Detect the bit line current / voltage difference and convert it into a digital signal. Logic control unit: Parse instructions and generate timing signals.
[0219] (2.2) Adopt the CMOS process and integrate high-voltage LDMOS devices (with a breakdown voltage ≥ 30V).
[0220] (3) Preparation of the computing circuit layer. Build a computing core on the third substrate to support edge AI tasks.
[0221] (3.1) Computing module design. Computing core: FinFET process, integrate ALU, tensor processing unit (TPU) and register file. L2 cache: 16MB SRAM, support ECC check. Interface unit: Support high-speed protocols such as PCIe 5.0 and LPDDR5.
[0222] (3.2) Heterogeneous integration adaptation: Design a TSV interface to achieve high-bandwidth interconnection with the lower-layer control circuit.
[0223] (4) Bonding of the storage layer and the control layer. Achieve high-density interconnection between the storage array layer and the control circuit layer through bonding technology.
[0224] (4.1) Surface treatment: Chemically mechanical polish (CMP) the bonding surface of the storage layer and the control layer, with a roughness < 0.5nm.
[0225] (4.2) Hybrid bonding: Copper-copper bonding points (diameter ≤ 1μm) and oxide bonding regions are coplanarly aligned (error < 50nm); After low-temperature pre-bonding (200°C / 10kN), high-temperature curing (400°C) is carried out to form a low-impedance interconnection.
[0226] (4.3) Electrical verification: Test the bit line resistance (< 10Ω), word line delay (< 5ns) and sense amplifier sensitivity.
[0227] (5) Thinning of the control layer and preparation of TSV. Thin the second substrate and form vertical interconnection channels (TSV).
[0228] (5.1) Substrate thinning: Thin the second substrate from 750μm to 50μm through mechanical grinding and wet etching.
[0229] (5.2) TSV etching and filling: DRIE etch through-holes with a diameter of 5μm and a depth of 50μm, electroplate copper for filling and CMP polishing.
[0230] (5.3) Redistribution Layer (RDL): Deposit copper RDL to expand the TSV pads to 10μm × 10μm.
[0231] (6) Bonding of the control layer and the computing layer. Achieve vertical interconnection between the control circuit layer and the computing circuit layer through TSV.
[0232] (6.1) Non-active alignment bonding: Align the computing circuit layer (the active surface of the third substrate) with the thinned back surface of the second substrate (including TSV). Use Sn-Ag eutectic bonding (melting point 221°C) to form a low-impedance interconnection.
[0233] (6.2) Interface strengthening: Inject polymer underfill to enhance mechanical stability and moisture resistance.
[0234] (6.3) Electrical testing: Verify the TSV conduction rate (>99.9%) and data transmission bandwidth (>1 Tbps).
[0235] (7) Preparation of the third substrate and integration of conductive pads. Construct an external interface to support the interaction between the chip and the system.
[0236] (7.1) Preparation of the second Through-Silicon Via (TSV):
[0237] Etch vias with a diameter of 2μm and a depth of 50μm, fill them with copper and polish by CMP.
[0238] (7.2) Conductive pad design: Copper-gold plated pads (50μm × 50μm, pitch 100μm), and the layout is divided into: data I / O area, power / ground area.
[0239] (7.3) Thermal management integration: Fill the TSV with a high thermal conductivity material (such as graphene) to conduct the heat of the computing layer to the package radiator.
[0240] (8) Thinning and interconnection layer packaging. Optimize the chip thickness and complete the external interconnection.
[0241] (8.1) Substrate thinning: Grind the first or third substrate to the target thickness (such as 100μm) to ensure package compatibility.
[0242] (8.2) Pad lead-out interconnection layer: Deposit aluminum / copper pads and connect them to the package substrate through wire bonding or flip-chip.
[0243] (8.3) Reliability testing: Perform thermal cycle testing (-40°C to 125°C) and high-accelerated stress test (HAST) to ensure chip stability.
[0244] Figure 8The schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure is shown.
[0245] As Figure 8 shown, the electronic device 300 includes an integrated circuit chip 400.
[0246] For example, the integrated circuit chip 400 can be the integrated circuit chip provided by any of the above embodiments. For example, the electronic device 300 can further include other devices, such as a central processing unit (CPU), a data bus, a memory, etc. The electronic device 300 can be a signal processing device, a computing device, etc. For example, it can be used for a controller, a terminal device, or a server device, etc.
[0247] In addition to the above exemplary description, the following points of the present disclosure need to be noted:
[0248] (1) The drawings of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.
[0249] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0250] (3) It should be understood that in the embodiments of the present disclosure, the magnitudes of the sequence numbers of the above steps do not mean the order of execution. The order of execution of each step should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.
[0251] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. An integrated circuit chip, comprising: A memory array layer, including at least one memory array, wherein each of the at least one memory arrays includes a plurality of NAND memory cells arranged in multiple rows and columns; A control circuit layer configured to perform data read and write operations on the at least one memory array; and A computing circuit layer configured to send data read and write instructions to the control circuit layer to receive data in the at least one memory array and perform computing tasks; Wherein, the memory array, the control circuit layer, and the computing circuit layer are stacked relative to a first substrate, the memory array layer is disposed on the first substrate, the control circuit layer is disposed on a side of the memory array layer away from the first substrate, and the computing circuit layer is disposed on a side of the control circuit layer away from the first substrate.
2. The integrated circuit chip according to claim 1, wherein, The control circuit layer includes a page buffer device to cache data of a target memory page in the at least one memory; array; the page buffer device includes a sensing circuit.
3. The integrated circuit chip according to claim 2, wherein, The control circuit layer further includes a bit line selector and a scan selector, The control circuit layer is further configured to dynamically activate the target memory page through the bit line selector and the scan selector.
4. The integrated circuit chip according to any one of claims 1-3 further comprises: A second substrate, Wherein, the control circuit layer is disposed on the second substrate and on a side of the second substrate facing the first substrate, the computing circuit layer is disposed on a side of the second substrate facing away from the first substrate, the second substrate includes a first through-silicon via, and the control circuit layer and the computing circuit layer are interconnected through the first through-silicon via, The computing circuit layer is further configured to send data read and write instructions to the control circuit layer through the first through-silicon via.
5. The integrated circuit chip according to claim 4 further comprises: A third substrate, Wherein, the computing circuit layer is disposed on a side of the third substrate facing the second substrate; The third substrate includes a second through-silicon via and a conductive pad electrically connected to the second through-silicon via is disposed on a side facing away from the computing circuit layer.
6. The integrated circuit chip according to claim 2, wherein, The computing circuit layer further includes a cache unit, The cache unit is configured to form a data path with the sensing circuit through the first through-silicon via.
7. The integrated circuit chip according to any one of claims 1 to 3, wherein, The memory array layer and the control circuit layer achieve active alignment through a bonding technique to form a first bonding interface.
8. The integrated circuit chip according to any one of claims 1 to 3, wherein, The control circuit layer and the computing circuit layer achieve non-active alignment of the control circuit layer through a bonding technique to form a second bonding interface.
9. A method for manufacturing an integrated circuit chip, comprising: Preparing a memory array layer on a first substrate; Preparing a control circuit layer on a side of the memory array layer away from the first substrate; And Preparing a computing circuit layer on a side of the control circuit layer away from the first substrate; Wherein, the memory array layer includes at least one memory array, Each of the at least one memory arrays includes a plurality of NAND memory cells arranged in multiple rows and columns, The control circuit layer is configured to perform data read and write operations on the at least one memory array, and The computing circuit layer is configured to send data read and write instructions to the control circuit layer to receive data in the at least one memory array and perform computing tasks.
10. The method for manufacturing an integrated circuit chip according to claim 9 further includes: Providing a second substrate, Wherein, the control circuit layer is disposed on the second substrate, and on a side of the second substrate facing the first substrate, the computing circuit layer is disposed on a side of the second substrate facing away from the first substrate. The second substrate includes a first through-silicon via, and the control circuit layer and the computing circuit layer are interconnected through the first through-silicon via. The computing circuit layer is further configured to send data read / write instructions to the control circuit layer through the first through-silicon via.
11. The method for manufacturing an integrated circuit chip as described in claim 10 further includes: Provide a third substrate. Wherein, the computing circuit layer is disposed on a side of the third substrate facing the second substrate. The third substrate includes a second through-silicon via and a conductive pad electrically connected to the second through-silicon via is disposed on a side facing away from the computing circuit layer.
12. An electronic device, comprising the integrated circuit chip according to any one of claims 1-8.
Citation Information
Cited By
Storage and calculation integrated chip architecture, chip stacking and packaging structure and terminal equipment
CN121009055A
Storage and calculation integrated chip, data access method, electronic equipment and storage medium
CN121301275A