Homogeneous / heterogeneous integrated systems with high performance computing and high storage capacity

The integrated system with self-aligned transistors and innovative interconnects addresses the miniaturization challenges of SRAM and logic circuits, achieving up to 3 times area reduction and improved computing performance in AI chips by efficiently packing more functional blocks and cache levels within a single die.

JP2025102809APending Publication Date: 2025-07-08INVENTION & COLLABORATION LAB PTE LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025041763
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-16
Filing Date
2025-03-14
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The miniaturization of SRAM and logic circuits in semiconductor manufacturing faces challenges due to increased die size, interference between metal wirings, and parasitic junctions, leading to inefficiencies and higher costs in high-performance computing chips like AI chips, where conventional miniaturization technologies like Moore's Law are reaching their limits.

Method used

An integrated system with monolithic dies containing processing devices and SRAM arrays, utilizing self-aligned transistor structures and innovative interconnects to reduce die size, allowing for more efficient packing of circuits and caches, and eliminating the need for high-bandwidth memory.

Benefits of technology

The integrated system significantly reduces the area of SRAM and logic circuits by up to 3 times, enabling higher device density and computing performance without shrinking technology nodes, and allows for more functional blocks and cache levels within a single monolithic die, enhancing AI chip capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025102809000001_ABST
    Figure 2025102809000001_ABST
Patent Text Reader

Abstract

To provide more powerful and efficient system-on-chip (SOC) or artificial intelligence (AI) chips and integrated systems based on monolithic integration.SOLUTION: An integrated system includes a first monolithic die and a second monolithic die. The first monolithic die includes a processing device circuit formed therein and the second monolithic die having a plurality of SRAM arrays formed therein. The second monolithic die has at least 2 GBytes, and the first monolithic die is electrically connected to the second monolithic die.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to semiconductor structures, and more particularly to integrated systems including logic chips with high-performance computing and SRAM chips with large memory capacity.

Background Art

[0002] Information technology (IT) systems are rapidly evolving in every enterprise and business, including those in factories, healthcare, and transportation. Today, system-on-chip (SOC) or artificial intelligence (AI) has become the core of IT systems that make factories smarter, improve patient prognosis, and enhance the safety of autonomous vehicles. Data from manufacturing equipment, sensors, and machine vision systems can easily total 1 petabyte per day. Therefore, to handle such petabyte-scale data, high-performance computing (HPC) SPC or AI chips are required.

[0003] Generally, AI chips can be classified by graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs). GPUs were originally designed to process graphical processing applications using parallel computing, but have increasingly been used for AI learning. The learning speed and efficiency of GPUs are generally 10 to 1000 times greater than those of general-purpose CPUs.

[0004] FPGAs have logic blocks that interact with each other and may be designed by engineers to support specific algorithms and are suitable for AI inference. Although FPGAs have drawbacks such as larger size, slower speed, and higher power consumption, they are preferred over ASIC design due to short time to market, low cost, and flexibility. Due to the flexibility of FPGAs, any part of the FPGA can be partially programmed according to requirements. The inference speed and efficiency of FPGAs are 10 to 100 times greater than those of general-purpose CPUs.

[0005] On the one hand, ASICs are directly adapted to the circuit and are generally more efficient than FPGAs. In the case of customized ASICs, their learning / inference speed and efficiency can be 10 to 1000 times greater than that of general-purpose CPUs. However, as AI algorithms continue to evolve, unlike FPGAs where customization is becoming easier, ASICs are gradually becoming obsolete as new AI algorithms are developed.

[0006] In any of GPUs, FPGAs, and ASICs (or similar SOCs, CPUs, NPUs, etc.), logic circuits and SRAM circuits are the two main circuits, and their combination generally accounts for approximately 90% of the AI chip size. The remaining 10% of the AI chip may include I / O pad circuits. However, the miniaturization process / technology node for manufacturing AI chips is becoming increasingly necessary to efficiently and rapidly train AI machines as they bring better efficiency and performance. The improvement of integrated circuit performance and cost has mainly been achieved by process miniaturization technology according to Moore's law, but such miniaturization technology up to 3nm to 5nm faces many technical problems, and thus, the investment cost and capital in the research and development of the semiconductor industry have increased exponentially.

[0007] For example, the miniaturization of SRAM devices for increased memory density, the reduction in operating voltage (VDD) for lower standby power consumption, and the improved yield required to achieve larger-capacity SRAM are becoming more difficult to achieve as the miniaturization to a 28nm (or lower) manufacturing process poses challenges.

[0008] FIG. 1 shows a SRAM cell architecture, i.e., a six-transistor (6-T) SRAM cell. It consists of two cross-coupled inverters (PMOS pull-up transistors PU-1 and PU-2, and NMOS pull-down transistors PD-1 and PD-2), and two access transistors (NMOS passgate transistors PG-1 and PG-2). A high-level voltage VDD is coupled to the PMOS pull-up transistors PU-1 and PU-2, and a low-level voltage VSS is coupled to the NMOS pull-down transistors PD-1 and PD-2. When the word line (WL) is enabled (i.e., the row is selected within the array), the access transistors are turned on to connect the storage nodes (Node-1 / Node-2) to the bit lines (BL and BL Bar) extending in the vertical direction.

[0009] FIG. 2 shows a "stick diagram" representing the layout and connection of the six transistors of the SRAM. The stick diagram typically includes only the active regions (vertical bars) and the gate lines (horizontal bars) to form the pull-down transistor PD and the pull-up transistor PU among the six transistors of the SRAM. Of course, on the one hand, there are still many contacts directly coupled to the six transistors, and on the other hand, coupled to the word line (WL), the bit lines (BL and BL Bar), the high-level voltage VDD, and the low-level voltage VSS, etc.

[0010] λ when the minimum feature size decreases 2 or F 2Part of the reason for the dramatic increase in the total area of the SRAM cell represented by can be explained as follows. Conventional 6T SRAM has six transistors connected by using multiple interconnects, and has a first interconnect layer M1 for connecting the gate level ("gate") of the transistors and the diffusion levels of the source and drain regions (regions generally called "diffusion"). In order to facilitate signal transmission without increasing the die size by only using M1, it is necessary to increase the second interconnect layer M2 and / or the third interconnect layer M3 (for example, word line (WL) and / or bit line (BL and BL Bar)), and therefore, a structure Via-1 composed of a specific type of conductive material is formed to connect M2 to M1.

[0011] Therefore, there exists a vertical structure formed from diffusion through the contact (Con) connection to M1, that is, "diffusion-Con-M1". Similarly, another structure for connecting the gate to M1 through the contact structure can be formed as "gate-Con-M1". Furthermore, when a connection structure for connecting from the M1 interconnect through via1 to the M2 interconnect needs to be formed, it is called "M1-via1-M2". A more complex structure from the gate level to the M2 interconnect can be represented as "gate-Con-M1-via1-M2". Furthermore, the stacked interconnect system can have a structure such as "M1-via1-M2-via2-M3" or "M1-via1-M2-via2-M3-via3-M4".

[0012] The gates and diffusions in the two access transistors (NMOS pass gate transistors PG-1 and PG-2 as shown in FIG. 1) are connected to the word line (WL) and / or the bit line (BL and BL Bar), which are arranged in the second interconnect layer M2 or the third interconnect layer M3. Therefore, in conventional SRAM, such metal connections must first pass through the interconnect layer M1. That is, the latest interconnect system in SRAM may not allow the gate or diffusion to be directly connected to M2 without bypassing the M1 structure.

[0013] As a result, the required space between one M1 interconnect and the other M1 interconnect increases the die size, and in some cases, the wiring connection may inhibit the intention of a specific efficient channeling to directly use M2 to cross the M1 region. Furthermore, it is difficult to form a self-alignment structure between via 1 and the contact, and at the same time, it is difficult for both via 1 and the contact to be connected to their respective interconnect systems.

[0014] Furthermore, in a conventional 6T SRAM, at least one NMOS transistor and one PMOS transistor are respectively arranged inside some adjacent regions of the p-substrate and the n-well, which are formed adjacent to each other within a close range. A parasitic junction structure called an n+ / p / n / p+ parasitic bipolar device is formed along its contour from the n+ region of the NMOS transistor to the p-well, to the adjacent n-well, and further to the p+ region of the PMOS transistor.

[0015] There is a large noise generated at the n+ / p junction or the p+ / n junction, and a very large current abnormally flows through this n+ / p / n / p+ junction, which may, in some cases, block the operation of some parts of the CMOS circuit and cause malfunction of the entire chip. Such an abnormal phenomenon called latch-up has an adverse effect on CMOS operation and must be avoided. Certainly, one way to enhance the tolerance to latch-up, which is indeed a weakness of CMOS, is to increase the distance from the n+ region to the p+ region. As a result, increasing the distance from the n+ region to the p+ region to avoid the latch-up problem also enlarges the size of the SRAM cell.

[0016] However, even with the miniaturization of the manufacturing process to 28 nm or less (so-called "minimum feature size", "lambda (λ)", or "F"), λ 2 or F 2The total area of the SRAM cell represented by is increased dramatically as the minimum feature size decreases, as shown in Fig. 3, due to interference between the sizes of the contacts and interference between the layouts of the metal wirings connecting the word line (WL), bit lines (BL and BL Bar), high-level voltage VDD, and low-level voltage VSS, etc. (quoted from "15.1 A 5nm 135Mb SRAM in EUV and High-Mobility-Channel FinFET Technology with Metal Coupling and Charge-Sharing Write-Assist Circuitry Schemes for High-Density and Low-VMIN Applications, 2020 IEEE International Solid- State Circuits Conference - (ISSCC), 2020, pp. 238-240)" by J. Chang et al.).

[0017] A similar situation also occurs in the miniaturization of logic circuits. The miniaturization of logic circuits for increased memory density, the reduction of the operating voltage (Vdd) for low standby power consumption, and the increased yield required to realize a larger-capacity logic circuit are becoming more difficult to achieve. Standard cells are generally used and are the basic elements within a logic circuit. Standard cells may include basic logic function cells (e.g., inverter cells, NOR cells, NAND cells).

[0018] Similarly, in the miniaturization of the manufacturing process to 28 nm or less, λ 2 or F 2 The total area of the standard cell represented by is increased dramatically as the minimum feature size decreases, due to the size of the contacts and interference between the layouts of the metal wirings.

[0019] Figure 4(a) shows a "stick diagram" representing the layout and connections of PMOS and NMOS transistors in a 5nm (UHD) standard cell of a semiconductor company. The stick diagram includes only the active regions (horizontal lines) and gate lines (vertical lines). Hereinafter, the active region may be referred to as a "fin". Of course, on the one hand, it is directly connected to PMOS and NMOS transistors, and on the other hand, there are still many contacts connected to input terminals, output terminals, high-level voltage Vdd, and low-level voltage VSS (or ground "GND"), etc. In particular, each transistor includes two active regions or fins (marked by the gray dashed rectangle) to form the channel of the transistor, so that the W / L ratio can be maintained within an acceptable range. The area size of the inverter cell is equal to X×Y, where X = 2×Cpp, Y = cell height, and Cpp is the distance of the contact to the poly pitch (Cpp).

[0020] Some of the active regions or fins between PMOS and NMOS (referred to as "dummy fins") are not utilized within the PMOS / NMOS of this standard cell, and it can be seen that the potential reason is likely related to the latch-up issue between PMOS and NMOS. Therefore, the latch-up distance between PMOS and NMOS in Figure 4(a) is 3×Fp, where Fp is the fin pitch. Based on the available data regarding Cpp (54nm) and cell height (216nm) in the 5nm standard cell, the cell area is 23328nm 2 (or 933.12λ 2 where lambda (λ) is the minimum feature size of 5nm) can be calculated by X×Y which is equal to it. Figure 4(b) shows the aforementioned 5nm standard cell and its dimensions. As shown in Figure 4(b), the latch-up distance between PMOS and NMOS is 15λ, Cpp is 10.8λ, and the cell height is 43.2λ.

[0021] The miniaturization trend of the area size (2Cpp × cell height) for three fundries with respect to different process technology nodes can be shown in Fig. 5. (For example, from 22nm to 5mm) As the technology node decreases, λ 2 In terms of conversion, it is obvious that the area size of the conventional standard cell (2Cpp × cell height) increases dramatically. Within the conventional standard cell, the smaller the technology process node, the higher the area size in terms of λ 2 conversion. Such a dramatic increase in λ 2 can be brought about by the difficulty of proportionally reducing the size of gate contacts / source contacts / drain contacts as λ decreases, whether in SRAM or logic circuits, the difficulty of proportionally reducing the latch-up distance between PMOS and NMOS, and interference within the metal layer in response to the decrease in λ.

[0022] From another perspective, any high-performance computing (HPC) chips such as SOC, AI, NPU (network processing unit), GPU, CPU, and FPGA, etc., currently, they are using monolithic integration to place as many and more circuits as possible. However, as shown in Fig. 6(a), maximizing the die area of each monolithic die is limited by the maximum reticle size of the lithography stepper, which is difficult to expand due to existing state-of-the-art photolithography exposure tools. For example, as shown in Fig. 6(b), current i193 and EUV lithography steppers have a maximum reticle size, thus, the monolithic SOC die has a scanner maximum field area (SMFA) of 26mm × 33mm, i.e., 858mm 2 . However, for high-performance computing or AI purposes, high-end consumer GPUs are 500 - 600mm 2It seems to reach. As a result, it has become more difficult or impossible to implement two or more major functional blocks (for example, GPU and FPGA) on a single die. Furthermore, since the most widely used six-transistor CMOS SRAM cell is extremely large, the size of sufficient embedded SRAM (eSRAM) for both major blocks is increased. In addition, although the external DRAM capacity needs to be expanded, discrete PoP (package-on-package, for example, HBM~SOC) or POD (package DRAM on SOC die) is still restricted by the difficulty of achieving the desired performance of worse inter-die or inter-package-chip signal interconnects.

[0023] Therefore, in the near future, there is a need to propose a new integrated system including a logic chip with HPC and an SRAM chip with high memory capacity that can solve the above problems so that a more powerful and efficient SOC or single chip of AI based on monolithic integration can be realized.

Summary of the Invention

[0024] One aspect of the present disclosure provides an integrated system, the integrated system comprising a first monolithic die and a second monolithic die. The first monolithic die has a processing device circuit formed therein, and the second monolithic die has a plurality of SRAM arrays formed therein. The second monolithic die comprises at least 2 gigabytes, and the first monolithic die is electrically connected to the second monolithic die.

[0025] In one embodiment of the present disclosure, the first monolithic die has a die area that is the same as or substantially the same as the maximum scanner image area defined by a specific technology process node, and the second monolithic die has a die area that is the same as or substantially the same as the maximum scanner image area defined by the specific technology process node.

[0026] In one embodiment of the present disclosure, the maximum drawing screen area of the scanner is 26 mm × 33 mm, or not greater than 858 mm 2 or less.

[0027] In one embodiment of the present disclosure, the first monolithic die and the second monolithic die are housed within a single package.

[0028] In one embodiment of the present disclosure, the plurality of SRAM arrays include at least 20 gigabytes.

[0029] In one embodiment of the present disclosure, the processing device circuit includes a first processing device circuit and a second processing device circuit. The first processing device circuit includes a plurality of first logic cores, each of the plurality of first logic cores includes a first SRAM size, the second processing device circuit includes a plurality of second logic cores, and each of the plurality of second logic cores includes a second SRAM size.

[0030] In one embodiment of the present disclosure, the main function performed by the first processing device circuit is different from the main function performed by the second processing device circuit.

[0031] In one embodiment of the present disclosure, the first processing device circuit or the second processing device circuit is selected from the group consisting of a graphics processing unit (GPU), a central processing unit (CPU), a tensor processing unit (TPU), a network processing unit (NPU), and a field programmable gate array (FPGA).

[0032] In one embodiment of the present disclosure, the first monolithic die further includes an L1 cache and an L2 cache utilized by the processing device circuit during operation of the first monolithic die, and the plurality of SRAM arrays include an L3 cache and an L4 cache utilized by the processing device circuit during operation of the first monolithic die.

[0033] In one embodiment of the present disclosure, the integrated system further includes a third monolithic die in which a plurality of SRAM arrays are formed. The third monolithic die includes at least 2 to 20 gigabytes. The first monolithic die, the second monolithic die, and the third monolithic die are housed within a single package. The first monolithic die has a die area that is the same as or substantially the same as the maximum scanner image area defined by a particular technology process node. The second monolithic die has a die area that is the same as or substantially the same as the maximum scanner image area defined by the particular technology process node. The third monolithic die has a die area that is the same as or substantially the same as the maximum scanner image area defined by the particular technology process node.

[0034] In one embodiment of the present disclosure, the first monolithic die, the second monolithic die, and the third monolithic die are stacked vertically.

[0035] Another aspect of the present disclosure provides an integrated system that includes a first monolithic die and a second monolithic die. The first monolithic die has processing device circuitry. The second monolithic die has a plurality of SRAM arrays. The plurality of SRAM arrays includes at least 2 gigabytes. The first monolithic die is physically separated from the second monolithic die. The first monolithic die is electrically connected to the second monolithic die. The integrated system does not include a high bandwidth memory (HBM).

[0036] In one embodiment of the present disclosure, the second monolithic die has an area size that is the same as or substantially the same as the maximum scanner image area defined by a particular technology process node. The first monolithic die has an area size that is the same as or substantially the same as the maximum scanner image area defined by the particular technology process node.

[0037] In one embodiment of the present disclosure, the first monolithic die and the second monolithic die are housed within a single package, and the first monolithic chip is electrically connected to the second monolithic chip by wire bonding, flip chip bonding, solder bonding, interposer through-silicon via (TSV) bonding, or micro copper pillar direct bonding.

[0038] Another aspect of the present disclosure provides an integrated system, the integrated system including a first monolithic circuit and a second monolithic circuit. The first monolithic circuit has a processing device circuit, the second monolithic circuit has a plurality of SRAM arrays, the plurality of SRAM arrays including at least 2 to 20 gigabytes, the first monolithic circuit being electrically connected to the second monolithic circuit, the first monolithic circuit being formed within a first monolithic die, the second monolithic circuit being formed within a second monolithic die, the first monolithic die and the second monolithic die being housed within a single package, or the first monolithic die and the second monolithic die being housed within a first package and a second package, respectively.

Brief Description of the Drawings

[0039] The above and other aspects of the present disclosure will be better understood with respect to the following detailed description of the preferred but non-limiting (plural) embodiments. The following description is made with reference to the accompanying drawings.

[0040]

Figure 1

Figure 2

Figure 3

Figure 4(a)

Figure 4(b)

Figure 5

Figure 6(a)

Figure 6(b)

Figure 7(a)

Figure 7(b)

Figure 7(c)

Figure 8(a)

Figure 8(b)

Figure 9(a)

Figure 9(b)

Figure 10

Figure 11(a)

Figure 11(b)

Figure 11(c)

Figure 11(d)

Figure 12(a)

Figure 12(b)

Figure 13(a)

Figure 13(b)

Figure 14

Figure 15

DETAILED DESCRIPTION OF THE INVENTION

[0041] The present disclosure provides an integrated system. The above and other aspects of the present disclosure will be better understood from the following detailed description of the preferred but non-limiting embodiments. The following description is made with reference to the accompanying drawings:

[0042] Some embodiments of the present disclosure are disclosed below with reference to the accompanying drawings. However, the structures and contents disclosed in the embodiments are for illustrative and explanatory purposes only, and the scope of protection of the present disclosure is not limited to the embodiments. The present disclosure does not show all possible embodiments, and it should be noted that those skilled in the art in the technical field of the present disclosure can make appropriate modifications or changes based on the present specification disclosed below to meet actual needs without departing from the spirit of the present disclosure. The present disclosure is applicable to other implementation forms not disclosed in this specification.

[0043] Embodiment 1

[0044] The present disclosure proposes incorporating the following invention. a. A new transistor (presented in U.S. Patent Application No. 17 / 138,918, filed on December 31, 2020, entitled "MINIATURIZED TRANSISTOR STRUCTURE WITH CONTROLLED DIMENSIONS OF SOURCE / DRAIN AND CONTACT-OPENING AND RELATED MANUFACTURE METHOD", the entire content of U.S. Patent Application No. 17 / 138,918 is incorporated herein by reference; presented in U.S. Patent Application No. 16 / 991,044, filed on August 12, 2020, entitled "TRANSISTOR STRUCTURE AND RELATED INVERTER", the entire content of U.S. Patent Application No. 16 / 991,044 is incorporated herein by reference; and presented in U.S. Patent Application No. 17 / 318,097, filed on May 12, 2021, entitled "COMPLEMENTARY MOSFET STRUCTURE WITH LOCALIZED ISOLATIONS IN SILICON SUBSTRATE TO REDUCE LEAKAGES AND PREVENT LATCH-UP", the entire content of U.S. Patent Application No. 17 / 318,097 is incorporated herein by reference), b. Interconnection with the transistor (presented in U.S. Patent Application No. 17 / 528,957, filed on November 17, 2021, entitled "INTERCONNECTION STRUCTURE AND MANUFACTURE METHOD THEREOF", the entire content of U.S. Patent Application No. 17 / 528,957 is incorporated herein by reference), c. SRAM cell (presented in U.S. Patent Application No. 17 / 395,922, filed on August 6, 2021, entitled "NEW SRAM CELL STRUCTURES", the entire content of U.S. Patent Application No. 17 / 395,922 is incorporated herein by reference), and d. Standard cell design (presented in U.S. Provisional Patent Application No. 63 / 238,826, filed on August 31, 2021, entitled "STANDARD CELL STRUCTURES", the entire content of which is incorporated herein by reference)

[0045] For example, FIG. 7(a) is a top view showing a MOSFET structure according to an embodiment of the present disclosure. FIG. 7(b) is a cross-sectional view obtained along the cut line C7J1 as represented in FIG. 7(a). FIG. 7(c) is a cross-sectional view obtained along the cut line C7J2 as represented in FIG. 7(a). In the proposed MOSFET, the silicon regions of the gate terminal (e.g., silicon region 702c) and the source / drain terminals are exposed, and the selective epitaxial growth technology (SEG) has seed regions for growing pillars (e.g., the first conductor pillar portion 731a and the third conductor pillar portion 731b) based on the seed regions.

[0046] Furthermore, the first conductor pillar portion 731a and the third conductor pillar portion 731b each further have a seed region or a seed pillar within their upper portions, and such seed regions or seed pillars can be used for subsequent selective epitaxial growth. Thereafter, the second conductor pillar portion 732a is formed on the first conductor pillar 731a by a second selective epitaxial growth, and the fourth conductor pillar 732b is formed on the third conductor pillar portion 731b.

[0047] As shown in FIGS. 7(a) to 7(c), this embodiment can be applied so that an M1 interconnect (a kind of conductive terminal) or a conductive layer can be self-alignedly connected directly to an MX interconnect layer (without connecting to conductive layers M2, M3,.. MX-1) through one vertical conductive or conductor plug as long as a seed portion or a seed pillar exists on top of the conductive terminal and the conductor pillar portion configured for the subsequent selective epitaxial growth method. The seed portion or the seed pillar is not limited to silicon, and any material that can be used as a seed configured for the subsequent selective epitaxial growth is acceptable.

[0048] FIG. 8(a) is a top view showing a combined structure of a PMOS transistor 52 and an NMOS transistor 51 according to an embodiment of this embodiment. FIG. 8(b) is a cross-sectional view of the PMOS transistor 52 and the NMOS transistor 51 obtained along the cutting line (X-axis) in FIG. 8(a). The structure of the PMOS transistor 52 is the same as that of the NMOS transistor 51. A gate structure 33 including a gate dielectric layer 331 and a gate conductive layer 332 (for example, a gate metal) is formed on a horizontal or initial surface of a semiconductor substrate (such as a silicon substrate). A dielectric cap 333 (such as a composite of an oxide layer and a nitride layer) is on the gate conductive layer 332. Further, a spacer 34 that may include a composite of an oxide layer 341 and a nitride layer 342 is used to cover the sidewalls of the gate structure 33. Trenches are formed in the silicon substrate, and all or at least a part of the source region 55 and the drain region 56 are respectively disposed in the corresponding trenches. The source (or drain) region in the MOS transistor 52 may include an N+ region or other suitable doping profile regions (for example, a gradual or stepwise change from a P-region and a P+ region).

[0049] Furthermore, localized isolation 48 (such as a nitride or other high-k dielectric material) is placed within one trench and disposed under the source region, and another localized isolation 48 is placed within another trench and disposed under the drain region. Such localized isolation 48 is below the horizontal silicon surface (HSS) of the silicon substrate and can be referred to as localized isolation in silicon substrate (LISS) 48. LISS 48 can be a thick nitride layer or a composite of dielectric layers. For example, the localized isolation or LISS 48 can comprise a composite localized isolation including an oxide layer 481 covering at least a part of the sidewalls of the trench and another oxide layer 482 covering at least a part of the bottom wall of the trench. The oxide layers 481 and 482 can be L-shaped oxide layers formed by thermal oxidation treatment.

[0050] The composite localized isolation 48 can further include a nitride layer 483 on the oxide layer 482 and / or the oxide layer 481. The shallow trench isolation (STI) region may comprise a composite STI 49 including an STI-1 layer 491 and an STI-2 layer 492, and the STI-1 layer 491 and the STI-2 layer 492 may each be made of a thick oxide material by different processes.

[0051] Furthermore, the source (or drain) region can comprise a composite source region 55 and / or a drain region 56. For example, in the NMOS transistor 52, the composite source region 55 (or drain region 56) comprises at least a low doped drain (LDD) 551 and an N+ highly doped region 552 within the trench. Note that the low doped drain (LDD) 551 is in particular in contact with the exposed silicon surface having a uniform (110) crystal orientation. The exposed silicon surface has a vertical boundary with an appropriate thickness that is embedded, in contrast to the edge of the gate structure. The exposed silicon surface is substantially aligned with the gate structure. The exposed silicon surface can be the end face of the channel of the transistor.

[0052] The low-doped drain (LDD) 551 and the N+ highly doped region 552 are used as crystal seeds for growing silicon from the exposed TEC region for forming a new orderly (110) lattice over the LISS region that has no seeding effect on the changing (110) crystal structure of the newly formed crystals in the composite source region 55 or drain region 56, based on selective epitaxial growth (SEG) technology (or other suitable techniques such as atomic layer deposition ALD or selective growth ALD-SALD). Such newly formed crystals (including the low-doped drain (LDD) 551 and the N+ highly doped region 552) can be referred to as TEC-Si.

[0053] In one embodiment, the TEC is aligned or substantially aligned with the edge of the gate structure 33, the length of the LDD 551 is adjustable, and the sidewalls of the LDD 551 on the side opposite the TEC are aligned or substantially aligned with the sidewalls of the spacer 34. The composite source (or drain) region may further include several tungsten (or other suitable metal materials, such as TiN / tungsten) plugs 553 formed in the lateral connection with the TEC-Si portion for the completion of the entire source / drain region. The active channel current flowing to future metal interconnections such as the metal-1 layer reaches the tungsten 553 (or other metal materials) directly connected to the metal-1 through the LDD 551 and the N+ highly doped region 552 by several good metal-metal ohmic contacts having a much lower resistance than conventional silicon-metal contacts.

[0054] The source / drain contact resistance of the NMOS transistor 52 can be maintained within an appropriate range by the structure of the fused metal-semiconductor junction utilized within the source / drain structure. This fused metal-semiconductor junction within the source / drain structure can improve the current crowding effect and reduce the contact resistance. Furthermore, since the bottom of the source / drain structure is insulated from the substrate by the bottom oxide (oxide layer 482), the isolation between n+~n+ or p+~p+ can be maintained within an appropriate range. Therefore, the spacing between two adjacent active regions of the PMOS transistor (not shown) can be reduced to 2λ. The bottom oxide (oxide layer 482) may significantly reduce the source / drain junction leakage current, thereby reducing the leakage current between n+~n+ or p+~p+.

[0055] TIFF2025102809000002.tif66164

[0056] Furthermore, the composite STI 49 can be heightened (for example, the STI-2 layer 492 extends from the initial semiconductor surface to above the upper surface of the gate structure), so that the selectively grown source / drain regions may be confined by the composite STI and not exceed the composite STI 49. The metal contact plug (such as the tungsten plug 553) can be deposited into the hole between the composite STI 49 and the gate structure without using another contact mask for forming the contact hole. Furthermore, the top surface and one sidewall of the high-concentration doped region 552 can be directly contacted with the metal contact plug, and the contact resistance of the source / drain region can be significantly reduced.

[0057] Furthermore, in the conventional design, the metal wirings for the high-level voltage Vdd and the low-level voltage Vss (or ground) are distributed on the initial silicon surface of the silicon substrate, and such distribution interferes with a plurality of other metal wirings when there is not enough space between those metal wirings. The present invention discloses a new standard cell or SRAM cell in which the metal wirings for the high-level voltage Vdd and / or the low-level voltage Vss can be distributed under the initial silicon surface of the silicon substrate, so that even if the size of the standard cell becomes small, interference between the sizes of the contacts and interference between the layouts of the metal wirings connecting the high-level voltage Vdd and the low-level voltage Vss, etc. can be avoided.

[0058] For example, within the drain region of NMOS51, tungsten or other metal material 553 is directly coupled (by removing LISS48) to the P-well that is electrically coupled to Vdd. Similarly, within the source region of NMOS51, tungsten or other metal material 553 is directly coupled (by removing LISS48) to the p-well or P-substrate that is electrically coupled to ground. Thus, the openings of the source / drain regions originally used to electrically couple the source / drain regions to the metal-2 layer (M2) or the metal-3 layer (M3) for Vdd or ground connection can be omitted in the new standard cell and the standard cell.

[0059] In summary, there are at least the following advantages. (1) The source, drain, and gate length dimensions of the transistors within the standard cell / SRAM can be precisely controlled, and the length dimensions can be as small as the minimum feature size lambda (λ) as shown in the incorporated U.S. Patent Application No. 17 / 138,918. Thus, when two adjacent transistors are connected to each other through the drain / source, the length dimension of the transistor is as small as about 3λ, and the distance between the edges of the gates of two adjacent transistors is as small as about 2λ. Of course, due to tolerances, the length dimension of the transistor can be about 3λ to 6λ or more, and the distance between the edges of the gates of two adjacent transistors can be 8λ or more. (2) The first metal interconnect (M1 layer) is directly connected to the gate, source, and / or drain regions through self-aligned miniaturized contacts without using a conventional contact hole opening mask and / or a metal-0 conversion layer for M1 connections. (3) The gate and / or diffusion (source / drain) regions are directly connected to the metal-2 (M2) interconnect layer in a self-aligned manner without connecting to the metal-1 layer (M1). Thus, the required space between one metal-1 layer (M1) interconnect layer and the other metal-1 layer (M1) interconnect layer, and the blocking problem in some wiring connections are reduced. Further, the same structure is applicable even when the lower metal layer is directly connected to the upper metal layer by a conductor pillar but the conductor pillar is not electrically connected to any intermediate metal layer between the lower metal layer and the upper metal layer. (4) The metal wiring of the high-level voltage Vdd and / or the low-level voltage VSS within the standard cell can be distributed under the initial silicon surface of the silicon substrate. Thus, the interference between the contact sizes, the high-level voltage Vdd, and the layout of the metal wiring connecting the low-level voltage Vss, etc., can be avoided even when the size of the standard cell is reduced. Further, for Vdd or ground connection, the openings in the source / drain regions originally used to electrically couple the source / drain regions to the metal-2 layer (M2) or the metal-3 layer (M3) can be omitted in the new standard cell and the standard cell.

[0060] Based on the above, FIG. 9(a) shows the SRAM bit cell size (λ 2 conversion) that can be seen across three different companies and different technology nodes from the present invention. FIG. 9(b) is a diagram showing the comparison result of the area size of the new standard cell provided by the present invention and those of conventional products provided by various other companies. As shown in FIG. 9(a), the area of the newly proposed SRAM cell (the present invention) is about 100λ 2 and it can be approximately 1 / 8 of the area of the conventional 5nm DRAM cell (of three different companies) shown in FIG. 3. Further, as shown in FIG. 9(b), the area of the newly proposed standard cell (for example, the inverter cell can be smaller than that of 200λ 2 to the same extent) is about 1 / 3.5 of the area of the conventional 5nm standard DRAM cell shown in FIG. 5.

[0061] Therefore, the innovation in the monolithic die design of the integrated scaling and / or stretching platform (ISSP) is proposed together with any combination of the proposed technologies (for example, new transistors, interconnections with transistors, SRAM cells, and standard cell designs) so that the circuit in the original layout of the die can be reduced by 2 to 3 times or more in terms of its area, to provide an integrated system.

[0062] From another perspective, more SRAM, or more major different functional blocks (CPU or GPU) can be formed with the original size of a single monolithic die. Thus, the device density and computing performance of an integrated system (such as an AI chip or SOC) can be significantly increased compared to the conventional one of the same size without shrinking the technology node for manufacturing the integrated system.

[0063] For example, if a 5nm technology process node is used, the CMOS 6-T SRAM cell size can be reduced to approximately 100F 2 (where F is the minimum feature size fabricated on the silicon wafer). That is, when F = 5nm, based on publications, it can occupy approximately 2500nm 2 as opposed to the latest cell area of approximately 800F 2 . Further, the 8-finger CMOS inverter (shown in FIGS. 4(a) and 4(b)) consumes a die area of 200F 2 based on the present invention, as opposed to exceeding 700F 2 of the published conventional CMOS inverter (5nm process node in FIG. 9(b)).

[0064] That is, if a single monolithic die has a circuit (e.g., an SRAM circuit, a logic circuit, a combination of SRAM and logic circuits, or a major functional block circuit such as a CPU, GPU, FPGA, etc.) that occupies a die area based on a technology process node (e.g., Ynm 2 ), with the help of the present invention, the total area of the monolithic die having the circuit on the same drawing can be shrunk even if the monolithic die is still manufactured by the same technology process node. The new die area occupied by the circuit on the same drawing in the monolithic die becomes smaller than the original die area such that it is 20% - 80% (or 30% - 70%) of Ynm 2 .

[0065] For example, FIG. 10 shows an integrated system 1000 based on the Integrated Scaling and Stretching Platform (ISSP) of the present invention in comparison with a conventional one. As shown in FIG. 10, the ISSP integrated system 1000 and the conventional system 1010 include at least one single monolithic die 1011 having at least one processing device / circuit or main functional block (e.g., logic circuit 1011A and SRAM circuit 1011B), as well as a pad area 1011C. The integrated system 1000 provided by the ISSP of the present invention also includes at least one single monolithic die 1001 having a logic circuit 1001A, an SRAM circuit 801B, and a pad area 1001C. By comparing the configurations of the monolithic dies 1011 and 1001 between the conventional system 1010 and the ISSP integrated system 1000, it can be shown that the ISSP of the present invention can reduce the size of the integrated system (monolithic die 1001) without degrading the conventional performance, or add more devices within the same scanner maximum drawing screen area (monolithic die 1001').

[0066] From the perspective of reducing the size of the ISSP integrated system 1000, as shown in the center of FIG. 10, the single monolithic die 1001 of the ISSP integrated system 1000 has the same circuits or main functional blocks as the conventional monolithic die 1011 (i.e., the logic circuit 1001A and the SRAM circuit 1010B of the single monolithic die 1001 are identical to the logic circuit 1011A and the SRAM circuit 1011B of the single monolithic die 1011), and the single monolithic die 1001 only occupies 20% - 80% (or 30% - 70%) of the scanner maximum drawing screen area of the conventional monolithic die 1011.

[0067] In one embodiment, the combined area of the SRAM circuit 1001B and the logic circuit 1001A within a single monolithic die 1001 shrinks the area by 3.4 times that of the conventional monolithic die 1011. That is, compared with the conventional monolithic die 1011, the ISSP of the present invention leads to the area of the logic circuit 1001A of a single monolithic die 1001 being reduced by 5.3 times, the area of the SRAM circuit 1001B of a single monolithic die 1001 being reduced by 5.3 times, and (as shown in the center of FIG. 10) the combined area of the SRAM circuit 1001B and the logic circuit 1001A within a single monolithic die 1001 can be reduced by 3.4 times.

[0068] From another perspective of adding more devices, as shown on the right side of FIG. 10, the single monolithic die 1001' and the conventional monolithic die 1001 have the same scanner maximum drawing area. That is, the single monolithic die 1001' is fabricated based on the same technology node as that of the conventional monolithic die 1011 (such as 5 nm or 7 nm), and the area of the SRAM circuit 1001B' within the single monolithic die 1001' can not only include more SRAM cells, but also include additional major functional blocks not present in the conventional monolithic die 1011. In another embodiment of the present disclosure, the die area of the single monolithic die 1001' (as shown on the right side of FIG. 10) can be the same as, or approximately the same as, the scanner maximum drawing area (SMFA) of the conventional single monolithic die 1011 defined by a specific technology process node. That is, based on the ISSP of the present invention, within the scanner maximum drawing area (SMFA), there is additional space for accommodating additional SRAM cells, or additional major functional blocks other than those included in the conventional monolithic die 1011 (logic circuit 1011A and SRAM circuit 1001B).

[0069] FIG. 11(a) is a diagram showing another ISSP integration system 1100 of the present disclosure. The ISSP integration system 1100 includes at least one monolithic die 1101 having the size of the SMFA. The monolithic die 1101 includes a processing device / circuit (e.g., XPU1101A), an SRAM cache (including high-level and low-level caches), and an I / O circuit 1101B. Each of the SRAM caches includes a set of SRAM arrays. The I / O circuit 1101B is electrically connected to a plurality of SRAM caches and / or the XPU1101A.

[0070] In this embodiment, the monolithic die 1101 of the ISSP integration system 1100 includes different levels of caches L1, L2, and L3 generally composed of SRAM. Caches L1 and L2 (collectively referred to as "low-level caches") are usually allocated one for each CPU or GPU core device. Cache L1 is divided into L1i and L1d, which are respectively used to store instructions and data. Cache L2 does not distinguish between instructions and data. Cache L3 (which can be one of the "high-level caches") is shared by a plurality of cores and usually does not distinguish between instructions and data. Caches L1 / L2 are usually one for each CPU or GPU core.

[0071] In the case of high-speed operation, therefore, according to the ISSP of the present disclosure, the die area of the monolithic die 1101 can be the same as or approximately the same as the scanner maximum draw screen area (SMFA) defined by a specific technology process node. However, the storage capacities of the caches L1 / L2 (low-level caches) and cache L3 (high-level cache) of the ISSP integration system 1100 can be increased. As shown in FIG. 11(a), a GPU having a plurality of cores can have an SMFA (e.g., 26 mm × 33 mm, i.e., 858 mm) in which the high-level cache can have SRAM of 64 MB or more (e.g., 128 MB, 256, or 512 MB). 2) may be provided. Further, additional logic cores of the GPU can be inserted into the same SMFA to improve performance. The same applies to a memory controller (not shown) within the wideband I / O 1101B for another embodiment.

[0072] Alternatively, in addition to the existing major functional blocks, different other major functional blocks such as an FPGA can be integrated together within the same monolithic die. FIG. 11(b) is a diagram showing a single monolithic die 1101' of an ISSP integrated system 1100' according to another embodiment of the present disclosure. In this embodiment, the monolithic die 1101' includes at least one wideband I / O 1101B', as well as a plurality of processing devices / circuits such as an XPU 1101A' and a YPU 1101C. The processing devices (XPU 1101A' and YPU 1101C) have major functional blocks, and each of them can serve as an NPU, a GPU, a CPU, an FPGA, or a TPU (tensor processing unit). The major functional blocks of the XPU 1101A' can be different from those of the YPU 1101C.

[0073] For example, the XPU 1101A' of the ISSP integrated system 1100' may serve as a CPU, and the YPU 1101C of the ISSP integrated system 1100' may serve as a GPU. The XPU 1101A' and the YPU 1101C each have a plurality of logic cores, and each core has a low-level cache (e.g., a cache L1 / L2 having 512K or 1M / 128 bits), and a high-capacity high-level cache (e.g., a cache L3 having 32MB, 64MB or more) shared by the XPU 1101A' and the YPU 1101C. These three levels of caches may each include a plurality of SRAM arrays.

[0074] Due to the fact that GPUs are becoming increasingly important for AI learning, FPGAs have blocks of logic that interact with each other and may be designed by engineers to support specific algorithms and are suitable for AI inference. Thus, in some embodiments of the present disclosure, the ISSP integrated system 1100” having a single monolithic die 1101” may include a GPU and an FPGA, as shown in FIG. 11(c). The configuration of the monolithic die 1101” in FIG. 11(c) is the same as that of the monolithic die 1101’ in FIG. 11(b), except that the XPU 1101A” of the monolithic die 1101” is a GPU or a CPU, and the YPU 1101C’ of the monolithic die 1101” is an FPGA. By this approach, the monolithic die 1101” on the one hand has excellent parallel computing, learning speed, and efficiency, and on the other hand, it also has excellent AI inference capabilities with shorter time to market, lower cost, and flexibility.

[0075] Furthermore, as shown in FIG. 11(c), the processing devices / circuits (i.e., XPU 1101A” and YPU 1101C’) share a high-level cache (e.g., cache L3). The high-level cache (e.g., cache L3) shared between 1101A” and YPU 1101C’ can be configured by setting it in another mode register (not shown) or can be adaptively configured during the operation of the monolithic die 1101”. For example, in one embodiment, by setting the mode register, 1 / 3 of the high-level cache may be used by XPU 1101A”, and 2 / 3 of the high-level cache may be used by YPU 1101C’. Such allocation capacities of the high-level cache (e.g., cache L3) for XPU 1101A” or YPU 1101C’ can also be dynamically changed based on the operation of the integrated scaling and / or stretching platform (ISSP) for forming the integrated system 1100”.

[0076] FIG. 11(d) is a diagram showing a single monolithic die 1101''' of an ISSP integration system 1100''' according to yet another embodiment of the present disclosure. The arrangement of the monolithic die 1101''' in FIG. 11(d) is the same as that of the monolithic die 1101' in FIG. 11(b), except that the high-level cache includes cache L3 and cache L4, each processing device / circuit (such as XPU1101A''' and YPU1101C”) has cache L3 shared by its own core, and cache L4 having 32 MB or more is shared by the XPU and the YPU.

[0077] In some embodiments of the present disclosure, a somewhat large-capacity shared SRAM (or embedded SRAM, “eSRAM”) can be designed into one monolithic (single) die due to the smaller area of the SRAM cell design according to the present invention. Since the high memory capacity of eSRAM can be used, it is faster and more effective compared to conventional embedded DRAM or external DRAM. Thus, it is reasonable and possible to have high-bandwidth / high-memory-capacity SRAM in a single monolithic die having the same or substantially the same (e.g., 80% - 99%) die size as the scanner maximum draw screen area (SMFA, e.g., 26 mm × 33 mm, i.e., 858 mm 2 ).

[0078] Accordingly, the integrated system 1200 provided by the integrated scaling and / or stretching platform (ISSP) of the present disclosure includes at least two single monolithic dies, and these two monolithic dies may have the same or substantially the same size. For example, FIG. 12(a) is a diagram showing another ISSP integrated system 1200 according to another embodiment of the present disclosure in comparison with a conventional one 1210. The ISSP integrated system 1200 includes a single monolithic die 1201 and a single monolithic die 1202 within a single package. The single monolithic die 1201 mainly has a logic processing device circuit and a low-level cache formed therein, and the second monolithic die 1202 has only a plurality of SRAM arrays and I / O circuits formed therein. The plurality of SRAM arrays include at least 2 to 20 gigabytes, for example, 2 to 10 gigabytes.

[0079] As shown in FIG. 12(a), the single monolithic die 1201 mainly includes a logic circuit and an I / O circuit 1201A, and a small low-level cache (such as L1 and L2 caches) composed of an SRAM array 1201B. The single monolithic die 1202 includes only a high-bandwidth SRAM circuit 1202B having 2 to 10 gigabytes or more (for example, 1 to 20 gigabytes), and an I / O circuit 1202A for the high-bandwidth SRAM circuit 1202B. In this embodiment, the SMFA of the single monolithic die 1201 and the single monolithic die 1202 can be about 26 mm × 33 mm. If 50% of the SMFA of the single monolithic die 1202 (50% SRAM cell utilization rate) is used for the SRAM cells of the high-bandwidth SRAM circuit 1202B, the remaining of the SMFA is used for the I / O circuit of the high-bandwidth SRAM circuit 1202B.

[0080] FIG. 12(b) is a diagram showing a comparison result of SRAM cell areas of the integrated system 1200 of the present invention and those of three foundries based on different technology nodes. The total number of bytes (1 bit per SRAM cell) in the 26 mm × 33 mm SMFA of a single monolithic die (e.g., single monolithic die 1202) can be estimated by referring to the SRAM cell area as shown in FIG. 12(b). For example, in the present embodiment, the SMFA (26 mm × 33 mm) of the single monolithic die 1200 can accommodate 21 GB SRAM at a 5 nm technology node, and can provide more than 24 GB when the SRAM cell utilization rate can be increased.

[0081] According to FIG. 12(b), the conventional SRAM cell area (of the three foundries) may be 2 to 8 times that of the SRAM cell area of the present invention. Therefore, the ISSP integrated system 1200 can accommodate more bytes (1 bit per SRAM cell) in the 26 mm × 33 mm SMFA than the prior art ones. The total number of bytes (1 bit per SRAM cell) in the 26 mm × 33 mm SMFA based on different technology nodes is shown in Table 1 below.

[0082] [Table 1]

[0083] Of course, considering the various technologies proposed herein and the selective use of conventional wiring process technologies, the SMFA (26 mm × 33 mm) of the single monolithic die 1202 can accommodate a smaller capacity of SRAM, e.g., 1 / 4 to 3 / 4 times the SRAM sizes at the various technology nodes in Table 1 above. For example, the single monolithic die 1202 can accommodate approximately 5 to 15 GB SRAM, or 2.5 GB to 7.5 GB, by the selective use of the various technologies proposed herein and conventional wiring process technologies.

[0084] FIG. 13(a) is a diagram showing a single monolithic die 1301 of another ISSP integrated system 1300 according to the present invention. The arrangement of the single monolithic die 1301 is such that the single monolithic die 1301 of the present embodiment can be a high-performance computing (HPC) monolithic die including two or more main functional blocks such as a broadband I / O circuit 1301A, an XPU 1301B and a YPU 1101C each having a plurality of cores, and each core of the XPU 1301B and the YPU 1301C has its own cache L1 and / or cache L2 (L1~128KB, and L2~512KB to 1MB), which is the same as that of the single monolithic die 1201 in FIG. 12(a) except for this. The main functional blocks of the XPU 1301B or the YPU 1301C in FIG. 13(a) can be an NPU, a GPU, a CPU, an FPGA, or a TPU (tensor processing unit), and each of them has a main functional block. The XPU 1301B or the YPU 1301C can have different main functional blocks.

[0085] FIG. 13(b) is a diagram showing a single monolithic die 1302 of the ISSP integrated system 1300. The arrangement of the single monolithic die 1302 is the same as that of the single monolithic die 1202 in FIG. 12(a) except that the single monolithic die 1302 is a high-bandwidth SRAM (HBSRAM). In the present embodiment, the single monolithic die 1302 has the same (or having an area of about 80~90%) SMFA as the latest SMFA, and includes only a cache L3 and / or L4 having a plurality of SRAM arrays and a SRAM I / O circuit 1302A having a broadband 1302B I / O for the SRAM I / O circuit 1302A. The total SRAM in the single monolithic die 1302 can be 2~5GB, 5~10GB, 10~15GB, 15~20GB, or more depending on the utilization rate of the SRAM cells. Such a single monolithic die 1302 can be a high-bandwidth DRAM (HBSRAM).

[0086] As shown in FIGS. 13(a) and 13(b), the single monolithic die 1301 and the single monolithic die 1302 each have a wideband I / O bus, e.g., a 64-bit, 128-bit, or 256-bit data bus. The single monolithic die 1301 and the single monolithic die 1302 may be within the same IC package or within separate IC packages. For example, in some embodiments, the single monolithic die 1301 (e.g., the HPC die) is coupled to the single monolithic die 1302 (e.g., by wire bonding, flip chip bonding, solder bonding, 2.5D interposer silicon through via (TSV) bonding, 3D micro copper pillar direct bonding) as shown in FIG. 14 and housed within a single package to form the integrated system 1400. In embodiments, both the single monolithic die 1301 and the single monolithic die 1302 have the same or substantially the same SMFA, and thus such coupling may be accomplished by directly coupling a wafer 14A having at least the single monolithic die 1301 (or having multiple dies) to another wafer 14B having at least the single monolithic die 1302 (or having multiple dies), and then the coupled wafers 14A and 14B are sliced into a plurality of SMFA blocks to form the integrated system 1400 provided by the ISSP of the present disclosure. It is contemplated that another interposer having TSVs may be inserted between the single monolithic die 1301 and the single monolithic die 1302.

[0087] FIG. 15 is a diagram showing another ISSP integration system 1500 according to the present disclosure. The integration system 1500 includes two or more single monolithic dies 1302 (i.e., two HBSRAM dies as shown in FIG. 13(b)) coupled to each other, and one of the two single monolithic dies 1302 is then coupled to a single monolithic die 1301 (e.g., an HPC die as shown in FIG. 13(a)), and then all three or more dies are housed within a single package. Thus, such a package may include an HPC die and HBSRAM of more than 42, 48, or 96 GB. Of course, two or more single monolithic dies 1302 and a single monolithic die 1301 having a broadband I / O bus may be stacked vertically and coupled to each other based on the latest bonding technology.

[0088] Of course, three, four, or five or more HBSRAM dies may be integrated within a single package of the integration system 1500, and in that case, the cache L3 and L4 within the integration system 1500 may be larger than 128 GB or 256 GB SRAM. In some embodiments of the present disclosure, the single monolithic dies 1301 and 1302 of the integration system 1500 may be housed within the same IC package.

[0089] Compared with currently available HBM DRAM memory including approximately 24 GB based on a stack of 12 DRAM chips, the present invention can replace HBM3 memory with more HBSRAM (e.g., one HBSRAM chip having approximately 5 - 10 GB or 15 - 20 GB). Thus, within the ISSP, HMB memory is not required or only a small amount of HBM memory (e.g., less than 4 GB or 8 GB of HBM) is required.

[0090] Monolithic integration on a single die that enables the achievement of Moore's Law is currently facing its limits, particularly due to the limitations of photolithography technology. On the one hand, the minimum feature size printed on the die is very costly to shrink in that dimension, while on the other hand, the die size is limited by the maximum field area of the scanner. However, more and more diverse functions of the processor are emerging, and it is difficult to integrate them on a monolithic die. Furthermore, the small eSRAM on each major functional die and the somewhat overlapping presence of external or embedded DRAM are not desirable and optimized solutions. Based on the Integrated Scaling and / or Stretching Platform (ISSP) within a monolithic die or SOC die, (1) A single major functional block such as an FPGA, TPU, NPU, CPU, or GPU can be scaled down to a much smaller size, (2) More SRAM can be formed within the monolithic die, and, (3) Two or more major functional blocks, also scaled down through this ISSP, such as GPUs and FPGAs (or other combinations thereof), can be integrated together within the same monolithic die. (4) More levels of cache can exist within the monolithic die. (5) Such ISSP monolithic dies can be combined with another die (e.g., eDRAM) based on heterogeneous integration. (6) An HPC die 1 with L1&L2 caches is electrically connected (e.g., wire bonding or flip chip bonding) to one or more HBSRAM dies 2 utilized as L3&L4 caches within a single package, and each of the HPC die 1 and the HBSRAM die 2 has SMFA. (7) Within the ISSP, HMB memory is not required or only a small amount of HBM memory is required.

[0091] The present invention has been described by way of example and with reference to preferred embodiments, but it is to be understood that the invention is not limited thereto. In contrast, various modifications, as well as similar arrangements and procedures, are intended to be included, and accordingly, the appended claims should be given the broadest interpretation so as to include all such modifications, as well as all similar arrangements and procedures.

Claims

1. A first monolithic die in which a processing device circuit is formed therein, wherein the processing device circuit includes a first processing device circuit and a second processing device circuit, the first processing device circuit includes a plurality of first logic cores, each of the plurality of first logic cores corresponds to a first low-level cache, the second processing device circuit includes a plurality of second logic cores, each of the plurality of second logic cores corresponds to a second low-level cache, a main function performed by the first processing device circuit is different from a main function performed by the second processing device circuit, and the first processing device circuit or the second processing device circuit is selected from the group consisting of a graphics processing unit (GPU), a central processing unit (CPU), a tensor processing unit (TPU), a network processing unit (NPU), and a field programmable gate array (FPGA), the first monolithic die, and a second monolithic die in which a plurality of SRAM arrays are formed therein comprising, a transistor of the first processing device circuit, the second processing device circuit, or the plurality of SRAM arrays includes a semiconductor substrate including an initial semiconductor surface, a gate structure in the semiconductor substrate, a spacer covering sidewalls of the gate structure, a first trench and a second trench formed in the semiconductor substrate, a source region disposed in the first trench, a drain region disposed in the first trench, and a shallow trench isolation (STI) region surrounding the first trench and the second trench and having an upper surface not lower than an upper surface of the source region or the drain region, the first monolithic die is electrically connected to the second monolithic die, an integrated system in which a die area of the first monolithic die is the same as or substantially the same as a die area of the second monolithic die.

2. The die area of the first monolithic die is the same as or substantially the same as the maximum scanner image area defined by a specific technology process node and does not exceed 858 mm 2 . The integrated system according to Claim 1.

3. A localized isolation region under the initial semiconductor surface, the localized isolation region disposed under the source region or the drain region, and an edge of the source region or the drain region is aligned or substantially aligned with an edge of the gate structure or the spacer. The integrated system according to claim 1.

4. The first monolithic die and the second monolithic die are housed within a single package, wherein the plurality of SRAM arrays comprise at least 2 to 20 gigabytes. The integrated system according to claim 1.

5. The first low-level cache of the first monolithic die includes at least an L1 cache utilized by the processing device circuit during operation of the first monolithic die, The second low-level cache of the first monolithic die includes at least an L1 cache utilized by the processing device circuit during operation of the first monolithic die, wherein the plurality of SRAM arrays include at least an L3 cache utilized by the processing device circuit during operation of the first monolithic die. The integrated system according to claim 1.

6. Further comprising a third monolithic die in which a plurality of SRAM arrays are formed therein, wherein the third monolithic die comprises at least 2 to 20 gigabytes, wherein the first monolithic die, the second monolithic die, and the third monolithic die are housed within a single package, wherein the die area of the third monolithic die is the same as or substantially the same as the die area of the first monolithic die. The integrated system according to claim 1.

7. wherein the first monolithic die, the second monolithic die, and the third monolithic die are stacked vertically. The integrated system according to claim 6.

8. The upper surface of the STI region is higher than the initial semiconductor surface, wherein the source region or the drain region has a substantially upright sidewall surrounded by the STI region. The integrated system according to claim 1.

9. An integrated system, A first monolithic die having a processing device circuit, wherein the processing device circuit includes a first processing device circuit and a second processing device circuit, the first processing device circuit includes a plurality of first logic cores, each of the plurality of first logic cores corresponds to a first SRAM size including an L1 cache and an L2 cache, the second processing device circuit includes a plurality of second logic cores, each of the plurality of second logic cores corresponds to a second SRAM size, a main function performed by the first processing device circuit is different from a main function performed by the second processing device circuit, and the first processing device circuit or the second processing device circuit is selected from the group consisting of a GPU, a CPU, a TPU, an NPU, and an FPGA, the first monolithic die, and a second monolithic die having a plurality of SRAM arrays are provided, One transistor of the first processing device circuit, the second processing device circuit, or the plurality of SRAM arrays has a semiconductor substrate having an initial semiconductor surface, a gate structure in the semiconductor substrate, a spacer covering sidewalls of the gate structure, a first trench and a second trench formed in the semiconductor substrate, a source region disposed in the first trench, a drain region disposed in the first trench, and a shallow trench isolation (STI) region surrounding the source region and the drain region and having an upper surface higher than the initial semiconductor surface. The first monolithic die is physically separated from the second monolithic die, The first monolithic die is electrically connected to the second monolithic die, The die area of the first monolithic die is the same as or substantially the same as the die area of the second monolithic die, and the die area of the first monolithic die does not exceed 858 mm 2 beyond.

10. further comprising a metal contact plug deposited in a hole between the STI region and the gate structure. The integrated system according to claim 9.

11. The die area of the first monolithic die is the same as or substantially the same as the maximum scanner image area defined by a specific technology process node, the first monolithic die and the second monolithic die are housed in a single package, The first monolithic die is electrically connected to the second monolithic die by wire bonding, flip chip bonding, solder bonding, interposer silicon through - via (TSV) bonding, or micro - copper pillar direct bonding. The integrated system according to claim 9.

12. The integrated system does not include a high - bandwidth memory (HBM). The integrated system according to claim 9.

13. A first monolithic die having a processing device circuit, wherein the processing device circuit includes a first processing device circuit and a second processing device circuit, the first processing device circuit includes a plurality of first logic cores, each of the plurality of first logic cores corresponds to a first SRAM size including an L1 cache and an L2 cache, the second processing device circuit includes a plurality of second logic cores, each of the plurality of second logic cores corresponds to a second SRAM size, the main function performed by the first processing device circuit is different from the main function performed by the second processing device circuit, and the first processing device circuit or the second processing device circuit is selected from the group consisting of a GPU, a CPU, a TPU, an NPU, and an FPGA, a first monolithic die, and A second monolithic die having a plurality of SRAM arrays comprising One transistor of the first processing device circuit, the second processing device circuit, or the plurality of SRAM arrays has a semiconductor substrate having an initial semiconductor surface, a gate structure in the semiconductor substrate, a spacer covering sidewalls of the gate structure, a first trench and a second trench formed in the semiconductor substrate, a source region disposed in the first trench, and a drain region disposed in the first trench, wherein the source region or the drain region includes an epitaxially doped semiconductor layer and a metal layer in contact with the outermost sidewall of the epitaxially doped semiconductor layer, the source region and the drain region, and a shallow trench isolation (STI) region surrounding the first trench and the second trench. The first monolithic die is electrically connected to the second monolithic die. The first monolithic die and the second monolithic die are housed within a single package, or the first monolithic die and the second monolithic die are respectively housed within a first package and a second package, The die area of the first monolithic die is the same as or substantially the same as the die area of the second monolithic die, and the die area of the first monolithic die does not exceed 858 mm 2 , integrated system.

14. The upper surface of the STI region is higher than the initial semiconductor surface, The integrated system according to claim 13.

15. The source region or the drain region has substantially upright sidewalls surrounded by the STI region, The integrated system according to claim 14.

Citation Information

Patent Citations

  • Semiconductor device

    JP2006165480A

  • Integrated circuit chip using top post-passivation technology and bottom structure technology

    JP2015073107A

  • Methods of fabricating integrated circuit devices including strained channel regions and related devices

    US20080191244A1

  • Managing Cache Coherence Using Information in a Page Table

    US20170337136A1

  • Logic drive using standard commodity programmable logic IC chips comprising non-volatile random access memory cells

    US20190238135A1