Homogeneous / Heterogeneous Integration System with High-Performance Computing and High Memory Capacity
The integrated system with self-aligned connections and localized isolations addresses the challenges of miniaturization and integration in AI chips, achieving reduced area and increased device density by integrating multiple functional blocks within a single die, thereby enhancing computing performance and memory capacity.
Patent Information
- Application Number
- JP2022201384
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-16
- Filing Date
- 2022-12-16
- Publication Date
- 2025-07-02
- Estimated Expiration
- 2042-12-16
Smart Images

Figure 0007701591000003 
Figure 0007701591000004 
Figure 0007701591000005
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to semiconductor structures, and more particularly to integrated systems including logic chips with high-performance computing and SRAM chips with large memory capacity.
Background Art
[0002] Information technology (IT) systems are evolving rapidly in every enterprise and business, including those in factories, healthcare, and transportation. Today, system-on-chip (SOC) or artificial intelligence (AI) has become the core of IT systems that make factories smarter, improve patient prognosis, and enhance the safety of autonomous vehicles. Data from manufacturing equipment, sensors, and machine vision systems can easily total 1 petabyte per day. Therefore, to handle such petabyte-scale data, high-performance computing (HPC) SPC or AI chips are required.
[0003] Generally, AI chips can be classified into graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs). GPUs were originally designed to process graphical processing applications using parallel computing, but have increasingly been used for AI learning. The learning speed and efficiency of GPUs are generally 10 to 1000 times greater than those of general-purpose CPUs.
[0004] FPGAs have logic blocks that interact with each other and may be designed by engineers to support specific algorithms and are suitable for AI inference. Although FPGAs have drawbacks such as larger size, slower speed, and higher power consumption, they are preferred over ASIC design due to the short time to market, low cost, and flexibility. Due to the flexibility of FPGAs, any part of the FPGA can be partially programmed according to requirements. The inference speed and efficiency of FPGAs are 10 to 100 times greater than those of general-purpose CPUs.
[0005] On the one hand, ASICs are directly adapted to circuits and are generally more efficient than FPGAs. In the case of customized ASICs, their learning / inference speed and efficiency can be 10 to 1000 times greater than those of general-purpose CPUs. However, as AI algorithms continue to evolve, unlike FPGAs where customization is becoming easier, ASICs are gradually becoming obsolete as new AI algorithms are developed.
[0006] In any of GPUs, FPGAs, and ASICs (or similar SOCs, CPUs, NPUs, etc.), logic circuits and SRAM circuits are two major circuits, and their combination generally accounts for approximately 90% of the AI chip size. The remaining 10% of the AI chip may include I / O pad circuits. However, the miniaturization process / technology nodes for manufacturing AI chips are becoming increasingly necessary to efficiently and quickly train AI machines as they bring better efficiency and performance. The improvement of integrated circuit performance and cost has mainly been achieved by process miniaturization technology according to Moore's law. However, such miniaturization technologies up to 3nm - 5nm are facing many technical problems, and thus, the investment costs and capital in the research and development of the semiconductor industry are increasing exponentially.
[0007] For example, the miniaturization of SRAM devices for increased memory density, the reduction in operating voltage (VDD) for lower standby power consumption, and the improved yield required to achieve larger-capacity SRAMs are becoming more difficult to realize as the miniaturization to a 28nm (or lower) manufacturing process poses challenges.
[0008] FIG. 1 shows a SRAM cell architecture, i.e., a 6-transistor (6-T) SRAM cell. It consists of two cross-coupled inverters (PMOS pull-up transistors PU-1 and PU-2, and NMOS pull-down transistors PD-1 and PD-2), and two access transistors (NMOS passgate transistors PG-1 and PG-2). A high-level voltage VDD is coupled to the PMOS pull-up transistors PU-1 and PU-2, and a low-level voltage VSS is coupled to the NMOS pull-down transistors PD-1 and PD-2. When the word line (WL) is enabled (i.e., the row is selected in the array), the access transistors are turned on to connect the storage nodes (Node-1 / Node-2) to the bit lines (BL and BL Bar) extending vertically.
[0009] FIG. 2 shows a "stick diagram" representing the layout and connection of the six transistors of the SRAM. The stick diagram typically includes only the active regions (vertical bars) and gate lines (horizontal bars) to form the pull-down transistor PD and pull-up transistor PU of the six transistors of the SRAM. Of course, on the one hand, there are still many contacts directly coupled to the six transistors, and on the other hand, coupled to the word line (WL), bit lines (BL and BL Bar), high-level voltage VDD, and low-level voltage VSS, etc.
[0010] λ when the minimum feature size decreases 2 or F 2Part of the reason for the dramatic increase in the total area of the SRAM cell represented by can be explained as follows. Conventional 6T SRAM has six transistors connected by using a plurality of interconnections, and has a first interconnect layer M1 for connecting the gate level ("gate") of the transistor and the diffusion levels of the source and drain regions (regions generally called "diffusion"). In order to facilitate signal transmission without increasing the die size by only using M1, it is necessary to increase the second interconnect layer M2 and / or the third interconnect layer M3 (for example, word line (WL) and / or bit line (BL and BL Bar)), and therefore, a structure Via-1 composed of a specific type of conductive material is formed to connect M2 to M1.
[0011] Therefore, there exists a vertical structure formed from diffusion through the contact (Con) connection to M1, that is, "diffusion-Con-M1". Similarly, another structure for connecting the gate to M1 through the contact structure can be formed as "gate-Con-M1". Furthermore, when a connection structure for connecting from the M1 interconnect to the M2 interconnect through via 1 needs to be formed, it is called "M1-via1-M2". A more complex structure from the gate level to the M2 interconnect can be represented as "gate-Con-M1-via1-M2". Furthermore, the stacked interconnect system can have a structure such as "M1-via1-M2-via2-M3" or "M1-via1-M2-via2-M3-via3-M4".
[0012] In the conventional SRAM, since the gates and diffusions in the two access transistors (NMOS pass gate transistors PG-1 and PG-2 as shown in FIG. 1) are connected to the word line (WL) and / or the bit line (BL and BL Bar), which are arranged in the second interconnect layer M2 or the third interconnect layer M3, such metal connections must first pass through the interconnect layer M1. That is, the latest interconnect system in the SRAM may not allow the gate or diffusion to be directly connected to M2 without bypassing the M1 structure.
[0013] As a result, the required space between one M1 interconnect and the other M1 interconnect increases the die size, and in some cases, the wiring connection may inhibit the intention of a specific efficient channeling that directly uses M2 to cross the M1 region. Furthermore, it is difficult to form a self-alignment structure between via 1 and the contact, and at the same time, both via 1 and the contact are connected to their respective interconnect systems themselves.
[0014] Furthermore, in a conventional 6T SRAM, at least, there is one NMOS transistor and one PMOS transistor respectively arranged inside some adjacent regions of the p-substrate and the n-well, which are formed adjacent to each other within a close range. A parasitic junction structure called an n+ / p / n / p+ parasitic bipolar device is formed along its contour starting from the n+ region of the NMOS transistor to the p-well, to the adjacent n-well, and further to the p+ region of the PMOS transistor.
[0015] There is a large noise generated at the n+ / p junction or the p+ / n junction, and a very large current abnormally flows through this n+ / p / n / p+ junction, which may, in some cases, block the operation of some parts of the CMOS circuit and cause malfunction of the entire chip. Such an abnormal phenomenon called latch-up affects the CMOS operation and must be avoided. Certainly, one way to enhance the resistance to latch-up, which is indeed a weakness of CMOS, is to increase the distance from the n+ region to the p+ region. As a result, increasing the distance from the n+ region to the p+ region to avoid the latch-up problem also enlarges the size of the SRAM cell.
[0016] However, even with the miniaturization of the manufacturing process to below 28 nm (so-called "minimum processing dimension", "lambda (λ)", or "F"), λ 2 or F 2The total area of the SRAM cell represented by is shown in Fig. 3 to increase exponentially as the minimum feature size decreases due to interference between the sizes of the contacts, interference between the layouts of the metal wiring connecting the word line (WL), bit lines (BL and BL Bar), high-level voltage VDD, and low-level voltage VSS, etc. (quoted from "15.1 A 5nm 135Mb SRAM in EUV and High-Mobility-Channel FinFET Technology with Metal Coupling and Charge-Sharing Write-Assist Circuitry Schemes for High-Density and Low-VMIN Applications, 2020 IEEE International Solid- State Circuits Conference - (ISSCC), 2020, pp. 238-240)" by J. Chang et al.).
[0017] A similar situation also occurs in the miniaturization of logic circuits. The miniaturization of logic circuits for increased memory density, the reduction of the operating voltage (Vdd) for low standby power consumption, and the increased yield required to realize a larger-capacity logic circuit are becoming increasingly difficult to achieve. Standard cells are commonly used and are the basic elements within a logic circuit. A standard cell may include basic logic function cells (e.g., inverter cell, NOR cell, NAND cell).
[0018] Similarly, in the miniaturization of the manufacturing process to below 28nm, the λ 2 or F 2 The total area of the standard cell represented by increases exponentially as the minimum feature size decreases due to the size of the contacts and interference between the layouts of the metal wiring.
[0019] Figure 4(a) shows a "stick diagram" representing the layout and connections of PMOS and NMOS transistors in a 5nm (UHD) standard cell of a semiconductor company. The stick diagram includes only the active regions (horizontal lines) and gate lines (vertical lines). Hereinafter, the active regions can be called "fins". Of course, on the one hand, they are directly connected to the PMOS and NMOS transistors, and on the other hand, there are still many contacts connected to the input terminals, output terminals, high-level voltage Vdd, and low-level voltage VSS (or ground "GND"), etc. In particular, each transistor includes two active regions or fins (marked by the gray dashed rectangles) to form the channel of the transistor, so that the W / L ratio can be maintained within an acceptable range. The area size of the inverter cell is equal to X×Y, where X = 2×Cpp, Y = cell height, and Cpp is the distance of the contact to the poly pitch (Cpp).
[0020] Some of the active regions or fins between the PMOS and NMOS (referred to as "dummy fins") are not utilized within the PMOS / NMOS of this standard cell, and it can be seen that the potential reason is likely related to the latch-up issue between the PMOS and NMOS. Therefore, the latch-up distance between the PMOS and NMOS in Figure 4(a) is 3×Fp, where Fp is the fin pitch. Based on the available data regarding Cpp (54nm) and cell height (216nm) in the 5nm standard cell, the cell area is 23328nm 2 (or 933.12λ 2 where lambda (λ) is the minimum feature size of 5nm) can be calculated by X×Y. Figure 4(b) shows the aforementioned 5nm standard cell and its dimensions. As shown in Figure 4(b), the latch-up distance between the PMOS and NMOS is 15λ, Cpp is 10.8λ, and the cell height is 43.2λ.
[0021] The miniaturization trend regarding the area size (2Cpp × cell height) of three fundries with respect to different process technology nodes can be shown in Fig. 5. (For example, from 22 nm to 5 mm), as the technology node decreases, λ 2 In terms of conversion, it is obvious that the area size of the conventional standard cell (2Cpp × cell height) increases dramatically. Within the conventional standard cell, the smaller the technology process node, the higher the area size in terms of λ 2 conversion. Such a dramatic increase in λ 2 can be brought about by the difficulty of proportionally reducing the size of gate contacts / source contacts / drain contacts as λ decreases, whether in SRAM or logic circuits, the difficulty of proportionally reducing the latch-up distance between PMOS and NMOS, and interference within the metal layer due to the decrease in λ, etc.
[0022] From another perspective, any high-performance computing (HPC) chips such as SOC, AI, NPU (network processing unit), GPU, CPU, and FPGA, etc. are currently using monolithic integration to place as many and more circuits as possible. However, as shown in Fig. 6(a), maximizing the die area of each monolithic die is limited by the maximum reticle size of the lithography stepper, which is difficult to expand due to existing state-of-the-art photolithography exposure tools. For example, as shown in Fig. 6(b), current i193 and EUV lithography steppers have a maximum reticle size, thus, the monolithic SOC die has a scanner maximum field area (SMFA) of 26 mm × 33 mm, that is, 858 mm 2 . However, for high-performance computing or AI purposes, high-end consumer GPUs are 500 - 600 mm 2It seems to extend to. As a result, it has become more difficult or impossible to implement two or more major functional blocks (e.g., GPU and FPGA) on a single die. Furthermore, since the most widely used six-transistor CMOS SRAM cell is extremely large, the size of the embedded SRAM (eSRAM) sufficient for both major blocks needs to be increased. Additionally, although the external DRAM capacity needs to be expanded, discrete PoP (package-on-package, e.g., HBM~SOC) or POD (package DRAM on SOC die) is still restricted by the difficulty of achieving the desired performance of the signal interconnections between die-chip or package-chip.
[0023] Therefore, in the near future, there is a need to propose a new integrated system including a logic chip with HPC and an SRAM chip with high memory capacity that can solve the above-mentioned problems so that a more powerful and efficient SOC or single chip of AI based on monolithic integration can be realized.
Summary of the Invention
[0024] One aspect of the present disclosure provides an integrated system, where the integrated system includes a first monolithic die and a second monolithic die. The first monolithic die has a processing device circuit formed therein, and the second monolithic die has a plurality of SRAM arrays formed therein. The second monolithic die is provided with at least 2 gigabytes, and the first monolithic die is electrically connected to the second monolithic die.
[0025] In one embodiment of the present disclosure, the first monolithic die has a die area that is the same as or substantially the same as the maximum scanner image area defined by a specific technology process node, and the second monolithic die has a die area that is the same as or substantially the same as the maximum scanner image area defined by the specific technology process node.
[0026] In one embodiment of the present disclosure, the maximum drawing area of the scanner is 26 mm × 33 mm, or not greater than 858 mm 2 or less.
[0027] In one embodiment of the present disclosure, the first monolithic die and the second monolithic die are housed within a single package.
[0028] In one embodiment of the present disclosure, the plurality of SRAM arrays include at least 20 gigabytes.
[0029] In one embodiment of the present disclosure, the processing device circuit includes a first processing device circuit and a second processing device circuit. The first processing device circuit includes a plurality of first logic cores, each of the plurality of first logic cores includes a first SRAM size, the second processing device circuit includes a plurality of second logic cores, and each of the plurality of second logic cores includes a second SRAM size.
[0030] In one embodiment of the present disclosure, the main function performed by the first processing device circuit is different from the main function performed by the second processing device circuit.
[0031] In one embodiment of the present disclosure, the first processing device circuit or the second processing device circuit is selected from the group consisting of a graphics processing unit (GPU), a central processing unit (CPU), a tensor processing unit (TPU), a network processing unit (NPU), and a field programmable gate array (FPGA).
[0032] In one embodiment of the present disclosure, the first monolithic die further includes an L1 cache and an L2 cache utilized by the processing device circuit during operation of the first monolithic die, and the plurality of SRAM arrays include an L3 cache and an L4 cache utilized by the processing device circuit during operation of the first monolithic die.
[0033] In one embodiment of the present disclosure, the integrated system further includes a third monolithic die in which a plurality of SRAM arrays are formed. The third monolithic die includes at least 2 to 20 gigabytes. The first monolithic die, the second monolithic die, and the third monolithic die are housed within a single package. The first monolithic die has a die area that is the same as or substantially the same as the maximum scanner image area defined by a specific technology process node. The second monolithic die has a die area that is the same as or substantially the same as the maximum scanner image area defined by the specific technology process node. The third monolithic die has a die area that is the same as or substantially the same as the maximum scanner image area defined by the specific technology process node.
[0034] In one embodiment of the present disclosure, the first monolithic die, the second monolithic die, and the third monolithic die are stacked vertically.
[0035] Another aspect of the present disclosure provides an integrated system that includes a first monolithic die and a second monolithic die. The first monolithic die has a processing device circuit. The second monolithic die has a plurality of SRAM arrays. The plurality of SRAM arrays include at least 2 gigabytes. The first monolithic die is physically separated from the second monolithic die. The first monolithic die is electrically connected to the second monolithic die. The integrated system does not include a high bandwidth memory (HBM).
[0036] In one embodiment of the present disclosure, the second monolithic die has an area size that is the same as or substantially the same as the maximum scanner image area defined by a specific technology process node. The first monolithic die has an area size that is the same as or substantially the same as the maximum scanner image area defined by the specific technology process node.
[0037] In one embodiment of the present disclosure, the first monolithic die and the second monolithic die are housed within a single package, and the first monolithic Die is electrically connected to the second monolithic Die by wire bonding, flip chip bonding, solder bonding, interposer silicon through electrode (TSV) bonding, or micro copper pillar direct bonding.
[0038] Another aspect of the present disclosure provides an integrated system, the integrated system including a first monolithic circuit and a second monolithic circuit. The first monolithic circuit has a processing device circuit, the second monolithic circuit has a plurality of SRAM arrays, the plurality of SRAM arrays including at least 2 to 20 gigabytes, the first monolithic circuit being electrically connected to the second monolithic circuit, the first monolithic circuit being formed within a first monolithic die, the second monolithic circuit being formed within a second monolithic die, the first monolithic die and the second monolithic die being housed within a single package, or the first monolithic die and the second monolithic die being housed within a first package and a second package, respectively.
Brief Description of the Drawings
[0039] The above and other aspects of the present disclosure will be better understood with respect to the following detailed description of the preferred but non-limiting (plural) embodiments. The following description is made with reference to the accompanying drawings.
[0040]
Figure 1
Figure 2
Figure 3
Figure 4(a)
Figure 4(b)
Figure 5
Figure 6(a)
Figure 6(b)
Figure 7(a)
Figure 7(b)
Figure 7(c)
Figure 8(a)
Figure 8(b)
Figure 9(a)
Figure 9(b)
Figure 10
Figure 11(a)
Figure 11(b)
Figure 11(c)
Figure 11(d)
Figure 12(a)
Figure 12(b)
Figure 13(a)
Figure 13(b)
Figure 14
Figure 15
[0041] The present disclosure provides an integrated system. The above and other aspects of the present disclosure will be better understood from the following detailed description of the preferred but non-limiting embodiments. The following description is made with reference to the accompanying drawings:
[0042] Some embodiments of the present disclosure are disclosed below with reference to the accompanying drawings. However, the structures and contents disclosed in the embodiments are for illustrative and explanatory purposes only, and the scope of protection of the present disclosure is not limited to the embodiments. The present disclosure does not show all possible embodiments, and it should be noted that those skilled in the art in the technical field of the present disclosure can make appropriate modifications or changes based on the present specification disclosed below to meet actual needs without departing from the spirit of the present disclosure. The present disclosure is applicable to other implementation forms not disclosed in this specification.
[0043] Embodiment 1
[0044] The present disclosure proposes incorporating the following invention. a. A new transistor (presented in U.S. Patent Application No. 17 / 138,918, filed on December 31, 2020, entitled "MINIATURIZED TRANSISTOR STRUCTURE WITH CONTROLLED DIMENSIONS OF SOURCE / DRAIN AND CONTACT-OPENING AND RELATED MANUFACTURE METHOD", the entire content of U.S. Patent Application No. 17 / 138,918 is incorporated herein by reference; presented in U.S. Patent Application No. 16 / 991,044, filed on August 12, 2020, entitled "TRANSISTOR STRUCTURE AND RELATED INVERTER", the entire content of U.S. Patent Application No. 16 / 991,044 is incorporated herein by reference; and presented in U.S. Patent Application No. 17 / 318,097, filed on May 12, 2021, entitled "COMPLEMENTARY MOSFET STRUCTURE WITH LOCALIZED ISOLATIONS IN SILICON SUBSTRATE TO REDUCE LEAKAGES AND PREVENT LATCH-UP", the entire content of U.S. Patent Application No. 17 / 318,097 is incorporated herein by reference), b. Interconnection with the transistor (presented in U.S. Patent Application No. 17 / 528,957, filed on November 17, 2021, entitled "INTERCONNECTION STRUCTURE AND MANUFACTURE METHOD THEREOF", the entire content of U.S. Patent Application No. 17 / 528,957 is incorporated herein by reference), c. SRAM cell (presented in U.S. Patent Application No. 17 / 395,922, filed on August 6, 2021, entitled "NEW SRAM CELL STRUCTURES", the entire content of U.S. Patent Application No. 17 / 395,922 is incorporated herein by reference), and d. Standard cell design (presented in U.S. Provisional Patent Application No. 63 / 238,826, filed on August 31, 2021, entitled "STANDARD CELL STRUCTURES", the entire content of which is incorporated herein by reference)
[0045] For example, FIG. 7(a) is a top view showing a MOSFET structure according to an embodiment of the present disclosure. FIG. 7(b) is a cross-sectional view taken along the cutting line C7J1 as represented in FIG. 7(a). FIG. 7(c) is a cross-sectional view taken along the cutting line C7J2 as represented in FIG. 7(a). In the proposed MOSFET, the silicon regions of the gate terminal (e.g., silicon region 702c) and the source / drain terminals are exposed, and the selective epitaxial growth technology (SEG) has a seed region for growing pillars (e.g., the first conductor pillar portion 731a and the third conductor pillar portion 731b) based on the seed region.
[0046] Furthermore, the first conductor pillar portion 731a and the third conductor pillar portion 731b each further have a seed region or a seed pillar within their upper portions, and such seed regions or seed pillars can be used for subsequent selective epitaxial growth. Thereafter, the second conductor pillar portion 732a is formed on the first conductor pillar 731a by a second selective epitaxial growth, and the fourth conductor pillar 732b is formed on the third conductor pillar portion 731b.
[0047] As shown in FIGS. 7(a) to 7(c), in this embodiment, as long as a seed portion or a seed pillar exists on top of a conductive terminal and a conductor pillar portion configured for a subsequent selective epitaxial growth method, where the M1 interconnect (a type of conductive terminal) or a conductive layer is involved, it can be applied to enable self-aligned connection directly to the MX interconnect layer (without connecting to the conductive layers M2, M3,.. MX-1) through one vertical conductive or conductor plug. The seed portion or the seed pillar is not limited to silicon, and any material that can be used as a seed configured for subsequent selective epitaxial growth is acceptable.
[0048] FIG. 8(a) is a top view showing a combined structure of a PMOS transistor 52 and an NMOS transistor 51 according to an embodiment of this embodiment. FIG. 8(b) is a cross-sectional view of the PMOS transistor 52 and the NMOS transistor 51 obtained along the cutting line (X-axis) in FIG. 8(a). The structure of the PMOS transistor 52 is the same as that of the NMOS transistor 51. A gate structure 33 including a gate dielectric layer 331 and a gate conductive layer 332 (for example, gate metal) is formed on a horizontal or initial surface of a semiconductor substrate (such as a silicon substrate). A dielectric cap 333 (such as a composite of an oxide layer and a nitride layer) is on the gate conductive layer 332. Further, a spacer 34 that may include a composite of an oxide layer 341 and a nitride layer 342 is used to cover the sidewalls of the gate structure 33. Trenches are formed in the silicon substrate, and all or at least a part of the source region 55 and the drain region 56 are respectively disposed in the corresponding trenches. The source (or drain) region in the MOS transistor 52 may include an N+ region, or other suitable doping profile regions (for example, a gradual or stepwise change from P- regions and P+ regions).
[0049] Furthermore, localized isolations 48 (such as nitride or other high-k dielectric materials) are placed within one trench and disposed under the source region, and another localized isolation 48 is placed within another trench and disposed under the drain region. Such localized isolations 48 are below the horizontal silicon surface (HSS) of the silicon substrate and can be referred to as localized isolations in silicon substrate (LISS) 48. LISS 48 can be a thick nitride layer or a composite of dielectric layers. For example, the localized isolation or LISS 48 can comprise a composite localized isolation including an oxide layer 481 covering at least a part of the sidewalls of the trench and another oxide layer 482 covering at least a part of the bottom wall of the trench. The oxide layers 481 and 482 can be L-shaped oxide layers formed by thermal oxidation treatment.
[0050] The composite localized isolation 48 can further include a nitride layer 483 on the oxide layer 482 and / or the oxide layer 481. The shallow trench isolation (STI) region may comprise a composite STI 49 including an STI-1 layer 491 and an STI-2 layer 492, and the STI-1 layer 491 and the STI-2 layer 492 may each be made of a thick oxide material by different processes.
[0051] Furthermore, the source (or drain) region can comprise a composite source region 55 and / or a drain region 56. For example, in the NMOS transistor 52, the composite source region 55 (or drain region 56) includes at least a lightly doped drain (LDD) 551 and an N+ highly doped region 552 in the trench. Note that the lightly doped drain (LDD) 551 particularly abuts an exposed silicon surface having a uniform (110) crystal orientation. The exposed silicon surface has a vertical boundary with an appropriate thickness embedded, in contrast to the edge of the gate structure. The exposed silicon surface is substantially aligned with the gate structure. The exposed silicon surface can be the end face of the channel of the transistor.
[0052] The low-doped drain (LDD) 551 and the N+ highly doped region 552 are used as crystal seeds for growing silicon from the exposed TEC region for forming a new orderly (110) lattice across the LISS region that has no seeding effect on the changing (110) crystal structure of the newly formed crystals in the composite source region 55 or drain region 56, based on selective epitaxial growth (SEG) technology (or other suitable techniques such as atomic layer deposition ALD or selective growth ALD-SALD). Such newly formed crystals (including the low-doped drain (LDD) 551 and the N+ highly doped region 552) can be referred to as TEC-Si.
[0053] In one embodiment, the TEC is aligned or substantially aligned with the edge of the gate structure 33, the length of the LDD 551 is adjustable, and the sidewalls of the LDD 551 on the side opposite to the TEC are aligned or substantially aligned with the sidewalls of the spacer 34. The composite source (or drain) region may further include several tungsten (or other suitable metal materials, such as TiN / tungsten) plugs 553 formed in the lateral connection with the TEC-Si part for the completion of the entire source / drain region. The active channel current flowing to future metal interconnections such as the metal-1 layer reaches the tungsten 553 (or other metal materials) directly connected to the metal-1 through the LDD 551 and the N+ highly doped region 552 by several good metal-metal ohmic contacts having a much lower resistance than the conventional silicon-metal contact.
[0054] The source / drain contact resistance of the NMOS transistor 52 can be maintained within an appropriate range by the structure of the fused metal-semiconductor junction utilized within the source / drain structure. This fused metal-semiconductor junction within the source / drain structure can improve the current crowding effect and reduce the contact resistance. Furthermore, since the bottom of the source / drain structure is insulated from the substrate by the bottom oxide (oxide layer 482), the isolation between n+~n+ or p+~p+ can be maintained within an appropriate range. Therefore, the spacing between two adjacent active regions of the PMOS transistor (not shown) can be reduced to 2λ. The bottom oxide (oxide layer 482) may significantly reduce the source / drain junction leakage current, thereby reducing the leakage current between n+~n+ or p+~p+.
[0055] TIFF0007701591000001.tif66164
[0056] Furthermore, the composite STI 49 is heightened (for example, the STI-2 layer 492 extends from the initial semiconductor surface to above the top surface of the gate structure), and thus, the selectively grown source / drain regions may be confined by the composite STI and not exceed the composite STI 49. The metal contact plug (such as the tungsten plug 553) can be deposited into the hole between the composite STI 49 and the gate structure without using another contact mask to form the contact hole. Furthermore, the top surface and one sidewall of the high-concentration doped region 552 can be directly contacted by the metal contact plug, and the contact resistance of the source / drain region can be drastically reduced.
[0057] Furthermore, in a conventional design, metal wirings for high-level voltage Vdd and low-level voltage Vss (or ground) are distributed on the initial silicon surface of a silicon substrate, and such distribution interferes with a plurality of other metal wirings when there is not enough space between those metal wirings. The present invention discloses a new standard cell or SRAM cell in which metal wirings for high-level voltage Vdd and / or low-level voltage Vss can be distributed under the initial silicon surface of a silicon substrate, so that even when the size of a standard cell is reduced, interference between contact sizes and interference between layouts of metal wirings connecting high-level voltage Vdd, low-level voltage Vss, etc. can be avoided.
[0058] For example, within the drain region of NMOS51, tungsten or other metal material 553 is directly coupled (by removing LISS48) to the P-well that is electrically coupled to Vdd. Similarly, within the source region of NMOS51, tungsten or other metal material 553 is directly coupled (by removing LISS48) to the p-well or P-substrate that is electrically coupled to ground. Thus, the openings of the source / drain regions originally used to electrically couple the source / drain regions to the metal-2 layer (M2) or metal-3 layer (M3) for Vdd or ground connection can be omitted in the new standard cell and the standard cell.
[0059] In summary, at least the following advantages exist. (1) The source, drain, and gate length dimensions of transistors within a standard cell / SRAM can be precisely controlled, and the length dimensions can be as small as the minimum feature size lambda (λ) as shown in the incorporated U.S. Patent Application No. 17 / 138,918. Thus, when two adjacent transistors are connected to each other through the drain / source, the length dimension of the transistor is as small as about 3λ, and the distance between the edges of the gates of two adjacent transistors is as small as about 2λ. Of course, due to tolerances, the length dimension of the transistor can be about 3λ to 6λ or more, and the distance between the edges of the gates of two adjacent transistors can be 8λ or more. (2) The first metal interconnect (M1 layer) is directly connected to the gate, source, and / or drain regions through self-aligned miniaturized contacts without using a conventional contact hole opening mask and / or a metal-0 conversion layer for M1 connections. (3) The gate and / or diffusion (source / drain) regions are directly connected to the metal-2 (M2) interconnect layer in a self-aligned manner without connecting to the metal-1 layer (M1). Thus, the required space between one metal-1 layer (M1) interconnect layer and the other metal-1 layer (M1) interconnect layer, and the blocking problem in some wiring connections are reduced. Furthermore, the same structure is applicable even when the lower metal layer is directly connected to the upper metal layer by a conductor pillar but the conductor pillar is not electrically connected to any intermediate metal layer between the lower metal layer and the upper metal layer. (4) The metal wiring of the high-level voltage Vdd and / or the low-level voltage VSS within the standard cell can be distributed under the initial silicon surface of the silicon substrate. Thus, the interference between the sizes of the contacts, the high-level voltage Vdd, and the layout of the metal wiring connecting the low-level voltage Vss, etc., can be avoided even when the size of the standard cell is reduced. Furthermore, the openings in the source / drain regions originally used to electrically couple the source / drain regions to the metal-2 layer (M2) or the metal-3 layer (M3) for Vdd or ground connections can be omitted in the new standard cell and the standard cell.
[0060] Based on the above, FIG. 9(a) shows the SRAM bit cell size (λ 2 conversion) that can be seen across three different companies and different technology nodes from the present invention. FIG. 9(b) is a diagram showing the comparison results of the area size of the new standard cell provided by the present invention and those of conventional products provided by various other companies. As shown in FIG. 9(a), the area of the newly proposed SRAM cell (the present invention) is about 100λ 2 , which can be approximately 1 / 8 of the area of the conventional 5nm DRAM cell (of three different companies) shown in FIG. 3. Further, as shown in FIG. 9(b), the area of the newly proposed standard cell (for example, the inverter cell may be smaller than that of 200λ 2 ) is about 1 / 3.5 of the area of the conventional 5nm standard DRAM cell shown in FIG. 5.
[0061] Therefore, the innovation in the monolithic die design of the integrated scaling and / or stretching platform (ISSP) is proposed together with any combination of the proposed technologies (for example, new transistors, interconnections with transistors, SRAM cells, and standard cell designs) so that the circuit in the original layout of the die can be reduced by 2 to 3 times or more in terms of its area, to provide an integrated system.
[0062] From another perspective, more SRAM, or more major different functional blocks (CPU or GPU), can be formed in the original size of a single monolithic die. Thus, the device density and computing performance of an integrated system (such as an AI chip or SOC) can be significantly increased compared to the conventional one of the same size without shrinking the technology node for manufacturing the integrated system.
[0063] For example, if a 5nm technology process node is used, the CMOS 6-T SRAM cell size can be reduced to approximately 100F 2 (where F is the minimum feature size fabricated on the silicon wafer). That is, when F = 5nm, based on publications, it can occupy approximately 2500nm 2 compared to the latest cell area of about 800F 2 . Further, the 8-finger CMOS inverter (shown in FIGS. 4(a) and 4(b)) consumes a die area of 200F 2 based on the present invention, as opposed to exceeding 700F 2 of the published conventional CMOS inverter (5nm process node in FIG. 9(b)).
[0064] That is, if a single monolithic die has a circuit (such as an SRAM circuit, a logic circuit, a combination of SRAM and logic circuits, or a major functional block circuit CPU, GPU, FPGA, etc.) occupying a die area based on a technology process node (e.g., Ynm 2 ), with the help of the present invention, the total area of the monolithic die having the circuit on the same drawing can be shrunk even if the monolithic die is still manufactured by the same technology process node. The new die area occupied by the circuit on the same drawing in the monolithic die becomes smaller than the original die area, such that it is 20% - 80% (or 30% - 70%) of Ynm 2 .
[0065] For example, FIG. 10 shows an integrated system 1000 based on the Integrated Scaling and Stretching Platform (ISSP) of the present invention in comparison with a conventional one. As shown in FIG. 10, the ISSP integrated system 1000 and the conventional system 1010 include at least one processing device / circuit or main functional block (e.g., logic circuit 1011A and SRAM circuit 1011B), and at least one single monolithic die 1011 having a pad region 1011C. The integrated system 1000 provided by the ISSP of the present invention also includes at least one single monolithic die 1001 having a logic circuit 1001A, an SRAM circuit 1001B , and a pad region 1001C. By comparing the configurations of the monolithic dies 1011 and 1001 between the conventional system 1010 and the ISSP integrated system 1000, it can be shown that the ISSP of the present invention can reduce the size of the integrated system (monolithic die 1001) without degrading the conventional performance, or add more devices within the same scanner maximum drawing area (monolithic die 1001').
[0066] From the perspective of reducing the size of the ISSP integrated system 1000, as shown in the center of FIG. 10, the single monolithic die 1001 of the ISSP integrated system 1000 has the same circuits or main functional blocks as the conventional monolithic die 1011 (i.e., the logic circuit 1001A and the SRAM circuit of the single monolithic die 1001 1001B are identical to the logic circuit 1011A and the SRAM circuit 1011B of the single monolithic die 1011), and the single monolithic die 1001 only occupies 20% - 80% (or 30% - 70%) of the scanner maximum drawing area of the conventional monolithic die 1011.
[0067] In one embodiment, the combined area of the SRAM circuit 1001B and the logic circuit 1001A within the single monolithic die 1001 shrinks the area by 3.4 times that of the area of the conventional monolithic die 1011. That is, compared with the conventional monolithic die 1011, the ISSP of the present invention results in the area of the logic circuit 1001A of the single monolithic die 1001 being reduced by 5.3 times, the area of the SRAM circuit 1001B of the single monolithic die 1001 being reduced by 5.3 times, and (as shown in the center of FIG. 10,) the combined area of the SRAM circuit 1001B and the logic circuit 1001A within the single monolithic die 1001 can be reduced by 3.4 times.
[0068] From another perspective of adding more devices, as shown on the right side of FIG. 10, the single monolithic die 1001’ and the conventional monolithic die 1001 have the same scanner maximum drawing area. That is, the single monolithic die 1001’ is fabricated based on the same technology node as that of the conventional monolithic die 1011 (such as 5 nm or 7 nm), and the area of the SRAM circuit 1001B’ within the single monolithic die 1001’ can not only include more SRAM cells, but also include additional major functional blocks not present within the conventional monolithic die 1011. In another embodiment of the present disclosure, the die area of the single monolithic die 1001’ (as shown on the right side of FIG. 10) can be the same as, or substantially the same as, the scanner maximum drawing area (SMFA) of the conventional single monolithic die 1011 defined by a specific technology process node. That is, based on the ISSP of the present invention, within the scanner maximum drawing area (SMFA), there is additional space for accommodating additional SRAM cells, or additional major functional blocks other than those included in the conventional monolithic die 1011 (logic circuit 1011A and SRAM circuit 1011B )).
[0069] FIG. 11(a) is a diagram showing another ISSP integration system 1100 of the present disclosure. The ISSP integration system 1100 includes at least one monolithic die 1101 having the size of SMFA. The monolithic die 1101 includes a processing device / circuit (e.g., XPU1101A), an SRAM cache (including high-level and low-level caches), and an I / O circuit 1101B. Each of the SRAM caches includes a set of SRAM arrays. The I / O circuit 1101B is electrically connected to a plurality of SRAM caches and / or the XPU1101A.
[0070] In this embodiment, the monolithic die 1101 of the ISSP integration system 1100 includes different levels of caches L1, L2, and L3 generally composed of SRAM. Caches L1 and L2 (collectively referred to as "low-level caches") are usually assigned one for each CPU or GPU core device. Cache L1 is divided into L1i and L1d respectively used for storing instructions and data. Cache L2 does not distinguish between instructions and data. Cache L3 (which can be one of the "high-level caches") is shared by multiple cores and usually does not distinguish between instructions and data. Caches L1 / L2 are usually one for each CPU or GPU core.
[0071] In the case of high-speed operation, therefore, according to the ISSP of the present disclosure, the die area of the monolithic die 1101 can be the same as or approximately the same as the scanner maximum draw screen area (SMFA) defined by a specific technology process node. However, the storage capacities of the caches L1 / L2 (low-level caches) and cache L3 (high-level cache) of the ISSP integration system 1100 can be increased. As shown in FIG. 11(a), a GPU having multiple cores can have an SMFA (e.g., 26 mm × 33 mm, i.e., 858 mm) in which the high-level cache can have SRAM of 64 MB or more (e.g., 128 MB, 256 or 512 MB). 2) may be provided. Further, additional logic cores of the GPU can be inserted into the same SMFA to improve performance. The same applies to a memory controller (not shown) within the wideband I / O 1101B for another embodiment.
[0072] Alternatively, in addition to the existing major functional blocks, different major functional blocks such as an FPGA can be integrated together within the same monolithic die. FIG. 11(b) is a diagram showing a single monolithic die 1101' of the ISSP integrated system 1100' according to another embodiment of the present disclosure. In this embodiment, the monolithic die 1101' includes at least one wideband I / O 1101B', as well as a plurality of processing devices / circuits such as an XPU 1101A' and a YPU 1101C. The processing devices (XPU 1101A' and YPU 1101C) have major functional blocks, and each of them can serve as an NPU, a GPU, a CPU, an FPGA, or a TPU (tensor processing unit). The major functional blocks of the XPU 1101A' may be different from those of the YPU 1101C.
[0073] For example, the XPU 1101A' of the ISSP integrated system 1100' may serve as a CPU, and the YPU 1101C of the ISSP integrated system 1100' may serve as a GPU. The XPU 1101A' and the YPU 1101C each have a plurality of logic cores, and each core has a low-level cache (for example, a cache L1 / L2 having 512K or 1M / 128 bits), and a high-capacity high-level cache (for example, a cache L3 having 32MB, 64MB or more) shared by the XPU 1101A' and the YPU 1101C. These three levels of caches may each include a plurality of SRAM arrays.
[0074] Due to the fact that GPUs are becoming increasingly important for AI learning, FPGAs have blocks of logic that interact with each other and may be designed by engineers to support specific algorithms and are suitable for AI inference. Thus, in some embodiments of the present disclosure, an ISSP integrated system 1100” having a single monolithic die 1101” may include a GPU and an FPGA, as shown in FIG. 11(c). The configuration of the monolithic die 1101” in FIG. 11(c) is the same as that of the monolithic die 1101’ in FIG. 11(b), except that the XPU 1101A” of the monolithic die 1101” is a GPU or a CPU and the YPU 1101C’ of the monolithic die 1101” is an FPGA. By this approach, the monolithic die 1101” on the one hand has excellent parallel computing, learning speed, and efficiency, and on the other hand, it also has excellent AI inference capabilities with shorter time to market, lower cost, and flexibility.
[0075] Furthermore, as shown in FIG. 11(c), the processing devices / circuits (i.e., XPU 1101A” and YPU 1101C’) share a high-level cache (e.g., cache L3). The high-level cache (e.g., cache L3) shared between 1101A” and YPU 1101C’ can be configured by setting it in another mode register (not shown) or can be adaptively configured during the operation of the monolithic die 1101”. For example, in one embodiment, by setting the mode register, 1 / 3 of the high-level cache may be used by XPU 1101A” and 2 / 3 of the high-level cache may be used by YPU 1101C’. Such allocation capacities of the high-level cache (e.g., cache L3) for XPU 1101A” or YPU 1101C’ can also be dynamically changed based on the operation of the integrated scaling and / or stretching platform (ISSP) for forming the integrated system 1100”.
[0076] FIG. 11(d) is a diagram showing a single monolithic die 1101''' of an ISSP integration system 1100''' according to yet another embodiment of the present disclosure. The arrangement of the monolithic die 1101''' in FIG. 11(d) is the same as that of the monolithic die 1101' in FIG. 11(b), except that the high-level cache includes cache L3 and cache L4, each processing device / circuit (such as XPU1101A''' and YPU1101C'') has cache L3 shared by its own core, and cache L4 having 32 MB or more is shared by the XPU and the YPU.
[0077] In some embodiments of the present disclosure, a somewhat large-capacity shared SRAM (or embedded SRAM, "eSRAM") can be designed into one monolithic (single) die due to the smaller area of the SRAM cell design according to the present invention. Since the high memory capacity of eSRAM can be used, it is faster and more effective compared to conventional embedded DRAM or external DRAM. Therefore, it is reasonable and possible to have high-bandwidth / high-memory-capacity SRAM in a single monolithic die having the same or substantially the same (e.g., 80% - 99%) die size as the scanner maximum draw screen area (SMFA, e.g., 26 mm × 33 mm, i.e., 858 mm 2 ).
[0078] Accordingly, the integrated system 1200 provided by the integrated scaling and / or stretching platform (ISSP) of the present disclosure includes at least two single monolithic dies, and the two monolithic dies may have the same or substantially the same size. For example, FIG. 12(a) is a diagram showing another ISSP integrated system 1200 according to another embodiment of the present disclosure in comparison with a conventional one 1210. The ISSP integrated system 1200 includes a single monolithic die 1201 and a single monolithic die 1202 within a single package. The single monolithic die 1201 mainly has a logic processing device circuit and a low-level cache formed therein, and the second monolithic die 1202 has only a plurality of SRAM arrays and I / O circuits formed therein. The plurality of SRAM arrays include at least 2 to 20 gigabytes, for example, 2 to 10 gigabytes.
[0079] As shown in FIG. 12(a), the single monolithic die 1201 mainly includes a logic circuit and an I / O circuit 1201A, and a small low-level cache (such as L1 and L2 caches) composed of an SRAM array 1201B. The single monolithic die 1202 includes only a high-bandwidth SRAM circuit 1202B having 2 to 10 gigabytes or more (for example, 1 to 20 gigabytes), and an I / O circuit 1202A for the high-bandwidth SRAM circuit 1202B. In this embodiment, the SMFA of the single monolithic die 1201 and the single monolithic die 1202 can be about 26 mm × 33 mm. If 50% of the SMFA of the single monolithic die 1202 (50% SRAM cell utilization rate) is used for the SRAM cells of the high-bandwidth SRAM circuit 1202B, the rest of the SMFA is used for the I / O circuit of the high-bandwidth SRAM circuit 1202B.
[0080] FIG. 12(b) is a diagram showing a comparison result of the SRAM cell area of the integrated system 1200 of the present invention and those of three foundries based on different technology nodes. The total number of bytes (1 bit per SRAM cell) in the 26 mm × 33 mm SMFA of a single monolithic die (e.g., a single monolithic die 1202) can be estimated by referring to the SRAM cell area as shown in FIG. 12(b). For example, in the present embodiment, the SMFA (26 mm × 33 mm) of the single monolithic die 1200 can accommodate 21 GB SRAM at a 5 nm technology node and can provide more than 24 GB if the SRAM cell utilization rate can be increased.
[0081] According to FIG. 12(b), the conventional SRAM cell area (of the three foundries) may be 2 to 8 times that of the SRAM cell area of the present invention. Therefore, the ISSP integrated system 1200 can accommodate a larger number of bytes (1 bit per SRAM cell) in the 26 mm × 33 mm SMFA than those of the prior art. The total number of bytes (1 bit per SRAM cell) in the 26 mm × 33 mm SMFA based on different technology nodes is shown in Table 1 below.
[0082] [Table 1]
[0083] Of course, considering the various technologies proposed herein and the selective use of conventional wiring process technologies, the SMFA (26 mm × 33 mm) of the single monolithic die 1202 can accommodate a smaller capacity of SRAM, e.g., 1 / 4 to 3 / 4 times the SRAM sizes at various technology nodes in Table 1 above. For example, the single monolithic die 1202 can accommodate approximately 5 to 15 GB SRAM, or 2.5 GB to 7.5 GB, by the selective use of the various technologies proposed herein and conventional wiring process technologies.
[0084] FIG. 13(a) is a diagram showing a single monolithic die 1301 of another ISSP integrated system 1300 according to the present invention. The arrangement of the single monolithic die 1301 is such that the single monolithic die 1301 of the present embodiment includes a broadband I / O circuit 1301A, an XPU 1301B having a plurality of cores, and a YPU 1301C and may be a high-performance computing (HPC) monolithic die including two or more main functional blocks such as, and each core of the XPU 1301B and the YPU 1301C has its own cache L1 and / or cache L2 (L1~128KB, and L2~512KB to 1MB), which is the same as that of the single monolithic die 1201 in FIG. 12(a) except for this. The main functional blocks of the XPU 1301B or the YPU 1301C in FIG. 13(a) can be an NPU, a GPU, a CPU, an FPGA, or a TPU (tensor processing unit), and each of them has a main functional block. The XPU 1301B or the YPU 1301C may have different main functional blocks.
[0085] FIG. 13(b) is a diagram showing a single monolithic die 1302 of the ISSP integrated system 1300. The arrangement of the single monolithic die 1302 is the same as that of the single monolithic die 1202 in FIG. 12(a) except that the single monolithic die 1302 is a high-bandwidth SRAM (HBSRAM). In the present embodiment, the single monolithic die 1302 has the same SMFA as the latest SMFA (or has an area of about 80~90%), and includes only a cache L3 and / or L4 having a plurality of SRAM arrays, and a SRAM I / O circuit 1302A having a broadband 1302B I / O for the SRAM I / O circuit 1302A. The total SRAM in the single monolithic die 1302 may be 2~5GB, 5~10GB, 10~15GB, 15~20GB, or more depending on the utilization rate of the SRAM cells. Such a single monolithic die 1302 can be a high-bandwidth DRAM (HBSRAM).
[0086] As shown in FIGS. 13(a) and 13(b), a single monolithic die 1301 and a single monolithic die 1302 each have a wideband I / O bus, e.g., a 64-bit, 128-bit, or 256-bit data bus. The single monolithic die 1301 and the single monolithic die 1302 may be within the same IC package or in separate IC packages. For example, in some embodiments, a single monolithic die 1301 (e.g., an HPC die) is coupled to a single monolithic die 1302 (e.g., by wire bonding, flip chip bonding, solder bonding, 2.5D interposer silicon through-via (TSV) bonding, 3D micro copper pillar direct die bonding) as shown in FIG. 14 and housed within a single package to form an integrated system 1400. In an embodiment, both the single monolithic die 1301 and the single monolithic die 1302 have the same or substantially the same SMFA, and thus such a coupling may be accomplished by directly coupling a wafer 14A having at least the single monolithic die 1301 (or having multiple dies) to another wafer 14B having at least the single monolithic die 1302 (or having multiple dies), and then the coupled wafers 14A and 14B are sliced into a plurality of SMFA blocks to form the integrated system 1400 provided by the ISSP of the present disclosure. It is contemplated that another interposer having TSVs may be inserted between the single monolithic die 1301 and the single monolithic die 1302.
[0087] FIG. 15 is a diagram showing another ISSP integration system 1500 according to the present disclosure. The integration system 1500 includes two or more single monolithic dies 1302 (i.e., two HBSRAM dies as shown in FIG. 13(b)) coupled to each other, and one of the two single monolithic dies 1302 is then coupled to a single monolithic die 1301 (e.g., an HPC die as shown in FIG. 13(a)), and then all three or more dies are housed within a single package. Thus, such a package may include an HPC die and HBSRAM of more than 42, 48, or 96 GB. Of course, two or more single monolithic dies 1302 and a single monolithic die 1301 having a broadband I / O bus may be stacked vertically and coupled to each other based on the latest bonding technology.
[0088] Of course, three, four, or five or more HBSRAM dies may be integrated within a single package of the integration system 1500, and in that case, the cache L3 and L4 within the integration system 1500 may be larger than 128 GB or 256 GB SRAM. In some embodiments of the present disclosure, the single monolithic dies 1301 and 1302 of the integration system 1500 may be housed within the same IC package.
[0089] Compared with currently available HBM DRAM memory including approximately 24 GB based on a stack of 12 DRAM chips, the present invention can replace HBM3 memory with more HBSRAM (e.g., one HBSRAM chip having approximately 5 - 10 GB or 15 - 20 GB). Thus, within the ISSP, HMB memory is not required or only a small amount of HBM memory (e.g., less than 4 GB or 8 GB of HBM) is required.
[0090] Monolithic integration on a single die that enables the achievement of Moore's Law is currently facing its limits, especially due to the limitations of photolithography technology. On the one hand, the minimum feature size printed on the die is very costly to shrink in that dimension, while on the other hand, the die size is limited by the maximum scanner field area. However, more and more diverse functions of processors are emerging, and it is difficult to integrate them on a monolithic die. Furthermore, the small eSRAM on each major functional die and the somewhat redundant presence of external or embedded DRAM are not desirable and optimized solutions. Based on the Integrated Scaling and / or Stretching Platform (ISSP) within a monolithic die or SOC die, (1) A single major functional block such as an FPGA, TPU, NPU, CPU, or GPU can be shrunk to a much smaller size, (2) More SRAM can be formed within the monolithic die, and, (3) Two or more major functional blocks, also shrunk through this ISSP, such as GPUs and FPGAs (or other combinations thereof), can be integrated together within the same monolithic die. (4) More levels of cache can exist within the monolithic die. (5) Such ISSP monolithic dies can be combined with another die (e.g., eDRAM) based on heterogeneous integration. (6) An HPC die 1 with L1&L2 caches is electrically connected (e.g., wire bonding or flip chip bonding) to one or more HBSRAM dies 2 utilized as L3&L4 caches within a single package, and each of the HPC die 1 and the HBSRAM die 2 has SMFA. (7) Within the ISSP, HMB memory is not required or only a small amount of HBM memory is required.
[0091] The present invention has been described by way of example and by way of preferred embodiments, but it is to be understood that the invention is not limited thereto. In contrast, various modifications, as well as similar arrangements and procedures, are intended to be included, and accordingly, the broadest interpretation should be given to the appended claims so as to include all such modifications, as well as all similar arrangements and procedures.
Claims
Claim 1 A first monolithic die in which a processing device circuit is formed therein, wherein the processing device circuit includes a first processing device circuit and a second processing device circuit, the first processing device circuit includes a plurality of first logic cores, each of the plurality of first logic cores includes a first SRAM, the second processing device circuit includes a plurality of second logic cores, each of the plurality of second logic cores includes a second SRAM, and a main function performed by the first processing device circuit is different from a main function performed by the second processing device circuit, the first monolithic die, and A second monolithic die in which a plurality of SRAM arrays are formed therein comprising the second monolithic die comprises at least 2 Gbytes the plurality of SRAM arrays or the processing device circuit includes PMOS transistors and NMOS transistors formed on a semiconductor substrate, the PMOS transistor includes a first source region, a first drain region, and a first insulating structure used to insulate a bottom of the first source region or the first drain region from the semiconductor substrate, and the NMOS transistor includes a second source region, a second drain region, and a second insulating structure used to insulate a bottom of the second source region or the second drain region from the semiconductor substrate The first monolithic die is electrically connected to the second monolithic die, the die area of the first monolithic die is the same as or substantially the same as the die area of the second monolithic die, and the die area of the first monolithic die does not exceed 858 mm 2 integrated system. Claim 2 the die area of the first monolithic die is the same as or substantially the same as a maximum scanner screen area defined by a specific technology process node The integrated system according to claim 1 Claim 3 The integrated system according to claim 2, wherein the first monolithic die and the second monolithic die are housed in a single package Claim 4 The integrated system according to claim 1, wherein the plurality of SRAM arrays comprise at least 20 Gbytes Claim 5 The integrated system according to claim 1, wherein the first SRAM has an L1 cache and an L2 cache, and the second SRAM has another L1 cache and another L2 cache Claim 6 The first processing device circuit or the second processing device circuit is the integrated system according to claim 5, which is selected from the group consisting of a graphics processing unit (GPU), a central processing unit (CPU), a tensor processing unit (TPU), a network processing unit (NPU), and a field programmable gate array (FPGA).
7. The first monolithic die further includes an L1 cache and an L2 cache used by the processing device circuit during operation of the first monolithic die. The integrated system according to claim 1, wherein the plurality of SRAM arrays include an L3 cache and an L4 cache used by the processing device circuit during operation of the first monolithic die.
8. The integrated system further includes a third monolithic die in which a plurality of SRAM arrays are formed. The third monolithic die has at least 2 to 20 gigabytes. The first monolithic die, the second monolithic die, and the third monolithic die are housed in a single package. The first monolithic die has a die area that is the same as or substantially the same as the maximum scanner image area defined by a specific technology process node. The second monolithic die has a die area that is the same as or substantially the same as the maximum scanner image area defined by the specific technology process node. The integrated system according to claim 1, wherein the third monolithic die has a die area that is the same as or substantially the same as the maximum scanner image area defined by the specific technology process node.
9. The integrated system according to claim 8, wherein the first monolithic die, the second monolithic die, and the third monolithic die are stacked vertically.
10. An integrated system A first monolithic die having a processing device circuit, wherein the processing device circuit includes a first processing device circuit and a second processing device circuit, the first processing device circuit includes a plurality of first logic cores, each of the plurality of first logic cores includes a first SRAM, the second processing device circuit includes a plurality of second logic cores, each of the plurality of second logic cores includes a second SRAM, and a main function performed by the first processing device circuit is different from a main function performed by the second processing device circuit, the first monolithic die, and A second monolithic die having a plurality of SRAM arrays each having at least 2 Gbytes comprising The plurality of SRAM arrays or the processing device circuit includes PMOS transistors and NMOS transistors formed on a semiconductor substrate, the PMOS transistor includes a first source region, a first drain region, and a first insulating structure used to insulate a bottom of the first source region or the first drain region from the semiconductor substrate, and the NMOS transistor includes a second source region, a second drain region, and a second insulating structure used to insulate a bottom of the second source region or the second drain region from the semiconductor substrate, The first monolithic die is physically separated from the second monolithic die, The first monolithic die is electrically connected to the second monolithic die, The integrated system does not include a high-bandwidth memory (HBM), the die area of the first monolithic die is the same as or substantially the same as the die area of the second monolithic die, and the die area of the first monolithic die does not exceed 858 mm 2 exceeding it.
11. The integrated system according to claim 10, wherein the first monolithic die has an area size that is the same as or substantially the same as a maximum scanner screen area defined by a specific technology process node.
12. The first monolithic die and the second monolithic die are housed in a single package, The integrated system according to claim 10, wherein the first monolithic die is electrically connected to the second monolithic die by wire bonding, flip chip bonding, solder bonding, interposer through-silicon via (TSV) bonding, or micro copper pillar direct bonding.
13. A first monolithic circuit having a processing device circuit, wherein the processing device circuit includes a first processing device circuit and a second processing device circuit, the first processing device circuit includes a plurality of first logic cores, each of the plurality of first logic cores includes a first SRAM, the second processing device circuit includes a plurality of second logic cores, each of the plurality of second logic cores includes a second SRAM, and a main function performed by the first processing device circuit is different from a main function performed by the second processing device circuit, the first monolithic circuit, and A second monolithic circuit having a plurality of SRAM arrays each having at least 2 to 20 gigabytes are provided, The plurality of SRAM arrays or the processing device circuit includes PMOS transistors and NMOS transistors formed on a semiconductor substrate, the PMOS transistor includes a first source region, a first drain region, and a first insulating structure used to insulate a bottom of the first source region or the first drain region from the semiconductor substrate, and the NMOS transistor includes a second source region, a second drain region, and a second insulating structure used to insulate a bottom of the second source region or the second drain region from the semiconductor substrate, The first monolithic circuit is electrically connected to the second monolithic circuit, The first monolithic circuit is formed in a first monolithic die, and the second monolithic circuit is formed in a second monolithic die, The first monolithic die and the second monolithic die are housed within a single package, or the first monolithic die and the second monolithic die are respectively housed within a first package and a second package, and the die area of the first monolithic die is the same as or substantially the same as the die area of the second monolithic die, and the die area of the first monolithic die does not exceed 858 mm 2 exceeding An integrated system.
Citation Information
Patent Citations
Manufacturing method of semiconductor device
JP2008311635A
Semiconductor device
JP2010080801A
Integrated circuit chip using top post-passivation technology and bottom structure technology
JP2015073107A
Semiconductor device and electronic equipment
JP2019047006A
Integrated scaling and stretching platform for optimization of monolithic and / or heterogeneous integration on single semiconductor die
JP2022140387A