Semiconductor device, manufacturing method thereof, storage device and system

By integrating memory devices and computing circuits and combining them through bonding technology, the problems of slow computing speed and high power consumption of memory devices and CPUs or GPUs are solved, and efficient AI computing is achieved.

CN120825928APending Publication Date: 2025-10-21YANGTZE MEMORY TECHNOLOGIES HOLDING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410437291.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

In the existing technology, the computing speed of memory devices and CPUs or GPUs is slow and the power consumption is high, which limits the computing efficiency of large AI models. The mismatch between storage and computing performance leads to high memory access latency, which makes it difficult to meet the needs of large AI models.

Method used

The storage array of the memory device and the interface unit of the computing circuit are integrated into the same semiconductor structure, and the execution unit of the computing circuit and the peripheral circuit of the memory device are integrated into another semiconductor structure. The two are combined through bonding technology to match the process nodes and reduce production costs and difficulty.

Benefits of technology

It improves computing speed, reduces data interaction latency and power consumption, enhances storage density, reduces production costs and complexity, and improves the efficiency of AI computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120825928A_ABST
    Figure CN120825928A_ABST
Patent Text Reader

Abstract

The invention provides a semiconductor device and a manufacturing method thereof, and a storage and calculation device and system, and relates to the technical field of semiconductors. The semiconductor device includes a first semiconductor structure and a second semiconductor structure. The first semiconductor structure includes an execution unit of a computing circuit and a peripheral circuit of a memory device. And a second semiconductor structure bonded to the first semiconductor structure, the second semiconductor structure including a memory array of the memory device and an interface cell of the computing circuit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of semiconductor technology, and more specifically, to a semiconductor device and a manufacturing method thereof, a storage and computing device, and a system. Background Art

[0002] In related technologies, data in a memory device needs to be transferred to a CPU (Central Processing Unit) or GPU (Graphics Processing Unit) for calculation. This method of using a memory device in conjunction with a CPU or GPU results in slow calculation speed and high power consumption, which limits calculation efficiency. With the advent of the era of large AI (Artificial Intelligence) models, large AI models require high computing power, but the mismatch between "storage" and "computing" performance leads to high memory access latency and low efficiency, making it difficult to meet demand. Summary of the Invention

[0003] The purpose of the present disclosure is to provide a semiconductor device and a manufacturing method thereof, a storage and computing device and a system.

[0004] An embodiment of the present disclosure provides a semiconductor device, comprising: a first semiconductor structure, the first semiconductor structure including an execution unit of a computing circuit and a peripheral circuit of a memory device; and a second semiconductor structure bonded to the first semiconductor structure, the second semiconductor structure including a storage array of the memory device and an interface unit of the computing circuit.

[0005] An embodiment of the present disclosure provides a method for manufacturing a semiconductor device, comprising: providing a first semiconductor structure, the first semiconductor structure comprising a first substrate and an execution unit of a computing circuit and a peripheral circuit of a memory device located on the first substrate; providing a second semiconductor structure, the second semiconductor structure comprising a storage array of the memory device and an interface unit of the computing circuit; and bonding the second semiconductor structure to the first semiconductor structure.

[0006] An embodiment of the present disclosure provides a storage and computing device, comprising a memory device and a computing circuit coupled to the memory device; the memory device comprises a memory array and a peripheral circuit coupled to the memory array; the computing circuit comprises an execution unit and an interface unit; the execution unit and the peripheral circuit are contained in a first semiconductor structure; the memory array and the interface unit of the computing circuit are contained in a second semiconductor structure; the first semiconductor structure is bonded to the second semiconductor structure.

[0007] An embodiment of the present disclosure provides a storage and computing system, comprising a memory device, a controller coupled to the memory device, and a computing circuit: the memory device comprises a storage array and a peripheral circuit coupled to the storage array; the computing circuit comprises an execution unit and an interface unit; the execution unit and the peripheral circuit are contained in a first semiconductor structure; the memory array and the interface unit of the computing circuit are contained in a second semiconductor structure; the first semiconductor structure is bonded to the second semiconductor structure; the controller is configured to control the memory device to perform storage operations and control the computing circuit to read data from the storage array to perform computing operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 A schematic diagram showing the structure of a semiconductor device in an embodiment of the present disclosure Figure 1 .

[0009] Figure 2 A schematic diagram showing the structure of a semiconductor device in an embodiment of the present disclosure Figure 2 .

[0010] Figure 3 Shown separately Figure 2 Top views of wafer A and wafer B in the semiconductor device before bonding are shown.

[0011] Figure 4 Show Figure 2 The semiconductor device is shown in a cross-sectional view taken along the dotted line CD.

[0012] Figure 5 1 and 2 show top views of wafer A and wafer B in another semiconductor device before bonding in an embodiment of the present disclosure.

[0013] Figure 6 A cross-sectional view of another semiconductor device taken along the dotted line CD is shown.

[0014] Figure 7 1 and 2 show top views of wafers A and B in another semiconductor device before bonding in an embodiment of the present disclosure.

[0015] Figure 8 A cross-sectional view of yet another semiconductor device taken along the dotted line CD is shown.

[0016] Figure 9 1 and 2 show top views of wafers A and B in another semiconductor device before bonding in an embodiment of the present disclosure.

[0017] Figure 10 A cross-sectional view of yet another semiconductor device taken along the dotted line CD is shown.

[0018] Figure 111 and 2 show top views of wafers A and B in another semiconductor device before bonding in an embodiment of the present disclosure.

[0019] Figure 12 A cross-sectional view of yet another semiconductor device taken along the dotted line CD is shown.

[0020] Figure 13 1 and 2 show top views of wafers A and B in another semiconductor device before bonding in an embodiment of the present disclosure.

[0021] Figure 14 A cross-sectional view of yet another semiconductor device taken along the dotted line CD is shown.

[0022] Figure 15 1 and 2 show top views of wafers A and B in another semiconductor device before bonding in an embodiment of the present disclosure.

[0023] Figure 16 A cross-sectional view of yet another semiconductor device taken along the dotted line CD is shown.

[0024] Figure 17 1 and 2 show top views of wafers A and B in another semiconductor device before bonding in an embodiment of the present disclosure.

[0025] Figure 18 A cross-sectional view of yet another semiconductor device taken along the dotted line CD is shown.

[0026] Figure 19 A schematic diagram showing the structure of a semiconductor device in an embodiment of the present disclosure Figure 3 .

[0027] Figure 20 Show Figure 19 Top views of wafer A and wafer B in the semiconductor device before bonding are shown.

[0028] Figure 21 Show Figure 19 The semiconductor device is shown in a cross-sectional view taken along the dotted line CD.

[0029] Figure 22 A schematic diagram showing the structure of another semiconductor device in the embodiment of the present disclosure is shown. Figure 4 .

[0030] Figure 23 Show Figure 22 The semiconductor device is shown in a cross-sectional view taken along the dotted line CD.

[0031] Figure 24 A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0032] Figure 25A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0033] Figure 26 A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0034] Figure 27 A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0035] Figure 28 A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0036] Figure 29 A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0037] Figure 30 A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0038] Figure 31 A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0039] Figure 32 A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0040] Figure 33 A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0041] Figure 34 A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0042] Figure 35 A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0043] Figure 36 A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0044] Figure 37 A cross-sectional view of another semiconductor device taken along the dotted line CD in an embodiment of the present disclosure is shown.

[0045] Figure 38 A schematic diagram of a computing circuit in an embodiment of the present disclosure is shown.

[0046] Figure 39 A schematic diagram showing another computing circuit in an embodiment of the present disclosure is shown.

[0047] Figure 40A block diagram of a system having a host and a storage and computing system in an embodiment of the present disclosure is shown.

[0048] Figure 41 A block diagram of a storage and computing device is shown as an example.

[0049] Figure 42 A block diagram of another storage and computing device is shown as an example.

[0050] Figure 43 A block diagram of a system is shown as an example.

[0051] Figure 44 A schematic diagram of a storage and computing device is shown as an example.

[0052] Figure 45 A schematic diagram of another storage and computing device is shown as an example.

[0053] Figure 46 A schematic diagram of another storage and computing device is shown as an example.

[0054] Figure 47 A schematic diagram of a storage array is shown as an example.

[0055] Figure 48 A schematic diagram of a storage and computing system is shown as an example.

[0056] Figure 49 A flow chart illustrating a method for manufacturing a semiconductor device in an embodiment of the present disclosure is shown.

[0057] Figures 50 to 54 The following is a schematic diagram showing various steps of a method for manufacturing a semiconductor device.

[0058] Figures 55 to 59 Schematic diagram illustrating various steps of another method for manufacturing a semiconductor device. DETAILED DESCRIPTION

[0059] The disclosure below provides many different embodiments or examples for implementing the different features of the provided subject matter. Of course, these are merely examples and are not intended to be limiting. In addition, spatially relative terms may be used herein for ease of description, for example, "below," "below," "below," "above," "above," "upper side," "first side," "second side," "lower side," etc., to describe the relationship of one element or feature to other elements or features as shown in the figures. Spatially relative terms are intended to encompass different orientations of the device in use or operation other than the orientation shown in the accompanying drawings. The device may have other orientations (rotated 90 degrees or in other orientations), and the spatially relative descriptors used herein shall be interpreted accordingly.

[0060] Example embodiments will now be described more fully with reference to the accompanying drawings. The accompanying drawings are schematic illustrations of the present disclosure only and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. In addition, the described features, structures or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure can be practiced while omitting one or more of the specific details, or that other methods, devices, steps, etc. can be adopted. In other cases, well-known structures, methods, devices, implementations or operations are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0061] Furthermore, the terms "first," "second," and the like are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of this disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined. The symbol " / " generally indicates an "or" relationship between the preceding and following objects.

[0062] In this disclosure, unless otherwise specified or limited, terms such as "coupled," "connected," and "coupled" should be understood broadly. For example, they may refer to electrical connection or mutual communication; they may refer to direct connection or indirect connection through an intermediary. Those skilled in the art will understand the specific meanings of these terms in this disclosure based on specific circumstances.

[0063] like Figure 1 As shown, the semiconductor device 100 provided in an embodiment of the present disclosure includes a first semiconductor structure 110 and a second semiconductor structure 120. The first semiconductor structure 110 includes an execution unit 111 of a computing circuit and a peripheral circuit 112 of a memory device. The second semiconductor structure 120 is bonded to the first semiconductor structure 110. The second semiconductor structure 120 includes a memory array 122 of the memory device and an interface unit 121 of the computing circuit.

[0064] exist Figure 1 In the embodiment, it is assumed that the stacking direction of the first semiconductor structure 110 and the second semiconductor structure 120 is a third direction (represented by the Z axis), and the first direction and the second direction perpendicular to the Z direction are represented by the X axis and the Y axis respectively. The first direction and the second direction can be perpendicular to each other, but the present disclosure is not limited to this.

[0065] The semiconductor device 100 of the embodiment of the present disclosure integrates both a memory device and a computing circuit. A memory device refers to any device with a data storage function, such as Flash (e.g., NAND Flash, NOR Flash, etc.), random access memory (RAM) (e.g., DRAM, SRAM, etc.), phase change memory (PCM) (e.g., storage class memory (SCM), etc.). The memory device further includes a memory array 122 and a peripheral circuit 112. The memory array 122 may include a plurality of memory cells. The peripheral circuit 112 may be configured to operate on the memory cells, for example, to write to or read from the memory cells.

[0066] A computing circuit refers to a circuit that performs computational processing on data, such as addition, multiplication, bitwise logical operations, and other related circuits. The computing circuit further includes an execution unit 111 and an interface unit 121 of the computing circuit. The interface unit 121 of the computing circuit can be configured to obtain data from a memory device, such as a storage array 122 or a peripheral circuit 112, and then send the obtained data to the execution unit 111. The execution unit 111 is configured to receive data from the interface unit 121 of the computing circuit, and perform computational processing on the data to obtain intermediate calculation results and / or final calculation results. Optionally, the interface unit 121 of the computing circuit can also be configured to send out the intermediate calculation results and / or final calculation results obtained by the execution unit 111. The interface unit 121 of the computing circuit can be equivalent to the IO (Input Output) component of the computing circuit.

[0067] Since the interface unit 121 of the computing circuit is more compatible with the manufacturing process node of the storage array 122 of the memory device, the execution unit 111 of the computing circuit is more compatible with the manufacturing process node of the peripheral circuit 112 of the memory device. Therefore, the embodiment of the present disclosure integrates the interface unit 121 of the computing circuit and the storage array 122 of the memory device into the same second semiconductor structure 120, and integrates the execution unit 111 of the computing circuit and the peripheral circuit 112 of the memory device into another first semiconductor structure 110. In this way, there is no need to switch process nodes during the manufacturing process of the same first semiconductor structure 110 or the same second semiconductor structure 120, that is, the execution unit 111 of the computing circuit and the peripheral circuit 112 of the memory device can be manufactured using the same process node, and the storage array 122 of the memory device and the interface unit 121 of the computing circuit can be manufactured using another identical process node. In the semiconductor field, the cost of switching process nodes is high and the difficulty is relatively high. Therefore, the method provided by the embodiment of the present disclosure can reduce the manufacturing cost of semiconductor devices.

[0068] In the embodiments of the present disclosure, on the one hand, integrating the memory device and the computing circuit into the same semiconductor device allows at least part of the computation to be performed within the semiconductor device. The data in the memory device can be directly processed by the computing circuit to perform at least part of the computation, reducing the latency, load, and power consumption of data interaction. Data no longer needs to be transmitted or completely transmitted to the CPU or GPU used in AI computing for processing, thereby increasing computing speed and reducing the power consumption and load of the GPU or CPU used in AI computing. On the other hand, by making the memory array of the memory device and the interface unit of the computing circuit into the same second semiconductor structure, and making the execution unit of the computing circuit and the peripheral circuit of the memory device into the same first semiconductor structure, the process requirements of the memory array and the interface unit of the computing circuit, as well as the execution unit and the peripheral circuit, can be matched, respectively, reducing the difficulty and cost of manufacturing. At the same time, the first semiconductor structure and the second semiconductor structure are manufactured separately and then bonded together by bonding, that is, the memory array and the peripheral circuit are manufactured separately, which can not only increase the storage density of the memory device, save area, but also improve the manufacturing efficiency of the semiconductor device.

[0069] In an exemplary embodiment, the storage array 122 comprises a DRAM (Dynamic Random Access Memory) storage array. When the storage array 122 comprises a DRAM storage array, the storage cells included in the storage array 122 comprise DRAM storage cells, the memory device comprises a DRAM memory device, and the peripheral circuit 112 comprises a DRAM peripheral circuit. The DRAM memory device can function as a memory. However, the present disclosure is not limited thereto. For example, the storage array 122 can comprise a NAND storage array, the storage cells included in the storage array 122 comprise NAND storage cells, and the memory device comprises a NAND flash memory.

[0070] In an exemplary embodiment, the first semiconductor structure 110 further includes a first bonding layer and a first substrate. An execution unit 111 and peripheral circuitry 112 are located between the first bonding layer and the first substrate. The second semiconductor structure 120 further includes a second bonding layer. A memory array 122 and an interface unit 121 for a computing circuit are located on a side of the second bonding layer away from the first semiconductor structure 110. The first semiconductor structure 110 and the second semiconductor structure 120 are bonded together via the first bonding layer and the second bonding layer.

[0071] In some embodiments, the second semiconductor structure 120 may further include a second substrate, and the memory array 122 and the interface unit 121 of the computing circuit may be located between the second bonding layer and the second substrate. In other embodiments, the second semiconductor structure 120 may further include a semiconductor layer, and the memory array 122 and the interface unit 121 of the computing circuit may be located between the second bonding layer and the semiconductor layer. For example, the memory array 122 and the interface unit 121 of the computing circuit may be first fabricated on the second substrate, and then a second bonding layer may be formed on the memory array 122 and the interface unit 121 of the computing circuit, and then the first semiconductor structure 110 and the second semiconductor structure 120 may be bonded through the first bonding layer and the second bonding layer. After bonding, the second substrate may be thinned or removed to form a semiconductor layer. A pad may be formed on the semiconductor layer to realize connection or electrical connection between the semiconductor device 100 or the second semiconductor structure 120 and other external circuits.

[0072] In some embodiments, the first substrate and the second substrate may be semiconductor substrates, for example, Si (silicon) substrates. However, the present disclosure is not limited thereto. For example, the first substrate and the second substrate may also include other semiconductors, such as one or more of germanium (Ge), silicon carbide (SiC), silicon germanium (SiGe), or diamond.

[0073] In some embodiments, the first semiconductor structure 110 and the second semiconductor structure 120 can be bonded in a front-to-front manner. In the embodiments of the present disclosure, when the substrate and the bonding layer are respectively arranged on opposite sides of the semiconductor structure (the device layer is sandwiched between the substrate and the bonding layer), the "front side" refers to the other side of the semiconductor structure opposite to the side where the substrate is located, and the "back side" refers to the side of the semiconductor structure where the substrate is located; when the substrate and the bonding layer are arranged on the same side of the semiconductor structure (the substrate is sandwiched between the device layer and the bonding layer), the "front side" refers to the other side of the semiconductor structure opposite to the side where the substrate and the bonding layer are located, and the "back side" refers to the side of the semiconductor structure where the substrate and the bonding layer are located.

[0074] In the following embodiments, it is assumed that the first semiconductor structure is wafer B and the second semiconductor structure is wafer A, that is, the first semiconductor structure and the second semiconductor structure are two independent wafers. Figure 2 and Figure 19 In the embodiment, the surface of the wafer B where the first substrate 113 is located is the back side of the wafer B, that is, the lower side of the wafer B is the back side of the wafer B, and the other surface opposite to the first substrate 113, that is, the surface where the first bonding layer 114 is located is the front side of the wafer B. For another example, Figure 2In the embodiment, the surface of wafer A where the second substrate 123 is located is the back side of wafer A, that is, the upper side of wafer A is the back side of wafer A, and the other surface opposite to the second substrate 123, that is, the surface where the second bonding layer 124 is located is the front side of wafer A. Figure 19 In the embodiment, the surface of wafer A where the second substrate 123 and the second bonding layer 124 are located is the back side of wafer A, that is, the lower side of wafer A is the back side of wafer A, and the other surface opposite to the second substrate 123 and the second bonding layer 124, that is, the upper side of wafer A is the front side of wafer A.

[0075] Continue to refer Figure 2 , wafer B includes a first substrate 113, a peripheral circuit 112 and an execution unit 111 of a computing circuit located on the first substrate 113, and a first bonding layer 114 located on the peripheral circuit 112 and the execution unit 111. Wafer A includes a second substrate 123, a storage array 122 and an interface unit 121 of a computing circuit located under the second substrate 123, and a second bonding layer 124 located under the storage array 122 and the interface unit 121 of the computing circuit. The front side of wafer B is bonded to the front side of wafer A. Here, it is assumed that a hybrid bonding method is adopted. In other embodiments, other bonding methods can also be adopted. In the following embodiments, hybrid bonding is used as an example, but the present disclosure is not limited to this.

[0076] It should be noted that, in some embodiments, the second substrate 123 on the wafer A can be removed. For example, a semiconductor layer (not shown) can be formed after the second substrate 123 is removed, or back-side lead-out can be performed after the second substrate 123 is removed, such as lead-out of the bit lines of the memory array 122.

[0077] In an exemplary embodiment, the interface unit 121 of the computing circuit includes a plurality of interface sub-units, and the memory array 122 includes a plurality of physical memory banks, wherein each interface sub-unit of the computing circuit is located between two adjacent physical memory banks.

[0078] In some embodiments, the interface unit 121 of the computing circuit includes a sense amplifier, which enables the interface unit 121 of the computing circuit to read data from the memory array 122. The interface unit 121 of the computing circuit may include at least one odd interface sub-unit and at least one even interface sub-unit. The memory array 122 includes multiple banks, which may be divided into odd physical memory banks 1221 and even physical memory banks 1222. The odd interface sub-unit may be used to connect to a corresponding odd physical memory bank 1221 to read data from the odd physical memory bank 1221; the even interface sub-unit may be used to connect to a corresponding even physical memory bank 1222 to read data from the even physical memory bank 1222.

[0079] In other embodiments, the interface unit 121 of the computing circuit may include multiple interface sub-units 1211, each interface sub-unit 1211 is located between two adjacent banks, each interface sub-unit can be connected to two adjacent banks, and can read data from the two banks connected to it.

[0080] In some embodiments, the interface unit 121 of the computing circuit may also send the intermediate result or the final result after the computing process to the peripheral circuit 112 , and the peripheral circuit 112 may store the intermediate result or the final result in the storage array 122 .

[0081] In the embodiment of the present disclosure, by locating each interface sub-unit 1211 between two adjacent banks, on the one hand, the distance between each interface sub-unit 1211 and the bank to which it is connected can be made equal, so that the distance when each interface sub-unit 1211 operates data from the odd physical storage body 1221 and the even physical storage body 1222, such as when reading data, is basically equal (there may be a certain deviation), thereby making the delay when operating data basically consistent.

[0082] In an exemplary embodiment, the interface unit 121 of the computing circuit is located on the same layer as the memory array 122 .

[0083] exist Figure 2 In the illustrated embodiment, the interface unit 121 of the computing circuit and the storage array 122 are located in the same layer, and each interface sub-unit 1211 is located between two adjacent banks. This allows each interface sub-unit 1211 to operate data at substantially the same distance from the odd-numbered physical storage bank 1221 and the even-numbered physical storage bank 1222, and also allows each interface sub-unit 1211 to operate data at as short a distance as possible from the bank to which it is connected, thereby reducing data transmission delay.

[0084] It should be noted that the phrase "the computing circuit interface unit 121 and the memory array 122 are located in the same layer" refers to the fact that the computing circuit interface unit 121 and the memory array 122 can be fabricated simultaneously using the same or approximately the same process steps, so that after fabrication, the heights of the computing circuit interface unit 121 and the memory array 122 along the Z-axis are approximately the same. The phrase "the same layer" does not limit the number of layers; that is, the computing circuit interface unit 121 and the memory array 122 can be fabricated in multiple layers. Nor does it limit the fact that the heights of the computing circuit interface unit 121 and the memory array 122 are identical. For example, the height of the storage devices (e.g., capacitors) included in the memory array 122 can be higher than the height of the transistors in the computing circuit interface unit 121. The phrase "the same or approximately the same process steps" does not require that the process steps of the computing circuit interface unit 121 and the memory array 122 are identical. The phrase "the computing circuit interface unit 121 and the memory array 122 are located in the same layer" is intended to distinguish the case where the computing circuit interface unit 121 and the memory array 122 are fabricated sequentially in a vertical stacking manner. The vertical stacking method means first making the interface unit 121 of the computing circuit and then making the storage array 122 on the interface unit 121 of the computing circuit, or first making the storage array 122 and then making the interface unit 121 of the computing circuit on the storage array 122.

[0085] In an exemplary embodiment, the interface unit 121 of the computing circuit is located on at least one side of the storage array 122. However, the present disclosure is not limited to the above Figure 2 The computing circuit interface unit 121 and the storage array 122 are arranged in the manner described above. In some embodiments, the computing circuit interface unit 121 can also be entirely fabricated on one side of wafer A. For example, the storage array 122 and the computing circuit interface unit 121 are still in the same layer, but the storage array 122 is located on the left side of the lower surface of the second substrate 123, and the computing circuit interface unit 122 is located on the right side of the lower surface of the second substrate 123. In other embodiments, the storage array 122 can be arranged in the central area of ​​the lower surface of the second substrate 123, and the computing circuit interface units 121 are distributed around the storage array 122.

[0086] In an exemplary embodiment, the execution unit 111 of the computing circuit and the peripheral circuit 112 are located on the same layer.

[0087] For example, continue to refer to Figure 2 The execution unit 111 and the peripheral circuit 112 of the computing circuit are both located in different areas of the upper surface of the first substrate 113 .

[0088] The above-mentioned “the execution unit 111 of the computing circuit and the peripheral circuit 112 are located in the same layer” means that the execution unit 111 of the computing circuit and the peripheral circuit 112 can be manufactured simultaneously in the same or approximately the same process steps, so that after the manufacturing is completed, the height of the execution unit 111 of the computing circuit and the peripheral circuit 112 along the Z axis are approximately the same. The “same layer” here does not limit the number of layers, that is, the execution unit 111 of the computing circuit and the peripheral circuit 112 can form multiple layers during the manufacturing process, nor does it limit the height of the execution unit 111 of the computing circuit and the peripheral circuit 112 to be exactly the same. The “same or approximately the same process steps” here do not require that the process steps of the execution unit 111 of the computing circuit and the peripheral circuit 112 are exactly the same. The “execution unit 111 of the computing circuit and the peripheral circuit 112” here are distinguished from the vertical stacking method of manufacturing the execution unit 111 and the peripheral circuit 112 of the computing circuit in sequence.

[0089] Figure 3 Shown is Figure 2 The top view of wafer A and wafer B before bonding, without considering the first bonding layer 114 and the second bonding layer 124. It should be noted that the top view of wafer A here means that before bonding, the second substrate 123 of wafer A is located on the bottom side and the second bonding layer 124 is located on the top side. Figure 3 As shown, in the top view of wafer A, it can be seen that the interface unit 121 of the computing circuit and the storage array 122 are in the same plane (XY axis plane), and the interface unit 121 of the computing circuit includes multiple interface sub-units, each of which is located between adjacent banks. In the top view of wafer B, it can be seen that the execution unit 111 of the computing circuit and the peripheral circuit 112 are in the same plane (XY axis plane), and the peripheral circuit 122 surrounds each execution unit 111. Then, wafer A is hybrid-bonded to the front of wafer B to form the following: Figure 2 The semiconductor device shown.

[0090] In an exemplary embodiment, first bonding layer 114 includes a first contact structure, and second bonding layer 124 includes a second contact structure. Memory array 122 is electrically connected to peripheral circuit 112 via a first portion of the first contact structure and a corresponding first portion of the second contact structure. Execution unit 111 is electrically connected to interface unit 121 of a computing circuit via a second portion of the first contact structure and a corresponding second portion of the second contact structure.

[0091] In the embodiments of the present disclosure, the first contact structure and the second contact structure may include a conductive material, such as Cu (copper), Ni (nickel), or other suitable bonding materials. The present disclosure does not limit the shape, size, or distribution of the first and second contact structures in the first and second bonding layers. In the following embodiments, the first and second contact structures are illustrated using first bonding contacts and second bonding contacts, respectively, but the present disclosure is not limited thereto.

[0092] In the embodiment of the present disclosure, the first contact structure and the second contact structure used to realize the connection between the storage array 122 and the peripheral circuit 112 are respectively referred to as the first-part first contact structure and the first-part second contact structure, and the first contact structure and the second contact structure used to realize the connection between the execution unit 111 and the interface unit 121 of the computing circuit are respectively referred to as the second-part first contact structure and the second-part second contact structure.

[0093] In some embodiments, the first contact structure and the second contact structure only include the first part and the second part mentioned above. In other embodiments, in addition to the first part and the second part mentioned above, the first bonding layer 114 and the second bonding layer 124 may also include a first contact structure and a second contact structure with other uses. In some other embodiments, the first bonding layer and the second bonding layer may also include a dummy first contact structure and a dummy second contact structure, that is, a first contact structure and a second contact structure that are not used to realize the electrical connection between the devices (such as transistors, etc.) in wafer A and wafer B.

[0094] For example, Figure 4 As shown, along Figure 2The cross-sectional view shown is a top-down view taken along the dashed line CD. Wafer A includes a second substrate 123. The lower surface of second substrate 123 includes a memory array 122 and a computing circuit interface unit 121 located on the same layer. It is assumed here that the memory array 122 and the computing circuit interface unit 121 are located on the second device layer 126. Below second device layer 126 is a second interconnect layer 125, which may include multiple second interconnects 1251. Second interconnects 1251 are made of a conductive material. Below second interconnect layer 125 is a second bonding layer 124, which includes multiple second bonding contacts 1241. Wafer B includes a first substrate 113. The first substrate 113 includes peripheral circuits 112 and a computing circuit execution unit 111 located on the same layer. It is assumed here that peripheral circuits 112 and the computing circuit execution unit 111 are located on the first device layer 116. Above the first device layer 116 is a first interconnect layer 115, which includes a plurality of first interconnects 1151. The first interconnects 1151 are made of a conductive material. Above the first interconnect layer 115 is a first bonding layer 114, which includes a plurality of first bonding contacts 1141.

[0095] Figure 4 In the process, the first bonding layer 114 and the second bonding layer 124 are hybrid-bonded, and a bonding interface 130 is formed between wafers A and A. The first bonding contacts 1141 are in contact with the corresponding second bonding contacts 1241 at the bonding interface 130 .

[0096] The first portion of first bonding contacts 1141 and the corresponding first portion of second bonding contacts 1241 are used to connect the memory array 122 in wafer A to the peripheral circuit 112 in wafer B. Part of the second interconnect 1251 can be used to interconnect devices (e.g., transistors and capacitors) in the memory array 122, part of the second interconnect 1251 can be used to interconnect devices (e.g., transistors) in the interface unit 121 of the computing circuit, and part of the second interconnect 1251 can be used to interconnect the memory array 122 and the interface unit 121 of the computing circuit. One end of the first portion of second bonding contacts 1241 is connected to one end of the corresponding first portion of first bonding contacts 1141, and the other end is connected to the corresponding second interconnect 1251 in the second interconnect layer 125. The other end of the first portion of first bonding contacts 1141 is connected to one end of the corresponding first interconnect 1151 in the first interconnect layer 115. Part of the first interconnect 1151 can be used to implement internal interconnection of devices (such as transistors, etc.) in the execution unit 111 of the computing circuit, part of the first interconnect 1151 can be used to implement internal interconnection of devices (such as transistors, etc.) in the peripheral circuit 112, and part of the first interconnect 1151 can be used to implement interconnection between the execution unit 111 of the computing circuit and the peripheral circuit 112. As a result, the memory array 122 in the wafer A can be electrically connected to the corresponding second interconnect 1251, the first part of the second bonding contact 1241, the first part of the first bonding contact 1141 and the first interconnect 1151, thereby enabling data transmission between the memory array 122 and the peripheral circuit 112. This can shorten the data transmission path between the memory array 122 and the peripheral circuit 112, thereby improving data transmission efficiency.

[0097] Second-portion first bonding contacts 1141 and corresponding second-portion second bonding contacts 1241 are used to connect the execution unit 111 of the computing circuit in wafer B with the interface unit 121 of the computing circuit in wafer A. One end of each second-portion second bonding contact 1241 is connected to one end of the corresponding second-portion first bonding contact 1141, and the other end is connected to the corresponding second interconnect 1251 in the second interconnect layer 125. The other end of each second-portion first bonding contact 1141 is connected to one end of the corresponding first interconnect 1151 in the first interconnect layer 115. This allows the interface unit 121 of the computing circuit in wafer A to be electrically connected to the execution unit 111 of the computing circuit via the corresponding second interconnect 1251, second-portion second bonding contacts 1241, second-portion first bonding contacts 1141, and first interconnect 1151, thereby enabling data transmission between the interface unit 121 of the computing circuit and the execution unit 111 of the computing circuit. This can shorten the data transmission path between the interface unit 121 of the computing circuit and the execution unit 111 of the computing circuit, thereby improving data transmission efficiency.

[0098] In some embodiments, the first interconnect layer 115 and the second interconnect layer 125 may be metal interconnect layers, but the present disclosure is not limited thereto. In addition to the first interconnect portion 1151, the second interconnect portion 1251, the first bonding layer 114, and the second bonding layer 124, the first interconnect layer 115 and the second interconnect layer 125, and the first bonding layer 114 and the second bonding layer 124 may also include a dielectric material. The dielectric material may include, for example, SiO, SiN, SiCN, or other suitable dielectric materials, but the present disclosure is not limited thereto.

[0099] In the embodiment of the present disclosure, a bonding process may be applied to bond wafer A to wafer B. The bonding process may be operated at a temperature above 220° C., for example, so that the first contact structure and the second contact structure can melt to form a connection between the second interconnect layer 125 in wafer A and the first interconnect layer 115 in wafer B.

[0100] It should be noted that the present disclosure is not limited to the following. Figure 4 In other embodiments, the first portion of first bonding contacts 1141 and the corresponding first portion of second bonding contacts 1241 are used to connect the storage array 122 in wafer A to the execution unit 111 of the computing circuit in wafer B. The second portion of first bonding contacts 1141 and the corresponding second portion of second bonding contacts 1241 are used to connect the peripheral circuit 112 in wafer B to the interface unit 121 of the computing circuit in wafer A.

[0101] The disclosed embodiment fabricates a memory array of a memory device and an interface unit of a computing circuit on a wafer A, with the memory array and the interface unit of the computing circuit on the same horizontal plane (2D (two-dimensional)); fabricates a peripheral circuit of the memory device and an execution unit of the computing circuit on another wafer B, with the peripheral circuit and the execution unit on the same horizontal plane; and then aligns the front side of wafer A with the front side of wafer B and bonds them together by bonding (e.g., hybrid bonding), so that the memory array on wafer A is electrically connected to the peripheral circuit on wafer B, and the interface unit of the computing circuit on wafer A is electrically connected to the execution unit on wafer B. This can increase storage density, improve data transmission speed, shorten data transmission path, reduce process difficulty and complexity, and accelerate data processing speed.

[0102] In other embodiments, wafer B may be located above wafer A. That is, the upper surface of the second substrate 123 included in wafer A has a storage array 122 and an interface unit 121 of a computing circuit, and the storage array 122 and the interface unit 121 of the computing circuit have a second bonding layer 124. The lower surface of the first substrate 113 of wafer B has a peripheral circuit 112 and an execution unit 111 of a computing circuit, and the lower surface of the peripheral circuit 112 and the execution unit 111 of the computing circuit has a first bonding layer 114. Wafer B is hybrid-bonded to wafer A front-to-front through the first bonding layer 114. The first substrate 113 may be thinned or removed, and then a semiconductor layer may be fabricated on the first device layer 116, but the present disclosure is not limited thereto.

[0103] Figure 5 FIG1 shows a top view of wafer A and wafer B before bonding in another semiconductor device provided by an embodiment of the present disclosure, wherein the first bonding layer 114 and the second bonding layer 124 are not considered. Figure 5 As shown, the top view of wafer A includes memory array 122 and second interconnect region 127 between memory array 122. The top view of wafer B includes execution unit 111 of the computing circuit, peripheral circuit 112 surrounding execution unit 111 of the computing circuit, and first interconnect region 117 corresponding to second interconnect region 127. Wafer A is then hybrid bonded to the front side of wafer B.

[0104] Figure 6 for Figure 5 The semiconductor device is shown as a cross-sectional view taken along the dotted line CD after wafer A and wafer B are bonded. Figure 6 Examples and Figure 4The difference between the embodiments is that the memory array 122 and the computing circuit interface unit 121 in wafer A are stacked vertically, no longer in the same plane, and the memory array 122 is stacked vertically on the computing circuit interface unit 121. By vertically stacking the memory array 122 and the computing circuit interface unit 121, the storage density can be further increased.

[0105] Continue to refer Figure 6 The wafer A may further include a second interconnection region 127 in the same plane or layer as the memory array 122 and the interface unit 121 of the computing circuit. The second interconnection region 127 includes at least one second interconnection structure 1271. The second bonding layer 124 has a second contact structure ( Figure 6 The first end of the second interconnect structure 1271 is connected to the memory array 122 and the interface unit 121 of the computing circuit, respectively, and the second end is connected to at least part of the second bonding contacts 1241. Wafer B includes a first interconnect region 117 corresponding to the second interconnect region 127 in wafer A. The first interconnect region 117 is in the same plane or layer as the peripheral circuit 112 and the execution unit 111 of the computing circuit. A first contact structure (here, the first bonding contact 1141 is used as an example) is provided in the area of ​​the first bonding layer 114 corresponding to the first interconnect region 117. The first bonding contact 1141 is connected to the corresponding first interconnect portion 1151 (extending to the first interconnect portion 1151 in the first interconnect region 117). Thus, electrical connection between wafer A and wafer B can be achieved through the first bonding contacts 1141, the second bonding contacts 1241, and the second interconnect structure 1271 corresponding to the first interconnect region 117 and the second interconnect region 127, respectively.

[0106] Figure 6 In the embodiment, on the one hand, by vertically stacking the storage array 122 and the interface unit 121 of the computing circuit, the storage density can be further improved; on the other hand, by setting the first interconnection area 117 and the second interconnection area 127, the number of alignments of the first bonding contacts 1141 and the second bonding contacts 1241 when bonding wafers A and wafer B can be reduced, thereby reducing the process difficulty.

[0107] Figure 6 In the embodiment, the second interconnect structure 1271 may be a through silicon via (TSV) connection structure, but the present disclosure is not limited thereto.

[0108] It should be noted that although Figure 5 and Figure 6In this embodiment, bonding between wafers A and B is achieved by separately providing first and second interconnect regions 117 and 127 on first and second substrates 113 and 123, respectively. However, the present disclosure is not limited thereto. In other embodiments, the first and second interconnect regions 117 and 127 may be omitted, and the second interconnect structure 1271 may be directly provided between the memory array 122 and the interface unit 121 of the computing circuit, extending through the vertically stacked memory array 122 and the interface unit 121 of the computing circuit to achieve electrical connection between the memory array 122 and the interface unit 121 of the computing circuit, and connected to the second bonding contact 1241 to achieve electrical connection with wafer B. In this case, the second interconnect layer structure may be omitted from wafer A, i.e., the present disclosure does not limit the location of the second interconnect structure 1271.

[0109] For further reference, Figure 6 Wafer A also includes a second interconnect layer (not shown). The second interconnect layer is located between the memory array 122 and the interface unit 121 of the computing circuit. The second interconnect layer includes a plurality of second interconnect portions 1251. Some of the second interconnect portions 1251 extend into the second interconnect region 127 to connect with the second interconnect structure 1271, thereby achieving electrical connection between the second interconnect structure 1271 and the memory array 122 and the interface unit 121 of the computing circuit.

[0110] In the embodiment of the present disclosure, a second interconnection layer is provided between the vertically stacked memory array 122 and the interface unit 121 of the computing circuit. On the one hand, the second interconnection portion 1251 in the second interconnection layer can realize the interconnection of the internal devices of the memory array 122. On the other hand, the interconnection of the internal devices of the interface unit 121 of the computing circuit can be realized. The devices between the memory array 122 and the interface unit 121 of the computing circuit can also be realized. At the same time, the electrical connection with the wafer A can also be realized through the second interconnection layer. This can reduce the number of interconnection layers in the wafer A, reduce the process steps, reduce the production cost, and reduce the difficulty of process preparation.

[0111] However, the present disclosure is not limited to this. In other embodiments, a second interconnect layer may be formed between the memory array 122 and the computing circuit interface unit 121. This interconnect layer is used to interconnect the internal devices of the memory array 122. The second interconnect portion of this interconnect layer extending to the second interconnect region 127 is connected to the second interconnect structure 1271 to achieve electrical connection between wafer B and the computing circuit interface unit 121. Another second interconnect layer is formed on the lower surface of the computing circuit interface unit 121. This second interconnect layer is used to interconnect the internal devices of the computing circuit interface unit 121. The second interconnect portion of this second interconnect layer extending to the second interconnect region 127 is connected to the second interconnect structure 1271 to achieve electrical connection between wafer B and the memory array 122.

[0112] In other embodiments, Figure 6 The difference between the embodiments is that wafer B is located on wafer A, that is, Figure 6 The semiconductor device is shown flipped upside down. At this point, the first substrate 113 may be thinned or removed to form a semiconductor layer.

[0113] Figure 7 FIG1 shows a top view of wafer A and wafer B before bonding in another semiconductor device provided by an embodiment of the present disclosure, where the first bonding layer 114 and the second bonding layer 124 are not considered. Figure 7 As shown, the top view of wafer A shows the interface units 121 of the computing circuit and the second interconnection regions 127 located between the interface units 121 of the computing circuit. The top view of wafer B shows the execution units 111 of the computing circuit and the peripheral circuits 112 surrounding the execution units 111, as well as the first interconnection regions 117 corresponding to the second interconnection regions 127.

[0114] It should be noted that the present disclosure does not limit the locations of the first interconnection region 117 and the second interconnection region 127. The figure is merely an example. As long as the first interconnection region 117 and the second interconnection region 127 can be aligned and bonded after wafers A and B are bonded, it is sufficient.

[0115] Figure 8 Shown Figure 7 The semiconductor device is shown as a cross-sectional view taken along the dotted line CD after wafer A and wafer B are bonded. Figure 8 In an embodiment, the interface unit 121 of the computing circuit is stacked on the memory array 122 .

[0116] exist Figure 6 and Figure 8In an embodiment, the interface unit of the peripheral circuit and the interface unit 121 of the computing circuit can be implemented on the same plane or the same layer, that is, the wafer A also includes the interface unit of the peripheral circuit, and the interface unit of the peripheral circuit is coupled to the storage array 122, and can be used to operate the data in the storage array 122, such as reading or writing. The embodiment of the present disclosure implements the interface unit of the peripheral circuit and the interface unit 121 of the computing circuit on the same wafer A, and both are manufactured on the same plane. On the one hand, units with similar process node requirements can be manufactured on the same wafer, eliminating the need to switch process nodes, reducing process difficulty, and thus reducing process manufacturing costs; on the other hand, the data transmission path between the interface unit of the peripheral circuit and the storage array 122 can be shortened, reducing data transmission delay and increasing data transmission speed.

[0117] In other embodiments, when the interface unit of the peripheral circuit is set on wafer A, the interface unit 121 of the computing circuit can reuse the interface unit of the peripheral circuit, thereby reducing the manufacturing steps of the interface unit 121 of the computing circuit, reducing costs and process complexity.

[0118] exist Figure 6 and Figure 8 In an embodiment, the second substrate 123 may be thinned or removed to form a semiconductor layer.

[0119] In other embodiments, Figure 8 The difference between the embodiments is that wafer B is located on wafer A. At this time, the first substrate 113 may be thinned or removed to form a semiconductor layer.

[0120] Figure 9 FIG. 1 shows a top view of wafer A and wafer B before bonding in another semiconductor device provided by an embodiment of the present disclosure, without considering the first bonding layer 114 and the second bonding layer 124. Figure 9 As shown, the top view of wafer A shows the interface unit 121 and storage array 122 of the computing circuit and the second interconnection area 127. The top view of wafer B shows the execution unit 111 and the first interconnection area 117 of the computing circuit.

[0121] Figure 10 Shown Figure 9 The cross-sectional view of the semiconductor device formed by bonding wafers A and B is taken along the dotted line CD. Figure 10In this embodiment, the memory array 122 and the interface unit 121 of the computing circuit in wafer A are located on the same plane or layer, the second interconnect layer is located on the second bonding layer 124, and at least a portion of the second interconnect portion 1251 in the second interconnect layer extends into the second interconnect region 127 to connect to the second bonding contact 1241. In wafer B, the peripheral circuit 112 and the execution unit 111 of the computing circuit are vertically stacked, and the execution unit 111 of the computing circuit is located on the peripheral circuit 112.

[0122] Figure 10 In an embodiment, wafer B further includes a first interconnect layer, which is located between the execution unit 111 of the computing circuit and the peripheral circuit 112. Thus, the first interconnect portion 1151 in the first interconnect layer can realize the interconnection of the internal devices of the peripheral circuit 112, the interconnection of the internal devices of the execution unit 111, and the interconnection of the devices between the peripheral circuit 112 and the execution unit 111. At the same time, the first interconnect portion 1151 in the first interconnect layer, which extends to the first interconnect region 117, is connected to the first interconnect structure 1171 in the first interconnect region 117, thereby realizing an electrical connection between wafer A and wafer B. This can reduce the number of layers of the first interconnect layer in wafer B, thereby reducing the process difficulty and cost. However, the present disclosure is not limited to this. In other embodiments, a first interconnect layer can be provided on the peripheral circuit 112, and another first interconnect layer can be provided on the execution unit 111 of the computing circuit. This first interconnect layer on the peripheral circuit 112 can be used to interconnect the internal devices of the peripheral circuit 112, and further realize electrical connection with the execution unit 111 and wafer A through the first interconnect portion 1151 extending to the first interconnect region 117. This other first interconnect layer on the execution unit 111 of the computing circuit can be used to interconnect the internal devices of the execution unit 111, and further realize electrical connection with the peripheral circuit 112 and wafer A through the first interconnect portion 1151 extending to the first interconnect region 117.

[0123] In other embodiments, the first interconnection layer and the first interconnection region 117 may not be provided in wafer B (the corresponding second interconnection region 127 may also not be provided in wafer A), and the vertically stacked peripheral circuit 112 and the execution unit 111 may be electrically connected through the first interconnection structure 1171. The first interconnection structure 1171 may be connected to the first bonding contact 1141, and the first bonding contact 1141 is connected to the second bonding contact 1241, thereby making the distribution of the first bonding contact 1141 and the second bonding contact 1241 more uniform.

[0124] In other embodiments, Figure 10 The difference between the embodiments is that wafer B may be located on wafer A.

[0125] Figure 11FIG2 shows a top view of wafer A and wafer B before bonding in another semiconductor device provided by an embodiment of the present disclosure, where the first bonding layer 114 and the second bonding layer 124 are not considered. Figure 11 As shown, the top view of wafer A includes memory array 122 and second interconnect region 127 , while the top view of wafer B includes peripheral circuit 112 and first interconnect region 117 .

[0126] Figure 12 Show Figure 11 The cross-sectional view of the semiconductor device formed by bonding wafers A and B is taken along the dotted line CD. Figure 12 In this embodiment, the peripheral circuit 112 in wafer B is stacked on the execution unit 111 of the computing circuit. The first interconnect layer in wafer B is located between the vertically stacked peripheral circuit 112 and execution unit 111, and at least a portion of the first interconnect portion 1151 in the first interconnect layer extends into the first interconnect region 117 to connect to the first interconnect structure 1171 in the first interconnect region 117. The first interconnect structure 1171 is also connected to the first bonding contact 1141. In the embodiment of the present disclosure, by vertically stacking the interface unit 121 of the computing circuit with the memory array 122, and vertically stacking the peripheral circuit 112 with the execution unit 111 of the computing circuit, the storage density can be further improved and the area occupied by the semiconductor device can be reduced.

[0127] In other embodiments, Figure 12 The difference between the embodiments is that wafer B is located on wafer A.

[0128] Figure 13 FIG. 1 shows a top view of wafer A and wafer B before bonding in another semiconductor device according to an embodiment of the present disclosure, where the first bonding layer 114 and the second bonding layer 124 are not considered. Figure 13 As shown, the top view of wafer A may include memory array 122 and second interconnect region 127 . The top view of wafer B may include execution unit 111 of computing circuit and first interconnect region 117 .

[0129] Figure 14 Shown Figure 13 The cross-sectional view of the semiconductor device formed by bonding wafers A and B is taken along the dotted line CD. Figure 14 In the embodiment, the execution unit 111 of the computing circuit in wafer B is stacked on the peripheral circuit 112 , and the interface unit 121 of the computing circuit in wafer A is stacked on the storage array 122 .

[0130] In other embodiments, Figure 14 The difference between the embodiments is that the peripheral circuit 112 is stacked on the execution unit 111 of the computing circuit.

[0131] Figure 15 FIG. 1 shows a top view of wafer A and wafer B before bonding in another semiconductor device according to an embodiment of the present disclosure, where the first bonding layer 114 and the second bonding layer 124 are not considered. Figure 15 As shown, the top view of wafer A includes the interface unit 121 of the computing circuit and the second interconnection area 127 , and the top view of wafer B includes the execution unit 111 of the computing circuit and the first interconnection area 117 .

[0132] Figure 16 Show Figure 15 The cross-sectional view of the semiconductor device formed by bonding wafers A and B is taken along the dotted line CD. Figure 16 In the embodiment, in wafer A, the storage array 122 is stacked on the interface unit 121 of the computing circuit; in wafer B, the execution unit 111 is stacked on the peripheral circuit 112 .

[0133] In other embodiments, Figure 16 The difference between the embodiments is that wafer B is located on wafer A.

[0134] Figure 17 FIG. 1 shows a top view of wafer A and wafer B before bonding in another semiconductor device according to an embodiment of the present disclosure, where the first bonding layer 114 and the second bonding layer 124 are not considered. Figure 17 As shown, the top view of wafer A includes the interface unit 121 of the computing circuit and the second interconnection area 127 , and the top view of wafer B includes the peripheral circuit 112 and the first interconnection area 117 .

[0135] Figure 18 Show Figure 17 The cross-sectional view of the semiconductor device formed by bonding wafers A and B is taken along the dotted line CD. Figure 18 In the embodiment, the peripheral circuit 112 in wafer B is stacked on the execution unit 111 of the computing circuit, and the storage array 122 in wafer A is stacked on the interface unit 121 of the computing circuit.

[0136] In other embodiments, Figure 18 The difference between the embodiments is that wafer B is located on wafer A.

[0137] On top Figures 2 to 18 In the embodiment, wafer A and wafer B are bonded in a front-to-front manner. However, the embodiments of the present disclosure are not limited thereto. In other embodiments, wafer A and wafer B can be bonded in a front-to-back manner or a back-to-front manner.

[0138] In an exemplary embodiment, the first semiconductor structure 110 further includes a first bonding layer and a first substrate, with the execution unit 111 and the peripheral circuit 112 located between the first bonding layer and the first substrate. The second semiconductor structure 120 further includes a second substrate and a second bonding layer. The second substrate includes opposing first and second sides, the second bonding layer is located on the second side of the second substrate, and the memory array 122 and the interface unit 121 for the computing circuit are located on the first side of the second substrate. The first semiconductor structure 110 and the second semiconductor structure 120 are bonded together via the first bonding layer and the second bonding layer.

[0139] like Figure 19 As shown, wafer A is hybrid bonded from the back side to the front side of wafer B. Figure 19 and Figure 2 The difference of the embodiment is that the lower surface of the second substrate 123 of wafer A has a second bonding layer 124, and the upper surface of the second substrate 123 has a storage array 122 and an interface unit 121 of a computing circuit. Figure 19 In the embodiment, the middle regions of wafers A and B are respectively provided with a second interconnection region 127 and a first interconnection region 117 to achieve electrical connection between wafers A and B. Wafers A and B are electrically connected via the second bonding layer 124 and the first bonding layer 114 .

[0140] Figure 19 In an embodiment, the memory array 122 may be stacked on the computing circuit interface unit 121 , that is, the computing circuit interface unit 121 is first fabricated on the second substrate 123 , and then the memory array 122 is fabricated on the computing circuit interface unit 121 .

[0141] Furthermore, a second interconnection layer 125 can be made between the vertically stacked interface unit 121 of the computing circuit and the storage array 122 to realize the interconnection of the internal devices of the interface unit 121 of the computing circuit, the interconnection of the internal devices of the storage array 122, and the interconnection of the devices between the interface unit 121 of the computing circuit and the storage array 122, and the second interconnection layer 125 is connected to the second interconnection area 127 to realize the interconnection between the interface unit 121 of the computing circuit and the storage array 122 and the execution unit 111 and the peripheral circuit 112 in wafer B.

[0142] Figure 20 Show Figure 19 The top view of wafer A and wafer B in the semiconductor device shown in FIG. 1 before bonding, where the first bonding layer 114 and the second bonding layer 124 are not considered. Figure 20As shown, the top view of wafer A shows the storage array 122 and the second interconnection area 127, and the top view of wafer B shows the execution unit 111 and peripheral circuit 112 of the computing circuit and the first interconnection area 117, and then the back side of wafer A is hybrid bonded with the front side of wafer B.

[0143] Figure 21 Show Figure 20 The cross-sectional view of the semiconductor device formed by bonding wafers A and B is taken along the dotted line CD. Figure 21 In this embodiment, the second bonding layer 124 is located on the lower surface of the second substrate 123, the computing circuit interface unit 121 is located above the second substrate 123, and the memory array 122 is located above the computing circuit interface unit 121. The second interconnect structure 1271 extends from the second interconnect region 127 through the second substrate 123 until it connects with the second bonding contact 1241. For other details, please refer to the above embodiment.

[0144] The semiconductor device provided by the embodiment of the present disclosure vertically stacks (3D (three-dimensional)) the storage array 122 on the interface unit 121 of the computing circuit to further improve the storage density, and the peripheral circuit 112 and the execution unit 111 are made on another wafer B. The back side of wafer A is aligned with the front side of wafer B and bonded through a second interconnect structure 1271 (for example, a TSV connection structure).

[0145] In an exemplary embodiment, the memory array 122 is stacked on a first side or a second side opposite to the first side of the computing circuit interface unit 121. The second semiconductor structure 120 further includes a second interconnect layer located between the memory array 122 and the computing circuit interface unit 121.

[0146] For example, Figure 21 As shown, the second interconnection layer is located between the vertically stacked memory array 122 and the interface unit 121 of the computing circuit. The second interconnection portion 1251 in the second interconnection layer realizes the interconnection of the internal devices of the memory array 122, the interconnection of the internal devices of the interface unit 121 of the computing circuit, and the interconnection of the devices between the memory array 122 and the interface unit 121 of the computing circuit. The second interconnection portion 1251 is extended to the second interconnection region 127 through at least a portion of the second interconnection portion 1251 to connect to the second interconnection structure 1271.

[0147] In an exemplary embodiment, the second semiconductor structure 120 further includes an interface unit of a peripheral circuit.

[0148] In an exemplary embodiment, the interface unit 121 of the computing circuit and the interface unit of the peripheral circuit are located on the same layer.

[0149] exist Figure 21In an embodiment, the interface unit of the peripheral circuit (also referred to as the IO unit) can also be set in wafer A. In this way, the communication distance between the peripheral circuit 112 and the interface unit 121 of the computing circuit is shorter, and devices with the same or similar process nodes are made on the same wafer. Furthermore, the interface unit of the peripheral circuit can be set in the same plane or the same layer as the interface unit 121 of the computing circuit. At this time, the interface unit 121 of the computing circuit and the interface unit of the peripheral circuit can be collectively referred to as the total interface unit (total IO).

[0150] In an exemplary embodiment, the computing circuit interface unit 121 and the peripheral circuit interface unit are located on the same layer, and the computing circuit interface unit 121 reuses the peripheral circuit interface unit. That is, the computing circuit interface unit 121 does not need to be manufactured, and the computing circuit execution unit 111 directly uses the peripheral circuit interface unit. After the peripheral circuit 112 reads data from the storage array 122, it sends it to the computing circuit execution unit 111 for calculation processing.

[0151] In an exemplary embodiment, the first semiconductor structure 110 further includes a first substrate and a first bonding layer. The first substrate includes opposing first and second sides. The execution unit 111 and peripheral circuit 112 are located on the first side of the first substrate, and the first bonding layer is located on the second side of the first substrate. The second semiconductor structure 120 further includes a second substrate and a second bonding layer. The memory array 122 and the interface unit 121 for the computing circuit are located between the second substrate and the second bonding layer. The first semiconductor structure 110 and the second semiconductor structure 120 are bonded together via the first bonding layer and the second bonding layer.

[0152] For example, Figure 22 As shown, wafer B can be stacked on wafer A. In this case, wafer B's first bonding layer 114 is located on the lower surface of the first substrate 113, while the peripheral circuit 112 and the computing circuit's execution unit 111 are located on the upper surface of the first substrate 113. That is, the backside of wafer B is hybrid-bonded to the frontside of wafer A. Furthermore, a first interconnect layer 115 can be formed on the peripheral circuit 112 and the computing circuit's execution unit 111. The computing circuit's interface unit 121 is formed on the second substrate 123 of wafer A, and a memory array 122 is formed on the computing circuit's interface unit 121. Furthermore, a second interconnect layer 125 can be formed between the memory array 122 and the computing circuit's interface unit 121. A second bonding layer 124 is formed on the second interconnect layer 125. The first bonding layer 114 is bonded to the second bonding layer 124. The first interconnect region 117 is located above the second interconnect region 127.

[0153] In an exemplary embodiment, the first bonding layer 114 includes a first contact structure, and the second bonding layer 124 includes a second contact structure. The first semiconductor structure 110 also includes a first interconnect structure, and the second semiconductor structure 120 also includes a second interconnect structure. The first end of the first interconnect structure is connected to the execution unit 111 and the peripheral circuit 112, respectively, and the second end is connected to at least a portion of the first contact structure. The first end of the second interconnect structure is connected to the memory array 122 and the interface unit 121 of the computing circuit, respectively, and the second end is connected to at least a portion of the second contact structure.

[0154] In an exemplary embodiment, the first interconnect structure and the second interconnect structure include through silicon via connection structures.

[0155] Figure 23 Show Figure 22 The semiconductor device shown is a cross-sectional view cut from top to bottom along the dotted line CD. Figure 23 As shown, wafer B is stacked on wafer A. The first bonding layer 114 of wafer B is located on the lower surface of the first substrate 113, and the peripheral circuit 112 and the execution unit 111 of the computing circuit are located on the same layer, both of which are located on the upper surface of the first substrate 113. The upper surface of the first substrate 113 also includes a first interconnection region 117, which can be located on the same layer as the peripheral circuit 112 and the execution unit 111 of the computing circuit. The first interconnect portion 1151 in the first interconnection layer is used to realize the internal device interconnection of the peripheral circuit 112, the internal device interconnection of the execution unit 111, and the device interconnection between the peripheral circuit 112 and the execution unit 111. At least a portion of the first interconnect portion 1151 extends to the first interconnection region 117 and connects to the first interconnect structure 1171 in the first interconnection region 117. The first interconnect structure 1171 extends from the first interconnection region 117 and penetrates the first substrate 113 until it contacts the first bonding contact 1141. First bonding contact 1141 is connected to corresponding second bonding contact 1241. On the second substrate 123 of wafer A, there is an interface unit 121 for the computing circuit. A memory array 122 is vertically stacked on the interface unit 121 for the computing circuit. A second bonding layer 124 is provided on the memory array 122. In some embodiments, a second interconnect layer exists between the memory array 122 and the interface unit 121 for the computing circuit. At least a portion of a second interconnect portion 1251 in the second interconnect layer extends to the second interconnect region 127 and connects to a second interconnect structure 1271 in the second interconnect region 127.

[0156] The second interconnect structure 1271 extends to connect with the second bonding contact 1241 .

[0157] Figure 24In the embodiment, wafer A is stacked vertically on wafer B, that is, the back side of wafer A is bonded to the front side of wafer B. The second bonding layer 124 in wafer A is arranged on the lower side of the second substrate 123, the memory array 122 is arranged on the upper side of the second substrate 123, and the interface unit 121 of the computing circuit is arranged above the memory array 122. The second interconnect layer is arranged between the memory array 122 and the interface unit 121 of the computing circuit. At least a portion of the second interconnect portion 1251 in the second interconnect layer extends to the second interconnect region 127 and is connected to the second interconnect structure 1271. The second interconnect structure 1271 extends and penetrates the second substrate 123 until it is connected to the second bonding contact 1241. For other contents of this embodiment, reference can be made to the other embodiments described above.

[0158] Figure 25 In this embodiment, wafer B is stacked vertically on wafer A, that is, the back side of wafer B is bonded to the front side of wafer A. In wafer B, the first bonding layer 114 is located on the lower side of the first substrate 113, and the peripheral circuit 112, the execution unit 111 of the computing circuit, and the first interconnection area 117 are located on the upper side of the first substrate 113. The first interconnection layer is located on the peripheral circuit 112 and the execution unit 111 of the computing circuit. The first interconnect portion 1151 in the first interconnection layer extends at least partially into the first interconnection area 117 and is connected to the first interconnection structure 1171. The first interconnection structure 1171 extends and penetrates the first substrate 113 until it connects to the first bonding contact 1141 in the first bonding layer 114. In wafer A, the memory array 122 is located on the second substrate 123, the interface unit 121 of the computing circuit is located on the memory array 122, and the second bonding layer 124 is located on the interface unit 121 of the computing circuit. In some embodiments, the second interconnect layer is located between the computing circuit interface unit 121 and the memory array 122. At least a portion of the second interconnect portion 1251 in the second interconnect layer extends to the second interconnect region 127 and connects to the second interconnect structure 1271. The second interconnect structure 1271 is also connected to the second bonding contact 1241.

[0159] In an exemplary embodiment, the execution unit 111 is stacked on a first side or a second side opposite the first side of the peripheral circuit 112. The first semiconductor structure 110 further includes a first interconnect layer located between the execution unit 111 and the peripheral circuit 112. That is, in some embodiments, the execution unit 111 can be vertically stacked with the peripheral circuit 112.

[0160] Figure 26In this embodiment, wafer A is vertically stacked on wafer B. In wafer B, peripheral circuit 112 is vertically stacked above the execution unit 111 of the computing circuit. A first interconnect layer is located between peripheral circuit 112 and the execution unit 111 of the computing circuit. At least a portion of first interconnect 1151 in the first interconnect layer extends into first interconnect region 117 and connects to first interconnect structure 1171. In wafer A, memory array 122 is vertically stacked above the interface unit 121 of the computing circuit. For other details, please refer to the other embodiments described above.

[0161] Figure 27 In this embodiment, wafer B is vertically stacked on wafer A. Peripheral circuit 112 in wafer B is vertically stacked on execution unit 111 of the computing circuit. Memory array 122 in wafer A is vertically stacked on interface unit 121 of the computing circuit. For other details, please refer to the other embodiments above.

[0162] Figure 28 In this embodiment, wafer A is vertically stacked on wafer B. In wafer B, the execution unit 111 of the computing circuit is vertically stacked on the peripheral circuit 112. The memory array 122 in wafer A is vertically stacked on the interface unit 121 of the computing circuit. For other details, please refer to the other embodiments above.

[0163] Figure 29 In this embodiment, wafer B is vertically stacked on wafer A. In wafer B, the execution unit 111 of the computing circuit is vertically stacked on the peripheral circuit 112. The memory array 122 in wafer A is vertically stacked on the interface unit 121 of the computing circuit. For other details, please refer to the other embodiments above.

[0164] Figure 30 In this embodiment, wafer A is vertically stacked on wafer B. In wafer A, the interface unit 121 of the computing circuit is vertically stacked on the storage array 122. In wafer B, the peripheral circuit 112 is vertically stacked on the execution unit 111 of the computing circuit. For other details, please refer to the other embodiments above.

[0165] Figure 31 In this embodiment, wafer B is vertically stacked on wafer A. The interface unit 121 of the computing circuit in wafer A is vertically stacked on the storage array 122. In wafer B, the peripheral circuit 112 is vertically stacked on the execution unit 111 of the computing circuit. For other details, please refer to the other embodiments above.

[0166] Figure 32 In this embodiment, wafer A is vertically stacked on wafer B. In wafer A, the interface unit 121 of the computing circuit is vertically stacked on the storage array 122. In wafer B, the execution unit 111 of the computing circuit is vertically stacked on the peripheral circuit 112. For other details, please refer to the other embodiments above.

[0167] Figure 33 In this embodiment, wafer B is vertically stacked on wafer A. In wafer A, the interface unit 121 of the computing circuit is vertically stacked on the storage array 122. In wafer B, the execution unit 111 of the computing circuit is vertically stacked on the peripheral circuit 112. For other details, please refer to the other embodiments above.

[0168] Figure 34 In this embodiment, wafer A is vertically stacked on wafer B. On wafer A, the computing circuit interface unit 121, storage array 122, and second interconnect region 127 are located on the same plane or layer, with the second interconnect layer located above the computing circuit interface unit 121 and storage array 122. On wafer B, peripheral circuit 112 is vertically stacked above the computing circuit execution unit 111. For other details, please refer to the other embodiments described above.

[0169] Figure 35 In this embodiment, wafer B is stacked vertically on wafer A. In wafer A, the computing circuit interface unit 121, storage array 122, and second interconnect region 127 are located on the same plane or layer, with the second interconnect layer located above the computing circuit interface unit 121 and storage array 122. Second interconnect 1251 extends to second interconnect region 127 and connects to second bonding contact 1241. In wafer B, peripheral circuit 112 is stacked vertically above the computing circuit execution unit 111. For other details, please refer to the other embodiments described above.

[0170] Figure 36 In this embodiment, wafer A is stacked vertically on wafer B. In wafer A, the computing circuit interface unit 121, storage array 122, and second interconnect region 127 are located in the same plane or layer, and the second interconnect layer is located above the computing circuit interface unit 121 and storage array 122. The second interconnect portion 1251 extends to the second interconnect region 127 and connects to the second interconnect structure 1271, which is connected to the second bonding contact 1241. In wafer B, the computing circuit execution unit 111 is stacked vertically above the peripheral circuit 112. For other details, please refer to the other embodiments described above.

[0171] Figure 37 In this embodiment, wafer B is vertically stacked on wafer A. In wafer A, the computing circuit interface unit 121, storage array 122, and second interconnect region 127 are located on the same plane or layer, and the second interconnect layer is located above the computing circuit interface unit 121 and storage array 122. Second interconnect portion 1251 extends to second interconnect region 127 and connects to second bonding contact 1241. In wafer B, the computing circuit execution unit 111 is vertically stacked above the peripheral circuit 112. For other details, please refer to the other embodiments described above.

[0172] In the above embodiment, the interface unit of the peripheral circuit can be provided on the same wafer as the interface unit 121 of the computing circuit. Furthermore, the interface unit of the peripheral circuit can be located on the same layer as the interface unit 121 of the computing circuit. In other embodiments, the interface unit 121 of the computing circuit can reuse the interface unit of the peripheral circuit.

[0173] It is understandable that although the above embodiment illustrates the example of separately setting the first interconnection area 117 and the second interconnection area 127 to realize the interconnection between wafer A and wafer B, the present disclosure is not limited to this. In other embodiments, the first interconnection area 117 and the second interconnection area 127 may not be set, but the execution units 111 and peripheral circuits 112 vertically stacked in the same wafer, and / or the interface unit 121 and the storage array 122 of the computing circuit are directly interconnected through the TSV connection structure, thereby reducing the setting of the first interconnection layer and / or the second interconnection layer. In other embodiments, when the first interconnection area 117 and the second interconnection area 127 are set, the setting position of the first interconnection area 117 and the second interconnection area 127 is not limited, as long as the electrical connection between wafer A and wafer B can be achieved.

[0174] The semiconductor device provided by the embodiments of the present disclosure can, on the one hand, save area and improve storage density by separating the storage array and the peripheral circuits and manufacturing them on two wafers; on the other hand, some computing circuits are performed inside the semiconductor device to assist in calculations, which reduces the data transmission distance, improves the calculation speed, and reduces the power consumption and load of AI calculations such as GPUs.

[0175] In an exemplary embodiment, execution unit 111 includes an arithmetic unit, which is connected to an interface unit 121 of a computing circuit. Execution unit 111 is configured to read data from a storage array 122 via interface unit 121 of the computing circuit and input the data to the arithmetic unit; the arithmetic unit is configured to perform arithmetic processing on the data to obtain a calculation result.

[0176] In the disclosed embodiments, the calculation result obtained by the execution unit 111 in the computing circuit can be an intermediate result, that is, the intermediate result can be used to obtain the final result later; or it can be the final result. After the execution unit 111 obtains the intermediate result, the intermediate result can be returned to the execution unit 111 for further processing, or it can be sent externally. For example, it can be sent externally through the interface unit of the peripheral circuit, or sent externally through the interface unit 121 of the computing circuit.

[0177] For example, Figure 38FIG2 is a schematic diagram of a computing circuit provided in an exemplary embodiment. The computing circuit may include an interface unit 121 of the computing circuit and an execution unit 111 of the computing circuit. The interface unit 121 of the computing circuit may be used to obtain data from a memory device and input the obtained data to the execution unit 111. The execution unit 111 includes an operator 1111 to perform calculations on the obtained data to obtain a calculation result. The execution unit 111 may output the obtained calculation result as data, for example, storing the calculation result in a memory device or sending it to a CPU or GPU.

[0178] In an exemplary embodiment, the execution unit 111 further includes a control circuit, and the calculation circuit further includes a first register circuit. The control circuit is configured to store the calculation result in the first register circuit and return the calculation result stored in the first register circuit to the operator for continued calculation processing. The control circuit is also configured to output the calculation result.

[0179] In an exemplary embodiment, the operator includes an adder and / or a multiplier. In some embodiments, the convolution operation in the AI ​​calculation can be implemented by the adder and the multiplier. In other embodiments, the convolution operation in the AI ​​calculation can be implemented by the adder. In the following embodiments, the operator includes both an adder and a multiplier for illustration, but the present disclosure is not limited to this.

[0180] For example, Figure 39 As shown, the execution unit 111 of the computing circuit includes an adder 1111a, a multiplier 1111b, a decoder 1112, and a control circuit 1113. The decoder 1112 is coupled to the interface unit 121 of the computing circuit and is configured to decode data input through the interface unit 121 of the computing circuit. The decoder 1112 is coupled to the adder 1111a, the multiplier 1111b, and the control circuit 1113, and is configured to input the decoded data to at least one of the adder 1111a, the multiplier 1111b, the control circuit 1113, etc., to perform computing processing, such as addition and / or multiplication operations. Figure 48In an embodiment, the execution unit of the computing circuit may further include a first register circuit 481, which is coupled to the control circuit 1113. The control circuit 1113 may store the calculation result output by the adder 1111a and / or the multiplier 1111b in the first register circuit 481. The control circuit 1113 may also return the calculation result stored in the first register circuit to the adder 1111a and / or the multiplier 1111b to continue the calculation process. The control circuit 1113 may also output the calculation result stored in the first register circuit 481 to the outside, for example, to a peripheral circuit, instruct the peripheral circuit to output the calculation result to a GPU or CPU for further calculation processing, or instruct the peripheral circuit to store the calculation result in a storage array. For another example, the calculation result may be output to the outside through the interface unit 121 of the computing circuit.

[0181] It should be noted that the control circuit in the embodiment of the present disclosure refers to a circuit provided in the execution unit of the computing circuit, and is used to control the computing and processing of data read from the storage array through the interface unit of the computing circuit.

[0182] The semiconductor device provided by the embodiment of the present disclosure includes a computing circuit and a memory device, the memory device includes a memory array and a peripheral circuit, the computing circuit includes an interface unit and an execution unit of the computing circuit and a first register circuit, the execution unit may further include a decoder, an adder, a multiplier and a control circuit, the execution unit and the first register circuit are more matched in manufacturing process nodes, and the interface unit of the computing circuit is more matched in manufacturing process nodes of the memory array, so the interface unit of the computing circuit and the memory array are made on the same wafer, for example, a 14nm process node can be used (for example only, not limited to this); the execution unit (including the decoder, adder, multiplier, control circuit), the first register circuit and the peripheral circuit are made on another wafer, for example, a 7nm process node can be used (for example only, not limited to this). The two wafers are then bonded, thereby improving process preparation efficiency, reducing preparation complexity, improving preparation efficiency and reducing costs.

[0183] The computing circuit provided by the embodiment of the present disclosure is suitable for AI computing. When the computing circuit is used for AI large model computing, that is, the computing circuit can pre-process at least a portion of the data read from the storage array 122, obtain the intermediate results, and then transmit them to the server for computing. For example, the interface unit 121 of the computing circuit sends the data to the decoder 1112, the decoder 1112 decodes or decodes the encoded data, and the decoded or decoded data is input to the adder 1111a and / or the multiplier 1111b, and the control circuit 1113 receives the intermediate result and determines whether the intermediate result is returned to the adder 1111a and / or the multiplier 1111b to continue computing, or is sent to the first register circuit 481 for caching to output the intermediate result. This is in line with the AI ​​large model computing method, that is, there are many intermediate processing results, not the final processing results.

[0184] Generative artificial intelligence (AI) reasoning involves AI computing. For example, transformer models, a common model in AI systems, typically use tensor processing units (TPUs) and memory for computation. Large transformer models require large amounts of data and computation, which requires high power consumption and sufficient memory. When memory access speed lags behind processor computation speed, memory bottlenecks prevent high-performance processors from operating efficiently and pose a significant constraint on high-performance computing (HPC). This problem is known as the memory wall.

[0185] To address one or more of the aforementioned issues and break down the memory wall, the present disclosure introduces a solution, including a semiconductor device comprising a memory device and a computing circuit for performing calculations on the data in the memory device. An execution unit 111 of the computing circuit and a peripheral circuit 112 of the memory device are provided in a first semiconductor structure 110, a storage array 122 of the memory device and an interface unit 121 of the computing circuit are provided in a second semiconductor structure 120, the first semiconductor structure 110 and the second semiconductor structure 120 are bonded together, and calculations are performed under the control of a control circuit 1113 of the execution unit 111. In this way, a portion of the computing tasks of the AI ​​system can be distributed to the computing circuit of the AI ​​system, especially tasks requiring a large data width. Without transferring large amounts of data from the memory device to the processor of the AI ​​system to perform calculations, the computing tasks are completed within the computing circuit, while the processor can handle other calculations. Therefore, the introduction of the computing circuit in the semiconductor device effectively improves the computing speed of the AI ​​system.

[0186] Figure 40A block diagram of a system 10 having a host 20 and a storage and computing system 30 in an embodiment of the present disclosure is shown. The system 10 may be a mobile phone, a desktop computer, a laptop computer, a tablet computer, a vehicle-mounted computer, a game controller, a printer, a positioning device, a wearable electronic device, a smart sensor, a virtual reality (VR) device, an augmented reality (AR) device, an artificial intelligence (AI) device, or any other suitable electronic device having a storage device therein. The system 10 may also be cloud computing, a distributed storage and computing system, a centralized storage and computing system, a server, etc. Figure 40 As shown, the system 10 may include a host 20 and a storage and computing system 30. The storage and computing system 30 may include one or more non-volatile memory devices 34 (e.g., Figure 40 NAND flash memory in ), one or more volatile memory devices 36 (e.g., Figure 40 DRAM in), and a memory controller 32 (also referred to as a controller for short). Figure 40 In an embodiment, the storage and computing system 30 may further include one or more computing circuits 38, each of which is coupled to at least one volatile memory device 36 (e.g., DRAM) and configured to read data from the at least one volatile memory device 36, perform computing on the read data, and obtain a computing result. In some embodiments, the computing circuit 38 may return the computing result to the volatile memory device 36 (e.g., DRAM) for storage, for example, via a peripheral circuit of the DRAM. In other embodiments, the computing circuit 38 may also be coupled to the memory controller 32 to send the computing result to the memory controller 32, and the memory controller 32 may send the computing result to the host 20.

[0187] The storage and computing system 30 may be configured to store data and perform computations using the data in response to requests from the host 20. A memory controller 32 may provide a physical connection between the host 20 and the storage and computing system 30. Specifically, the memory controller 32 may provide a data interface between the host and the storage and computing system 30 based on the format of the host's data bus. The memory controller 32 may decode instructions provided by the host 20 and access one or more non-volatile memory devices 34 and one or more volatile memory devices 36.

[0188] One or more non-volatile memory devices 34 may be configured to store data, providing persistent storage for the data.

[0189] One or more volatile memory devices 36 can temporarily store programming data provided from the host or data read from the non-volatile memory device 34. When a read request is sent from the host 20, if the requested data in the non-volatile memory device 34 is cached in the volatile memory device 36, the memory controller 32 can send the data cached in the volatile memory device 36 directly to the host 20.

[0190] In some embodiments, the volatile memory device 36 may also be configured to store a mapping table between logical addresses and physical addresses of data stored in the non-volatile memory device 34. In some embodiments, the memory controller 32 may communicate with the volatile memory device 36 using at least one communication protocol or technology standard (e.g., associated with dual-inline-memory modules (DIMMs), registered DIMMs (RDIMMs), load-reduced DIMMs (LRDIMMs), unregistered DIMMs (UDIMMs), etc.).

[0191] In some embodiments, the host 20 may include a processor (e.g., a tensor processing unit (TPU), a central processing unit (CPU)), or a system on chip (SoC) (e.g., an application processor (AP)). The host 20 may be configured to send data to the storage and computing system 30 or receive data from the storage and computing system 30.

[0192] The non-volatile memory device 34 may include, but is not limited to, NAND flash memory, resistive random access memory (RRAM), nano random access memory (NRAM), phase change random access memory (PCRAM), ferroelectric random access memory (FRAM), magnetoresistive random access memory (MRAM), etc. The volatile memory device 36 may include, but is not limited to, dynamic random access memory (DRAM), static random access memory (SRAM), etc.

[0193] According to some embodiments, the memory controller 32 is coupled to the non-volatile memory device 34 and the host 20 and is configured to control the non-volatile memory device 34. The memory controller 32 can manage data stored in the non-volatile memory device 34 and communicate with the host 20. In some embodiments, the memory controller 32 is designed to operate in a low duty cycle environment, such as a secure digital (SD) card, a compact flash (CF) card, a universal serial bus (USB) flash drive, or other media used in electronic devices such as personal computers, digital cameras, mobile phones, etc. In some embodiments, the memory controller 32 is designed to operate in a high duty cycle environment, such as an SSD or an embedded multimedia card (eMMC), which is used as a data storage device for mobile devices such as smartphones, tablets, laptops, etc., as well as enterprise storage arrays. The memory controller 32 can be configured to control the operations of the non-volatile memory device 34 (e.g., read operations, erase operations, and program operations). The memory controller 32 may also be configured to manage various functions regarding data stored or to be stored in the non-volatile memory device 34, including, but not limited to, bad block management, garbage collection, logical-to-physical address translation, wear leveling, and the like. In some embodiments, the memory controller 32 may also be configured to process error checking and correcting (ECC) codes for data read from or written to the non-volatile memory device 34. The memory controller 32 may also perform any other appropriate functions, such as formatting the non-volatile memory device 34. The memory controller 32 may communicate with an external device (e.g., the host 20) according to a specific communication protocol. For example, the memory controller 32 can communicate with an external device through at least one of various interface protocols, such as a USB protocol, an MMC protocol, a peripheral component interconnect (PCI) protocol, a high-speed PCI (PCI-E) protocol, an advanced technology attachment (ATA) protocol, a serial ATA protocol, a parallel ATA protocol, a small computer miniature interface (SCSI) protocol, an enhanced minidisk interface (ESDI) protocol, an integrated drive electronics (IDE) protocol, a FireWire protocol, etc.

[0194] The memory controller 32 and one or more volatile memory devices 36 and the computing circuit 38 can be integrated into various types of memory devices, for example, included in the same package. That is, the memory computing system 30 can be implemented and packaged into different types of terminal electronic products. Figure 41In one example shown, the memory controller 32, the single volatile memory device 36, and the single computing circuit 38 may be integrated into a memory card 40. The memory card 40 may include a PC card (PCMCIA, Personal Computer Memory Card International Association), a CF card, a Smart Media (SM) card, a memory stick, a multimedia card (MMC, RS-MMC, MMCmicro), an SD card (SD, miniSD, microSD, SDHC), a UFS, etc. The memory card 40 may also include a computer that connects the memory card 40 to a host (e.g., Figure 40 A memory card connector 42 is coupled to the host 20 in the memory card connector.

[0195] Figure 42 The block diagram of another storage and calculation device is shown as an example. Figure 42 In another example shown, the memory controller 32 and the plurality of volatile memory devices 36 and the plurality of computing circuits 38 may be integrated into a memory computing device 50. The memory computing device 50 may also include a computer that connects the memory computing device 50 to a host (e.g., Figure 40 In some embodiments, the storage capacity and / or operating speed of the storage computing device 50 is greater than the storage capacity and / or operating speed of the memory card 40.

[0196] In such Figure 43 In the illustrated example of a system with DRAM, the system includes a system-on-chip (SoC) on a printed circuit board (PCB), one or more memories, and one or more computing circuits 38. The memories include one or more DRAMs 36, and the SoC includes a processor (such as a graphics processing unit (GPU)) 434, a DRAM controller 433, and a DRAM physical layer 435. The DRAM controller 433 is responsible for scheduling read and write instructions and controlling the timing of the DRAMs 36. The DRAM physical layer 435 is responsible for encoding the scheduled instructions according to the requirements of the DRAMs 36, sending the corresponding write data to the DRAMs 36, and receiving data read from the DRAMs 36. Under the control of the DRAM controller 433, the computing circuit 38 reads data from the DRAMs 36, performs computations on the data, obtains computation results, and outputs the computation results, for example, to the GPU 434.

[0197] Furthermore, embodiments of the present disclosure provide a storage and computing device comprising a memory device and a computing circuit coupled to the memory device. The memory device comprises a memory array and peripheral circuitry coupled to the memory array. The computing circuitry comprises an execution unit and an interface unit. The execution unit and the peripheral circuitry are contained in a first semiconductor structure. The interface unit between the memory array and the computing circuitry is contained in a second semiconductor structure. The first semiconductor structure is bonded to the second semiconductor structure.

[0198] In an exemplary embodiment, the first semiconductor structure further includes a first bonding layer and a first substrate. The execution unit and the peripheral circuit are located between the first bonding layer and the first substrate. The second semiconductor structure further includes a second bonding layer. An interface unit between the memory array and the computing circuit is located on a side of the second bonding layer away from the first semiconductor structure. The first and second semiconductor structures are bonded via the first and second bonding layers.

[0199] In an exemplary embodiment, the first bonding layer includes a first contact structure, and the second bonding layer includes a second contact structure. The memory array is connected to the peripheral circuit via a first portion of the first contact structure and a corresponding first portion of the second contact structure. The execution unit is connected to the interface unit of the computing circuit via a second portion of the first contact structure and a corresponding second portion of the second contact structure.

[0200] In an exemplary embodiment, the first semiconductor structure further includes a first bonding layer and a first substrate, the execution unit and the peripheral circuit being located between the first bonding layer and the first substrate. The second semiconductor structure further includes a second substrate and a second bonding layer, the second substrate including opposing first and second sides, the second bonding layer being located on the second side of the second substrate, and the interface unit between the memory array and the computing circuit being located on the first side of the second substrate. The first semiconductor structure and the second semiconductor structure are bonded to each other via the first bonding layer and the second bonding layer.

[0201] In an exemplary embodiment, the first bonding layer includes a first contact structure, and the second bonding layer includes a second contact structure. The first semiconductor structure also includes a first interconnect structure, and the second semiconductor structure also includes a second interconnect structure. A first end of the first interconnect structure is connected to the execution unit and the peripheral circuit, respectively, and a second end is connected to at least a portion of the first contact structure. A first end of the second interconnect structure is connected to the memory array and the interface unit of the computing circuit, respectively, and a second end is connected to at least a portion of the second contact structure.

[0202] In an exemplary embodiment, the memory array is stacked on a first side of the interface unit of the computing circuit or on a second side opposite to the first side. The second semiconductor structure further includes a second interconnect layer located between the memory array and the interface unit of the computing circuit.

[0203] In an exemplary embodiment, the second semiconductor structure further includes an interface unit of the peripheral circuit, and the interface unit of the computing circuit and the interface unit of the peripheral circuit are located in the same layer.

[0204] Figure 44 A schematic diagram of a memory computing device 60 is shown as an example. The memory computing device 60 includes a memory device and a computing circuit 38 coupled to the memory device. The memory device includes a memory array 122 and a peripheral circuit 112 coupled to the memory array 122. The memory array 122 may include physical banks 66 of memory cells. Each bank 66 of memory cells may include a memory cell 622. Each memory cell 622 includes a transistor 624 and a storage device 626 coupled to the vertical transistor 624. In some embodiments, the memory array 122 is a DRAM cell array, and the storage device 626 is a capacitor for storing charge as binary information stored by the corresponding DRAM cell. In some embodiments, the memory array 122 is a PCM cell array, and the storage device 626 is a PCM element (e.g., including a chalcogenide alloy) for storing binary information of the corresponding PCM cell based on the different resistivity of the PCM element in the amorphous phase and the crystalline phase. In some embodiments, memory array 122 is an array of FRAM cells, and storage device 626 is a ferroelectric capacitor for storing binary information of a corresponding FRAM cell based on switching between two polarization states of a ferroelectric material under an external electric field. In the following embodiments, the storage device is exemplified as a capacitor that stores charge as binary information stored by a corresponding DRAM cell, but the present disclosure is not limited thereto.

[0205] like Figure 44 As shown, the memory cells 622 can be arranged in a two-dimensional (2D) array having rows and columns. The memory computing device 60 may include: word lines 627, which couple the peripheral circuit 112 with the memory array 122 for controlling the switching of the transistors 624 in the memory cells 622 located in a row, and bit lines 629, which couple the peripheral circuit 112 with the memory array 122 for sending data to the memory cells 622 located in a column and / or receiving data from the memory cells 622 located in a column. That is, each word line 627 is coupled to the memory cells 622 in a corresponding row, and each bit line 629 is coupled to the memory cells 622 in a corresponding column.

[0206] The memory device 626 may include any device capable of storing binary data (e.g., 0s and 1s), including, but not limited to, capacitors for DRAM cells and FRAM cells, and PCM elements for PCM cells. In some embodiments, the transistor 624 controls the selection and / or state switching of the corresponding memory device 626 coupled to the transistor 624. The peripheral circuit 112 may be coupled to the memory array 122 via the bit lines 629, the word lines 627, and any other appropriate metal connections. As described above, the peripheral circuit 112 may include any appropriate circuitry for applying voltage and / or current signals to each memory cell 622 via the word lines 627 and the bit lines 629, and sensing voltage and / or current signals from each memory cell 622 to facilitate the operation of the memory array 122. The peripheral circuit 112 may include any appropriate analog, digital, or mixed-signal circuitry for facilitating the associated operation of the array of memory cells by applying voltage and / or current signals to each target memory cell and sensing voltage and / or current signals from each target memory cell. Additionally, the peripheral circuit 64 may include various types of peripheral circuits formed using metal oxide semiconductor (MOS) technology.

[0207] In an exemplary embodiment, the peripheral circuit includes a second register circuit, which at least partially reuses the first register circuit. Generally, the number of registers in the peripheral circuit may be less than the number of registers in the execution unit. The second register circuit in the peripheral circuit can at least partially reuse the first register circuit in the execution unit of the computing circuit, thereby reducing the number of registers in the semiconductor device by sharing registers in the peripheral circuit and the execution unit.

[0208] refer to Figure 45 , the second register circuit of the peripheral circuit 112 includes an address register (not shown) and / or a data register 77. The peripheral circuit 112 includes a sense amplifier 71 of the peripheral circuit, a column decoder / bit line (BL) driver 72, a row decoder / word line (WL) driver 73, a voltage generator 74, a control logic 75, an address register, a data register 77 and an interface unit 79 of the peripheral circuit. It should be understood that in some other examples, the peripheral circuit 112 may also include Figure 45 Additional circuitry not shown.

[0209] Figure 45In an embodiment, the computing circuit 38 may include a computing circuit interface unit 121 and a computing circuit execution unit 111. The computing circuit interface unit 121 may obtain data stored in the storage array 122 via the peripheral circuit sense amplifier 71 (e.g., via the data register 77), and send the obtained data to the computing circuit execution unit 111. The computing circuit execution unit 111 performs computation on the input data to obtain a computation result, and may store the computation result in the storage array 122 via the computing circuit interface unit 121, or send the computation result out via the peripheral circuit interface unit 79.

[0210] The sense amplifier 71 of the peripheral circuit may be configured to read data from the memory array 122 according to a control signal from the control logic 75. The column decoder / bit line driver 72 may be configured to be controlled by the control logic 75 and to select one or more memory cells by applying a bit line voltage generated from the voltage generator 74.

[0211] The row decoder / word line driver 73 may be configured to be controlled by the control logic 75 and to select / deselect a bank 66 of memory cells of the memory array 62 and to select / deselect a word line of the bank 66 of memory cells. The row decoder / word line driver 73 may also be configured to drive a word line using a word line voltage generated from a voltage generator 74. As described in detail below, the row decoder / word line driver 73 is configured to apply a read voltage to a selected word line in a read operation on a memory cell coupled to the selected word line.

[0212] The voltage generator 74 may be configured to be controlled by the control logic 75 and generate word line voltages (eg, read voltage, program voltage, refresh voltage, etc.) and bit line voltages to be supplied to the memory array 122 .

[0213] The control logic 75 can be coupled to each peripheral circuit described above and is configured to control the operation of each peripheral circuit. The address register and data register 77 can be coupled to the control logic 75 and configured to store status information, command operation code (OP code) and command address for controlling the operation of each peripheral circuit. The interface unit 79 for the peripheral circuit can be coupled to the control logic 75 and act as a control buffer to buffer control commands received from a host (not shown) and relay them to the control logic 75, as well as buffer status information received from the control logic 75 and relay them to the host. The interface unit 79 for the peripheral circuit can also be coupled to the column decoder / bit line driver 72 and can act as a data input / output (IO) interface.

[0214] In an exemplary embodiment, the interface unit of the computing circuit includes a sense amplifier coupled to the memory array.

[0215] In the above Figure 45 In this embodiment, the computing circuit can reuse the sense amplifier 71 of the peripheral circuit, thereby simplifying the circuit structure and reducing costs. However, the present disclosure is not limited to this. In other embodiments, the computing circuit can include its own independent sense amplifier, so that the computing circuit does not need to read data from the sense amplifier of the peripheral circuit, thereby improving data transmission speed.

[0216] For example, Figure 46 As shown, it is consistent with the above Figure 45 The difference is that the interface unit of the computing circuit includes a sensor amplifier (SA) 801 of the computing circuit and a data interface 802 of the computing circuit. The SA 801 of the computing circuit is coupled to the storage array 122, can read data from the storage array 122, and send the read data to the execution unit 111 of the computing circuit through the data interface 802 of the computing circuit.

[0217] In the storage and computing device provided by the embodiments of the present disclosure, the interface unit of the computing circuit and the peripheral circuit of the memory device each have their own sense amplifier. This is because the computing circuit is used to perform logical operations only rarely, and most of the time DRAM is used for data storage. Data storage and data logic operations are different when reading data from the storage array. Data storage may be batch reading, while logical operations may be reading a small amount of data from a specified location. Therefore, by having the computing circuit and the memory device each have their own sense amplifier, the normal data storage function of the DRAM can be maintained unaffected. However, the present disclosure is not limited to this. The interface unit of the computing circuit can also reuse the sense amplifiers in the peripheral circuit, so that the peripheral circuit first reads the data, then sends it to the interface unit of the computing circuit, and then sends it to the execution unit of the computing circuit to obtain the calculation result, and the calculation result is then sent to the peripheral circuit.

[0218] The storage and computing device provided by the embodiment of the present disclosure has a storage array (such as a DRAM storage array) and a peripheral circuit made on two wafers respectively, which can save the peripheral circuit area, increase the storage array area, and improve the storage density. At the same time, there are computing circuits on both wafers, and preliminary calculations can be completed in the computing circuits, and some auxiliary calculations can be done inside the semiconductor device, so that at least part of the data does not need to be transmitted to the CPU or GPU, etc., and the transmission distance is short, which can increase the computing speed. In addition, the functions of the computing circuits on the two wafers are different. The two wafers are bonded by Hybrid Bonding in a front-to-front or front-to-back manner, so that the process node does not need to be switched on the same wafer, which can reduce costs and improve efficiency.

[0219] Figure 47A schematic diagram of a storage array is shown as an example. Figure 45 As shown, the memory array 122 includes a bank (abbreviated as B) 66 of memory cells, and the bank 66 of memory cells is coupled to an interface unit 121 of at least one computing circuit via a data register 77 of a peripheral circuit. Figure 46 As shown, the bank 66 of memory cells is coupled to a data interface 802 of at least one computing circuit via a SA 801 of the computing circuit. Figure 47 As shown, each memory array 122 may include a bank 66 of sixteen memory cells, such as Figure 47 A bank 66 of multiple memory cells is coupled to peripheral circuitry 112 .

[0220] In some embodiments, the number of at least one computing circuit is equal to the number of banks 66 of memory cells, and each computing circuit corresponds to a corresponding bank. Figure 47 In some embodiments, the number of at least one computing circuit is less than the number of banks 66 of memory cells. In some embodiments, the number of at least one computing circuit is half the number of banks of memory cells, and each computing circuit corresponds to two corresponding banks of memory cells. For example, in Figure 47 In some embodiments, the number of at least one computing circuit is one-fourth the number of banks of memory cells, and each computing circuit corresponds to four corresponding banks 66 of memory cells. For example, in Figure 47 In the embodiment, the number of computing circuits may be four. The number of computing circuits may be designed as needed in practice, and the above embodiments in the present disclosure are intended to be illustrative and should not be construed as limiting the present disclosure.

[0221] AI systems are primarily used in two areas: training and reasoning, and the present disclosure can be primarily used for AI reasoning, where data is input into a trained AI module for recognition and analysis to obtain the expected results of the input data. In AI reasoning, calculations are performed based on the input data (which may be referred to as first data) and reference data pre-stored in the AI ​​system (which may be referred to as second data) to confirm one or more properties of the input data.

[0222] In AI reasoning, in many cases, the input data may be one-dimensional data, and the reference data may be two-dimensional data, where the first data is a one-dimensional vector and the second data is a two-dimensional matrix. In the AI ​​system provided in the embodiments of the present disclosure, a computing circuit is provided to implement computing near a memory device, where the computing is performed outside the memory device. However, the present disclosure is not limited to this, and computing can also be implemented by a storage unit of a memory device.

[0223] In some embodiments, the one-dimensional first data may be equivalent to a row of data having a length of a, and the two-dimensional second data (i.e., an a×b matrix) may be equivalent to b columns (e.g., b=6, i.e., having six columns, col1 to col6, but this is only for illustration), each column having a length of a, and a and b are both positive integers greater than or equal to 1. In some embodiments, the first data may be a two-dimensional matrix including more than one row of equal data length, and dimensionality reduction may be performed on the more than one row of first data to decompose the first data into multiple single rows, thereby applying the present disclosure.

[0224] In some embodiments, the first data and the second data can be obtained from the interface unit 79 of the peripheral circuit of the memory device and programmed into the multiple banks of memory cells in the memory array 122. For example, the first data can be stored in one bank 66 of memory cells in the multiple banks 66 of memory cells in the memory array 122, and the second data can be stored in the other bank 66 of memory cells in the multiple banks 66 of memory cells in the memory array 122. The first data can be updated after each calculation. The second data can be stored in the bank 66 of memory cells for multiple calculations using different first data and can be updated according to instructions from the host 20.

[0225] In some implementations, the second data may be stored in the non-volatile memory device 34. The memory controller 32 may read the second data from the non-volatile memory device 34 and program the second data to the volatile memory device 36 before performing the calculation.

[0226] In some embodiments, the first data comprises a row, and the control logic 75 is configured to program the row of the first data into one of the plurality of banks of memory cells. For example, the first data may be programmed into B1 of the plurality of banks of memory cells.

[0227] In some embodiments, the second data includes M columns, where M is a positive integer and M ≥ 2. The control logic 75 is configured to program each of the M columns into memory cells in M ​​banks of a plurality of banks of memory cells, where the number of the plurality of banks of memory cells is greater than M. In some embodiments, the number of banks of memory cells in the memory array 122 is sixteen, and the number of columns of the second data is six. The six columns of the second data can then be programmed into memory cells in six of the sixteen banks. For example, column 0 can be programmed into B2, column 1 can be programmed into B3, column 2 can be programmed into B4, column 3 can be programmed into B5, column 4 can be programmed into B6, and column 5 can be programmed into B7.

[0228] In some embodiments, the control logic 75 is further configured to send the first data from the plurality of banks of memory cells to the at least one computing circuit. In some embodiments, the control logic 75 is further configured to send the second data from the plurality of banks of memory cells to the at least one computing circuit.

[0229] In some embodiments, each execution unit 111 is configured to perform a convolution operation on the first data with the second data and obtain a first calculation result. The control circuit 1113 then sends the first calculation result to the interface unit 79 of the peripheral circuit or the interface unit 121 of the computing circuit to perform further operations, and then performs the convolution operation and generates a second calculation result. The calculation result may include the first calculation result and the second calculation result. The control circuit 1113 then continuously sends the second calculation result to the interface unit 79 of the peripheral circuit or the interface unit 121 of the computing circuit. In some embodiments, each of the calculation results can be stored in the first register circuit 481, the data register 77 of the peripheral circuit 112, or in one or more memory cells of the memory array 122.

[0230] In some embodiments, the storage and computing device includes a computing circuit, and the computing circuit sequentially performs the convolution operation between the first data and the second data. In some embodiments, the storage and computing device includes more than one computing circuit, and the convolution operation between the first data and the second data is performed simultaneously by different computing circuits.

[0231] At least one computing circuit is provided independently of the peripheral circuits. As the number of computing circuits within a semiconductor device increases, the computing speed of the memory device improves. In some embodiments, the memory array 122 is divided into more than one bank of memory cells, each bank including multiple memory cells. The number of computing circuits is equal to the number of banks of memory cells, meaning that the computing circuits correspond to multiple banks of memory cells. For example, the memory array 122 is divided into 128 banks, and the number of computing circuits is also 128. In some embodiments, the number of computing circuits is less than the number of banks of memory cells. For example, the memory array 122 is divided into 128 banks, and the number of computing circuits can be 100, 64, 50, or 40, or other numbers less than 128. In some embodiments, the number of computing circuits is half the number of banks of memory cells, with one computing circuit corresponding to each of the two banks of memory cells. For example, the memory array 122 is divided into 128 banks, and the number of computing circuits is 64. In some embodiments, the number of computing circuits is one-fourth the number of banks of memory cells, with one computing circuit corresponding to each of the four banks of memory cells. For example, the memory array 122 is divided into 128 banks, and the number of computing circuits is 32. The number of computing circuits can be set and adjusted based on the requirements of the AI ​​system, and the embodiments of the present disclosure are intended to illustrate the present disclosure and should not be construed as limiting.

[0232] Furthermore, an embodiment of the present disclosure provides a storage and computing system, comprising a memory device, a controller coupled to the memory device, and a computing circuit. The memory device comprises a memory array and a peripheral circuit coupled to the memory array. The computing circuit comprises an execution unit and an interface unit. The execution unit and the peripheral circuit are contained in a first semiconductor structure. The memory array and the interface unit of the computing circuit are contained in a second semiconductor structure. The first semiconductor structure is bonded to the second semiconductor structure. The controller is configured to control the memory device to perform storage operations and control the computing circuit to read data from the memory array to perform computing operations.

[0233] The data transmission speed between traditional DRAM and GPU / CPU limits the computing speed. The data transmission distance between DRAM and GPU is long and the transmission speed is slow, resulting in slow computing speed and high power consumption. Figure 48As shown, in the storage and computing device provided by the embodiment of the present disclosure, the memory device and the computing circuit are integrated in the same semiconductor device. Specifically, the storage array 122 of the memory device and the interface unit 121 of the computing circuit are arranged in the second semiconductor structure, and the peripheral circuit 112 of the memory device and the execution unit 111 of the computing circuit are arranged in the first semiconductor structure. Then, the first semiconductor structure and the second semiconductor structure are bonded, so that at least part of the calculation can be completed inside the semiconductor device, with fast computing speed and low power consumption. That is, this DRAM near-memory computing architecture can improve computing speed and has high storage density.

[0234] Furthermore, an embodiment of the present disclosure also provides a system comprising a memory device, a controller, and a computing circuit. The memory device comprises a plurality of banks of storage units and a peripheral circuit coupled to the storage units. The peripheral circuit comprises a control logic configured to program first data and second data into the plurality of banks. The computing circuit is configured to obtain first data and second data from at least one bank, and perform calculations based on the first data and second data to obtain calculation results. A data path bus is coupled to the computing circuit to transmit the first data and second data to the computing circuit. The computing circuit comprises a control circuit configured to receive the calculation results, and return the calculation results to the execution unit for further processing, or output the calculation results to the outside.

[0235] In some embodiments, the system can be any electrical system to which an AI system is applied, such as a computer, a digital camera, a mobile phone, a smart appliance, the Internet of Things (IoT), a server, a base station, and the like. In the present disclosure, data processing and calculations of the AI ​​system can be performed by a computing circuit. In some embodiments, computing tasks that consume a lot of resources can be distributed to the computing circuit instead of a TPU or GPU by integrating at least one computing circuit and a memory device in the same semiconductor device, thereby improving the performance of the AI ​​system. The number of computing circuits can be designed based on the needs of the AI ​​system. By integrating more computing circuits and memory devices into the same semiconductor device, the AI ​​system will be more efficient.

[0236] like Figure 49 As shown, the embodiment of the present disclosure also provides a method for manufacturing a semiconductor device, which may include the following steps.

[0237] In S631 , a first semiconductor structure is provided, where the first semiconductor structure includes a first substrate and an execution unit of a computing circuit and a peripheral circuit of a memory device located on the first substrate.

[0238] In S632 , a second semiconductor structure is provided, the second semiconductor structure including a memory array of a memory device and an interface unit of a computing circuit.

[0239] It is understandable that the execution order of the above steps S631 and S632 is not limited, and the two can be executed in parallel or one after the other.

[0240] In S633 , the second semiconductor structure is bonded to the first semiconductor structure.

[0241] In an exemplary embodiment, providing a first semiconductor structure includes: forming the execution unit and the peripheral circuit on the first substrate; and forming a first bonding layer on the execution unit and the peripheral circuit. Providing a second semiconductor structure includes: forming an interface unit between the memory array and the computing circuit on the second substrate; and forming a second bonding layer on the interface unit between the memory array and the computing circuit. Bonding the second semiconductor structure to the first semiconductor structure includes: bonding the first bonding layer to the second bonding layer.

[0242] In an exemplary embodiment, forming the execution unit and the peripheral circuit on the first substrate includes: forming the execution unit and the peripheral circuit on the first substrate by using a synchronous process.

[0243] In the embodiments of the present disclosure, the use of a simultaneous process to form the execution unit and peripheral circuits means that the peripheral circuits can be fabricated simultaneously with the execution unit fabrication process, thereby placing the execution unit and the peripheral circuits on the same plane or layer. However, this does not require that the fabrication process steps for the execution unit and the peripheral circuits be identical. A simultaneous process refers to a process in which the execution unit is first formed on a first substrate and then the peripheral circuits are formed on top of the execution unit.

[0244] In an exemplary embodiment, the memory array includes a plurality of memory cells, each of which includes a transistor and a memory device. Forming an interface unit between the memory array and the computing circuit on a second substrate includes: forming the interface unit between the transistors in the memory array and the computing circuit on the second substrate; and forming the memory device in the memory array on the second substrate.

[0245] Taking dynamic random access memory (DRAM) as an example, it is a random access semiconductor memory that stores each bit of data in a memory cell having a capacitor and a transistor (also referred to as an array transistor), both of which can be based on metal oxide semiconductor (MOS) technology. The capacitor can be set to a charged state or a discharged state. These two states represent the two values ​​of the bit, which are conventionally referred to as 0 and 1. DRAM also includes peripheral circuits, which include transistors (to distinguish them, they can be referred to as peripheral transistors). The peripheral circuits and array transistors handle data input / output (IO) and memory cell operations (e.g., write or read). The capacitor can be formed in a planar configuration, a stacked configuration, or a trench configuration, depending on the manufacturing method. The capacitor can be coupled to the first doped region (e.g., the drain region) of the array transistor to be charged or discharged through the first doped region. The word line can be coupled to the gate of the array transistor to turn the array transistor on or off. The bit line can be coupled to the second doped region (e.g., the source region) of the array transistor and serves as a path for charging or discharging the capacitor.

[0246] In the disclosed embodiments, capacitors, array transistors, and interface units of a computing circuit of a DRAM memory device are processed on one wafer (e.g., an array wafer), and peripheral transistors of a DRAM memory device and execution units of a computing circuit are processed on another wafer (e.g., a peripheral wafer). The array wafer and peripheral wafer can be processed separately using logic technology nodes that achieve the desired IO speed and functionality. After the array wafer and peripheral wafer are processed, electrical connection between the two wafers is achieved by crossing the bonding interface between the wafers in one process step, thereby achieving higher storage density, a simpler process flow, and a shorter cycle time.

[0247] In an exemplary embodiment, the first bonding layer includes a first contact structure, and the second bonding layer includes a second contact structure. Bonding the first bonding layer to the second bonding layer includes: connecting a first portion of the first contact structure corresponding to the memory array to a first portion of the second contact structure corresponding to the peripheral circuit; and connecting a second portion of the first contact structure corresponding to the execution unit to a second portion of the second contact structure corresponding to the interface unit of the computing circuit.

[0248] Figures 50 to 54 The following is a schematic diagram illustrating various steps of a method for manufacturing a semiconductor device. The following example is given in which the memory cell in the memory array is a 1T1C structure (i.e., one array transistor and one capacitor), but the present disclosure is not limited thereto. Figure 50As shown, the transistors (eg, 1T) of the memory array 122 and the interface unit 121 of the computing circuit are first formed on the second substrate 123 of the wafer A. Figure 51 As shown, the capacitors of the memory array 122 are then formed on the transistors of the memory array 122 on the wafer A. In some embodiments, a memory cell includes a transistor and a capacitor, for example, Figure 51 As shown, a capacitor is formed on each transistor of the storage array 122. However, the present disclosure is not limited to this. In other embodiments, for example, a 3T1C structure (i.e., a storage unit includes 3 transistors and 1 capacitor) may also be used. The present disclosure does not limit the structure of the storage unit. In the embodiment of the present disclosure, considering that both the transistors of the storage array and the interface unit of the computing circuit include transistors, the transistors of the storage array and the interface unit of the computing circuit may be manufactured simultaneously in a synchronous process, thereby improving the manufacturing efficiency and reducing the number of process steps. However, the present disclosure is not limited to this. In other embodiments, the transistors of the storage array 122 may be manufactured on the second substrate 123 first, and then the interface unit 121 of the computing circuit may be manufactured. As shown in FIG. Figure 52 As shown, in Figure 51 On this basis, a second bonding layer 124 can be formed on the interface unit of the memory array and the computing circuit.

[0249] like Figure 53 As shown, the execution unit 111 of the computing circuit and the peripheral circuit 112 of the memory device are formed on the first substrate 113 of the wafer B. Since the execution unit 111 of the computing circuit and the peripheral circuit 112 of the memory device are both transistors, they can also be manufactured by a simultaneous process, but the present disclosure is not limited to this. Figure 54 As shown, in Figure 53 On the basis of the first bonding layer 114 is formed on the execution unit 111 of the computing circuit and the peripheral circuit 112 of the memory device. Figure 54 Wafer B shown and Figure 52 Wafer A is shown facing front-side, connected by hybrid bonding or TSV direct bonding.

[0250] In an exemplary embodiment, providing a first semiconductor structure includes: forming the execution unit and the peripheral circuit on the first substrate; and forming a first bonding layer on the execution unit and the peripheral circuit. Providing a second semiconductor structure includes: forming an interface unit between the memory array and the computing circuit on a second substrate; and forming a second bonding layer under the second substrate.

[0251] Bonding the second semiconductor structure to the first semiconductor structure includes: bonding the first bonding layer to the second bonding layer.

[0252] In an exemplary embodiment, forming the execution unit and the peripheral circuit on the first substrate includes: forming the execution unit, the peripheral circuit, and a first interconnect region on the first substrate; forming a first interconnect layer on the execution unit and the peripheral circuit; and forming a first interconnect structure in the first interconnect region, wherein a first end of the first interconnect structure is connected to the execution unit and the peripheral circuit, respectively, through the first interconnect layer. Forming a first bonding layer on the execution unit and the peripheral circuit includes: forming the first bonding layer on the execution unit and the peripheral circuit, the first bonding layer including a first contact structure, connecting a second end of the first interconnect structure to at least a portion of the first contact structure.

[0253] In an exemplary embodiment, forming the interface unit between the memory array and the computing circuit on a second substrate includes: forming the interface unit of the computing circuit and a second interconnect region on the second substrate; forming a second interconnect layer on the interface unit of the computing circuit and the second interconnect region; forming the memory array on the second interconnect layer; and forming a second interconnect structure in the second interconnect region that passes through the second substrate, such that a first end of the second interconnect structure is connected to the memory array and the interface unit of the computing circuit, respectively, through the second interconnect layer. Forming a second bonding layer under the second substrate includes: forming the second bonding layer under the second substrate, the second bonding layer including a second contact structure, such that a second end of the second interconnect structure is connected to at least a portion of the second contact structure.

[0254] Figures 55 to 58 Schematic diagram showing various steps of another method for manufacturing a semiconductor device. Figure 55 As shown, the interface unit 121 of the computing circuit is first completed on the second substrate 123 of the wafer A, and the middle area of ​​the second substrate 123 is reserved as the second interconnection area 127. Figure 56 As shown in FIG. 1 , a second interconnect layer 125 is fabricated on the interface unit 121 of the computing circuit. The transistors of the memory array 122 are then fabricated on the second interconnect layer 125. Figure 57 As shown, the capacitors of the memory array 122 are fabricated on the transistors of the memory array 122, and then a second bonding layer 124 is formed on the lower side of the second substrate 123. A TSV interconnection structure (second interconnection structure) is fabricated in the second interconnection region 127 to penetrate the back of the second substrate 123 and connect to the second bonding layer 124. Figure 57 As shown, the peripheral circuit 112 and the execution unit 111 of the computing circuit are made on the first substrate 113 of wafer B (compared with wafer A, a more advanced process technology can be used), and the middle area is reserved for TSV interconnection, namely the first interconnection area 117. Figure 58As shown, a first interconnection layer 115 is made on the top of wafer B, and a TSV interconnection structure is made in the first interconnection area 117, namely the first interconnection structure. Figure 57 Wafer A and Figure 58 Wafer B is shown face-to-back, connected via hybrid bonding and TSVs.

[0255] It is understandable that the manufacturing process of the semiconductor device provided by the embodiment of the present disclosure is not limited to the above-mentioned methods. When the structure of the semiconductor device changes, such as the vertical stacking of the peripheral circuit and the execution unit, the corresponding process steps will also change.

[0256] In order to produce the wafers A and B mentioned above, various semiconductor manufacturing processes can be applied. These semiconductor manufacturing processes may include: deposition process, photolithography process, etching process, wet cleaning process, metrology measurement process, real-time defect analysis, surface planarization process or implantation process, etc. The deposition process may also include chemical vapor deposition (CVD) process, physical vapor deposition (PVD) process, diffusion process, sputtering process, atomic layer deposition (ALD) process, electroplating process, etc.

[0257] For example, an implantation process may be applied to form a p-type doped well (PW), a first doped region, and a second doped region of an array transistor. A deposition process may be applied to form a gate dielectric layer, a gate structure, and a spacer of the array transistor. The first dielectric stack may be formed by a deposition process. In order to form a contact structure in the first dielectric stack, a plurality of contact openings may be formed by applying a photolithography process and an etching process. A deposition process may then be applied to fill the contact openings with a conductive material. Thereafter, a surface planarization process may be used to remove excess conductive material above the top surface of the first dielectric stack.

[0258] In subsequent process steps, various additional interconnect structures (e.g., metallization layers with conductive lines and / or vias) may be formed above the DRAM memory device. Such interconnect structures electrically connect the DRAM memory device to other contact structures and / or active devices to form functional circuits. Additional device features such as passivation layers and input / output structures may also be formed.

[0259] The foregoing description of specific embodiments can be readily modified and / or adapted for various applications. Therefore, based on the teaching and guidance provided herein, such adaptations and modifications are intended to be within the meaning and range of equivalents of the disclosed embodiments.

[0260] Although specific configurations and arrangements have been discussed, it should be understood that this is for illustrative purposes only. Therefore, other configurations and arrangements may be used without departing from the scope of this disclosure. In addition, the subject matter as described in this disclosure may also be used in a variety of other applications. The functions and structural features as described in this disclosure may be combined, adjusted, modified, and rearranged with one another in a manner consistent with the scope of this disclosure.

Claims

1. A semiconductor device, characterized in that: include: a first semiconductor structure including an execution unit of a computing circuit and a peripheral circuit of a memory device; The second semiconductor structure is bonded to the first semiconductor structure, and the second semiconductor structure includes a memory array of the memory device and an interface unit of the computing circuit.

2. The semiconductor device according to claim 1, wherein The first semiconductor structure further includes a first bonding layer and a first substrate; the execution unit and the peripheral circuit are located between the first bonding layer and the first substrate; The second semiconductor structure further includes a second bonding layer; the interface unit between the memory array and the computing circuit is located on a side of the second bonding layer away from the first semiconductor structure; The first semiconductor structure and the second semiconductor structure are bonded via the first bonding layer and the second bonding layer.

3. The semiconductor device according to claim 2, wherein: The first bonding layer includes a first contact structure, and the second bonding layer includes a second contact structure; The memory array is electrically connected to the peripheral circuit via the first portion of the first contact structure and the corresponding first portion of the second contact structure; The execution unit is connected to the interface unit of the computing circuit via a second portion of the first contact structure and a corresponding second portion of the second contact structure.

4. The semiconductor device according to claim 1, wherein The first semiconductor structure further includes a first bonding layer and a first substrate, and the execution unit and the peripheral circuit are located between the first bonding layer and the first substrate; The second semiconductor structure further includes a second substrate and a second bonding layer, the second substrate including a first side and a second side opposite to each other, the second bonding layer is located on the second side of the second substrate, and the interface unit between the memory array and the computing circuit is located on the first side of the second substrate; The first semiconductor structure and the second semiconductor structure are bonded via the first bonding layer and the second bonding layer.

5. The semiconductor device according to claim 1, wherein The first semiconductor structure further includes a first substrate and a first bonding layer, the first substrate includes a first side and a second side opposite to each other, the execution unit and the peripheral circuit are located on the first side of the first substrate, and the first bonding layer is located on the second side of the first substrate; The second semiconductor structure further includes a second substrate and a second bonding layer, and the interface unit between the memory array and the computing circuit is located between the second substrate and the second bonding layer; The first semiconductor structure and the second semiconductor structure are bonded via the first bonding layer and the second bonding layer.

6. The semiconductor device according to claim 4 or 5, characterized in that The first bonding layer includes a first contact structure, and the second bonding layer includes a second contact structure; The first semiconductor structure further includes a first interconnect structure, and the second semiconductor structure further includes a second interconnect structure; A first end of the first interconnect structure is connected to the execution unit and the peripheral circuit respectively, and a second end is connected to at least part of the first contact structure; A first end of the second interconnect structure is connected to the memory array and the interface unit of the computing circuit respectively, and a second end is connected to at least a portion of the second contact structure.

7. The semiconductor device according to claim 6, wherein: The first interconnect structure and the second interconnect structure include through silicon via connection structures.

8. The semiconductor device according to claim 4 or 5, characterized in that The memory array is stacked on a first side of the interface unit of the computing circuit or on a second side opposite to the first side; The second semiconductor structure further includes a second interconnect layer located between the memory array and an interface unit of the computing circuit.

9. The semiconductor device according to claim 1, wherein The execution units are stacked on a first side of the peripheral circuit or a second side opposite to the first side; The first semiconductor structure further includes a first interconnect layer located between the execution unit and the peripheral circuit.

10. The semiconductor device according to claim 1, wherein The second semiconductor structure further includes an interface unit for the peripheral circuit.

11. The semiconductor device according to claim 10, wherein: The interface unit of the computing circuit and the interface unit of the peripheral circuit are located on the same layer.

12. The semiconductor device according to claim 1, wherein The interface unit of the computing circuit and the interface unit of the peripheral circuit are located in the same layer, and the interface unit of the computing circuit reuses the interface unit of the peripheral circuit.

13. The semiconductor device according to claim 1, wherein The interface unit of the computing circuit includes a sense amplifier coupled to the memory array.

14. The semiconductor device according to claim 1, wherein The execution unit includes an arithmetic unit, and the execution unit is coupled to the interface unit of the computing circuit; The execution unit is configured to read data from the storage array through the interface unit of the computing circuit and input the data to the operator; The operator is used to perform calculations on the data to obtain calculation results.

15. The semiconductor device according to claim 14, wherein: The execution unit further includes a control circuit, and the calculation circuit further includes a first register circuit; The control circuit is configured to store the calculation result in the first register circuit, and return the calculation result stored in the first register circuit to the operator to continue calculation processing; the control circuit is also configured to output the calculation result to the outside.

16. The semiconductor device according to claim 15, wherein: The peripheral circuit includes a second register circuit, and the second register circuit at least partially reuses the first register circuit.

17. The semiconductor device according to claim 14, wherein: The operator includes an adder and / or a multiplier.

18. The semiconductor device according to claim 1, wherein The interface unit of the computing circuit includes a plurality of interface subunits, and the storage array includes a plurality of physical storage banks; Wherein, each interface subunit of the computing circuit is located between two adjacent physical storage bodies.

19. The semiconductor device according to claim 1, wherein The interface unit of the computing circuit is located on at least one side of the storage array.

20. The semiconductor device according to claim 1, wherein The interface unit of the computing circuit and the storage array are located in the same layer.

21. The semiconductor device according to claim 1, wherein The execution unit of the computing circuit and the peripheral circuit are located in the same layer.

22. The semiconductor device according to claim 1, wherein The memory array includes a DRAM memory array.

23. A method for manufacturing a semiconductor device, characterized in that: include: Providing a first semiconductor structure, the first semiconductor structure including a first substrate and an execution unit of a computing circuit and a peripheral circuit of a memory device located on the first substrate; providing a second semiconductor structure comprising a memory array of the memory device and an interface unit of the computing circuit; The second semiconductor structure is bonded to the first semiconductor structure.

24. The method according to claim 23, wherein Providing a first semiconductor structure includes: forming the execution unit and the peripheral circuit on the first substrate; forming a first bonding layer on the execution unit and the peripheral circuit; Wherein, a second semiconductor structure is provided, comprising: forming an interface unit between the memory array and the computing circuit on a second substrate; forming a second bonding layer on the interface unit between the memory array and the computing circuit; The step of bonding the second semiconductor structure to the first semiconductor structure comprises: The first bonding layer and the second bonding layer are bonded.

25. The method according to claim 24, characterized in that The execution unit and the peripheral circuit are formed on the first substrate, comprising: The execution unit and the peripheral circuit are formed on the first substrate by adopting a synchronous process.

26. The method according to claim 24, characterized in that The memory array includes a plurality of memory cells, each of which includes a transistor and a memory device; The interface unit between the memory array and the computing circuit is formed on the second substrate, including: forming an interface unit between the transistors in the memory array and the computing circuit on the second substrate; Memory devices in the memory array are formed on the second substrate.

27. The method according to claim 24, characterized in that The first bonding layer includes a first contact structure, and the second bonding layer includes a second contact structure; The step of bonding the first bonding layer and the second bonding layer comprises: connecting a first portion of the first contact structure corresponding to the memory array to a first portion of the second contact structure corresponding to the peripheral circuit; The second portion of the first contact structure corresponding to the execution unit is connected to the second portion of the second contact structure corresponding to the interface unit of the computing circuit.

28. The method according to claim 23, wherein Providing a first semiconductor structure includes: forming the execution unit and the peripheral circuit on the first substrate; forming a first bonding layer on the execution unit and the peripheral circuit; Wherein, a second semiconductor structure is provided, comprising: forming an interface unit between the memory array and the computing circuit on a second substrate; forming a second bonding layer under the second substrate; The step of bonding the second semiconductor structure to the first semiconductor structure comprises: The first bonding layer and the second bonding layer are bonded.

29. The method according to claim 28, characterized in that The execution unit and the peripheral circuit are formed on the first substrate, comprising: forming the execution unit, the peripheral circuit, and a first interconnection region on the first substrate; forming a first interconnect layer on the execution unit and the peripheral circuit, and forming a first interconnect structure in the first interconnect region, wherein a first end of the first interconnect structure is connected to the execution unit and the peripheral circuit respectively through the first interconnect layer; Wherein, forming a first bonding layer on the execution unit and the peripheral circuit includes: The first bonding layer is formed on the execution unit and the peripheral circuit, wherein the first bonding layer includes a first contact structure, so that the second end of the first interconnection structure is connected to at least a portion of the first contact structure.

30. The method according to claim 28, wherein An interface unit between the memory array and the computing circuit is formed on a second substrate, comprising: forming an interface unit and a second interconnection region of the computing circuit on the second substrate; forming a second interconnect layer on the interface unit of the computing circuit and the second interconnect region; forming the memory array on the second interconnect layer; forming a second interconnect structure penetrating the second substrate in the second interconnect region, so that a first end of the second interconnect structure is connected to the memory array and the interface unit of the computing circuit respectively through the second interconnect layer; The step of forming a second bonding layer under the second substrate includes: The second bonding layer is formed under the second substrate, wherein the second bonding layer includes a second contact structure, so that the second end of the second interconnect structure is connected to at least a portion of the second contact structure.

31. A storage and calculation device, characterized in that: comprising a memory device and a computing circuit coupled to the memory device; The memory device includes a memory array and a peripheral circuit coupled to the memory array; The computing circuit includes an execution unit and an interface unit; The execution unit and the peripheral circuit are included in a first semiconductor structure; An interface unit between the memory array and the computing circuit is contained in a second semiconductor structure; The first semiconductor structure is bonded to the second semiconductor structure.

32. The device according to claim 31, characterized in that The first semiconductor structure further includes a first bonding layer and a first substrate; the execution unit and the peripheral circuit are located between the first bonding layer and the first substrate; The second semiconductor structure further includes a second bonding layer; the interface unit between the memory array and the computing circuit is located on a side of the second bonding layer away from the first semiconductor structure; The first semiconductor structure and the second semiconductor structure are bonded via the first bonding layer and the second bonding layer.

33. The device according to claim 32, characterized in that The first bonding layer includes a first contact structure, and the second bonding layer includes a second contact structure; The memory array is connected to the peripheral circuit via a first portion of the first contact structure and a corresponding first portion of the second contact structure; The execution unit is connected to the interface unit of the computing circuit via a second portion of the first contact structure and a corresponding second portion of the second contact structure.

34. The device according to claim 31, characterized in that The first semiconductor structure further includes a first bonding layer and a first substrate, and the execution unit and the peripheral circuit are located between the first bonding layer and the first substrate; The second semiconductor structure further includes a second substrate and a second bonding layer, the second substrate including a first side and a second side opposite to each other, the second bonding layer is located on the second side of the second substrate, and the interface unit between the memory array and the computing circuit is located on the first side of the second substrate; The first semiconductor structure and the second semiconductor structure are bonded via the first bonding layer and the second bonding layer.

35. The device according to claim 34, characterized in that The first bonding layer includes a first contact structure, and the second bonding layer includes a second contact structure; The first semiconductor structure further includes a first interconnect structure, and the second semiconductor structure further includes a second interconnect structure; A first end of the first interconnect structure is connected to the execution unit and the peripheral circuit respectively, and a second end is connected to at least part of the first contact structure; A first end of the second interconnect structure is connected to the memory array and the interface unit of the computing circuit respectively, and a second end is connected to at least a portion of the second contact structure.

36. The device according to claim 34 or 35, characterized in that The memory array is stacked on a first side of the interface unit of the computing circuit or on a second side opposite to the first side; The second semiconductor structure further includes a second interconnect layer located between the memory array and an interface unit of the computing circuit.

37. The device according to claim 31, characterized in that The second semiconductor structure further includes an interface unit of the peripheral circuit, and the interface unit of the computing circuit and the interface unit of the peripheral circuit are located in the same layer.

38. A storage and calculation system, characterized in that: comprising a memory device, a controller coupled to the memory device, and a computing circuit; The memory device includes a memory array and a peripheral circuit coupled to the memory array; The computing circuit includes an execution unit and an interface unit; The execution unit and the peripheral circuit are included in a first semiconductor structure; An interface unit between the memory array and the computing circuit is contained in a second semiconductor structure; The first semiconductor structure is bonded to the second semiconductor structure; The controller is configured to control the memory device to perform a storage operation and control the computing circuit to read data from the storage array to perform a computing operation.