Memory circuit and operating method thereof

By vertically stacking memory arrays in different physical layers, multiple arrays are simultaneously accessed to calculate MAC values, the IR voltage drop problem caused by the increase in the access line length in the existing in-memory computing system is solved, and system performance and computing efficiency are improved.

CN119993227APending Publication Date: 2025-05-13TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510071539.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-06-06
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing in-memory computing system only scales horizontally in the memory array, resulting in an increase in the access line length and an increase in the IR voltage drop, affecting system performance.

Method used

By forming multiple memory arrays in different physical layers and stacking them vertically, accessing multiple memory arrays simultaneously to calculate multiplication accumulation (MAC) values ​​is achieved.

Benefits of technology

Without increasing the IR voltage drop, the number of memory cells is significantly increased, and the computing efficiency and performance of the system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993227A_ABST
    Figure CN119993227A_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a memory circuit comprising a first memory array comprising a plurality of first memory cells, the plurality of first memory cells configured to store a first data element; a second memory array including a plurality of second memory cells configured to store a second data element; and a control circuit operatively coupled to both the first memory array and the second memory array, the control circuit configured to provide a multiplication accumulation (MAC) value based at least on simultaneously multiplying the third data element by the first data element and multiplying the third data element by the second data element. An embodiment of the invention provides a method of operating a memory circuit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention generally relate to the field of electronic circuits, and more particularly, to in-memory computing circuits and methods of operating the same. Background Art

[0002] The semiconductor industry has experienced rapid growth due to the increasing integration density of various electronic components such as transistors, diodes, resistors, capacitors, etc. In large part, the increase in integration density comes from the repeated reduction of the minimum feature size, which allows more components to be integrated into a given area. Summary of the invention

[0003] One embodiment of the present invention provides a memory circuit, comprising: a first memory array, including a plurality of first memory cells, the plurality of first memory cells being configured to store a first data element; a second memory array, vertically spaced from the first memory array, and including a plurality of second memory cells, the plurality of second memory cells being configured to store a second data element; and a control circuit, operatively coupled to both the first memory array and the second memory array, and configured to provide a multiply-accumulate (MAC) value based at least on simultaneously multiplying a third data element by the first data element and multiplying the third data element by the second data element.

[0004] Another embodiment of the present invention provides a memory circuit, comprising: a first memory array, formed in a first physical layer, and comprising a plurality of first storage cells, the plurality of first storage cells being configured to store a first data element; a second memory array, formed in a second physical layer, and comprising a plurality of second storage cells, the plurality of second storage cells being configured to store a second data element, wherein the first physical layer and the second physical layer are vertically spaced apart from each other; and a control circuit, operatively coupled to both the first memory array and the second memory array, wherein the control circuit is configured to: receive a third data element; and based on simultaneously accessing the first memory array and the second memory array, provide a multiply-accumulate (MAC) value of the third data element multiplied by each of the first data element and the second data element.

[0005] Yet another embodiment of the present invention provides a method of operating a memory circuit, comprising: forming a first memory array in a first physical layer; forming a second memory array in a second physical layer; and coupling the first physical layer to the second physical layer, the first physical layer and the second physical layer being vertically spaced apart from each other; wherein the first memory array and the second memory array are accessed simultaneously to retrieve a first data element from the first memory array and a second data element from the second memory array, and the retrieved first data element and the second data element are multiplied by a third data element. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Various aspects of the present invention will be best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be noted that, in accordance with standard practice in the industry, the various components are not drawn to scale. In fact, the sizes of the various components may be arbitrarily increased or reduced for clarity of discussion.

[0007] Figure 1 An exemplary neural network is shown in accordance with some embodiments.

[0008] Figure 2 A block diagram of a Computing-in-Memory (CiM) system is shown in accordance with some embodiments.

[0009] Figure 3 According to some embodiments Figure 2 Schematic diagram of the CiM system.

[0010] Figure 4 According to some embodiments Figure 2 Example implementation of parts of a CiM system.

[0011] Figure 5 It shows that according to some embodiments, Figure 4 Signal waveform of the implementation scheme.

[0012] Figure 6 According to some embodiments Figure 2 Another example implementation of a portion of a CiM system.

[0013] Figure 7 It shows that according to some embodiments, Figure 6 Signal waveform of the implementation scheme.

[0014] Figure 8 According to some embodiments Figure 2 Yet another example implementation of a portion of a CiM system.

[0015] Fig. 9 It shows that according to some embodiments, Figure 8 Signal waveform of the implementation scheme.

[0016] Fig.10 An example flow chart for operating a CiM system according to some embodiments is shown.

[0017] Fig.11 is a perspective diagram of a CiM system formed across multiple physical layers according to some embodiments.

[0018] Fig.12 A flow chart is shown of an example method of forming a memory system having different tiers in accordance with some embodiments.

[0019] Fig.13A , Fig. 13B , Fig. 13C , Fig.13D and Fig.13E According to some embodiments, Fig.12 Various cross-sectional views of devices formed by the method.

[0020] Fig.14 A flow chart is shown of another example method of forming a memory system having different tiers in accordance with some embodiments.

[0021] Fig.15A , Fig. 15B , Fig. 15C , Fig.15D and Fig.15E According to some embodiments, Fig.14 Various cross-sectional views of a memory device formed by the method.

[0022] Fig.16 A flow chart of yet another example method of forming a memory system having different tiers in accordance with some embodiments is shown.

[0023] Fig.17 A cross-sectional view of an example transistor formed in a metallization layer according to some embodiments is shown.

[0024] Fig.18 A cross-sectional view of another example transistor formed in a metallization layer in accordance with some embodiments is shown. DETAILED DESCRIPTION

[0025] The present invention provides many different embodiments or examples for realizing the different features of the present disclosure. Specific examples of components and arrangements are described below to simplify the present invention. Of course, these are merely examples and are not intended to limit the present invention. For example, in the following description, forming a first component above or on a second component may include an embodiment in which the first component and the second component are formed in direct contact, and may also include an embodiment in which an additional component may be formed between the first component and the second component so that the first component and the second component may not be in direct contact. In addition, the present invention may repeat reference numerals and / or characters in various examples. This repetition is for the purpose of simplicity and clarity, and does not itself indicate the relationship between the various embodiments and / or configurations discussed.

[0026] Furthermore, for ease of description, spatially relative terms such as "below," "beneath," "lower," "above," "upper," etc. may be used herein to describe the relationship of one element or component to another (or additional) elements or components as illustrated in the figures. The spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. The device may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein should likewise be interpreted accordingly.

[0027] With the advancement of modern semiconductor manufacturing processes and the increasing amount of data generated, there is an increasing need to store and process large amounts of data. Therefore, there is a drive to find improved methods for storing and processing large amounts of data. Although it is possible to process large amounts of data in software using traditional computer hardware, existing computer hardware is inefficient for certain data processing applications.

[0028] In this regard, machine learning has emerged as an effective way to analyze and extract value from large amounts of data. In general, machine learning is a field of computer science that involves algorithms that allow computers to "learn" (e.g., improve the performance of a task) without being explicitly programmed. Machine learning can use different techniques to analyze data to improve tasks. One such technique, such as deep learning, is based on neural networks. However, machine learning performed on traditional computer systems can involve excessive data transfer between memory and processors, resulting in high power consumption and slow computation.

[0029] Computing in memory (CiM), also known as processing in memory, involves performing computational operations in a memory array memory array. In other words, computational operations are performed directly on data read from memory cells, rather than transferring the data to a digital processor for processing. By avoiding the need to transfer some data to a digital processor, bandwidth limitations associated with transferring data back and forth between the processor and memory in a conventional computer system are reduced.

[0030] One application of CiM is artificial intelligence (AI), particularly machine learning. For example, a computing system (such as a CiM system) may use multiple layers of computing nodes, where lower layers perform computations based on the results of computations performed by higher layers. These computations may sometimes rely on the computation of dot products and absolute differences of vectors, typically computed using MACs (operations) performed on parameters (e.g., input data elements and weight data elements). The term "MAC" may refer to multiply-accumulate, multiply / accumulate, or multiply-accumulator, and typically refers to operations involving the multiplication of two values ​​and the accumulation of a sequence of multiplications.

[0031] In the prior art, the memory array of a CiM system is usually expanded only in the horizontal direction. Therefore, in order to include more and more memory cells in the memory array (e.g., with the trend of processing more and more data volumes), the access lines of the memory array can only be extended in the horizontal direction. The increase in the length of the access line increases the IR drop occurring on the access line, which has an adverse effect on the performance (e.g., speed) of the CiM system. In other words, there is usually a trade-off between the IR drop and the size of the memory array of the prior art CiM system. Therefore, the prior art CiM system is not completely satisfactory in some aspects.

[0032] The present invention provides various embodiments of a CiM system that is capable of efficiently outputting multiple MAC values ​​based on simultaneously accessing multiple memory arrays formed in different physical layers, respectively. In one aspect, a CiM system as disclosed herein may include multiple memory arrays formed in different physical layers, respectively. Such physical layers may be vertically spaced from each other, which allows the memory arrays of the disclosed CiM system to be stacked on top of each other. Using memory arrays formed in different physical layers, multiple data elements (or multiple bits of a first data element) programmed into the memory arrays, respectively, may be read out simultaneously. Thus, the CiM system may perform multiple MAC operations simultaneously to generate one or more MAC values. For example, when reading a first data element (e.g., a first weight data element) and a second data element (e.g., a second weight data element), the CiM system may simultaneously multiply a third data element (e.g., an input data element) by the first data element and the second data element. Thus, the number of memory cells of the disclosed CiM system may be significantly increased without increasing the IR drop (e.g., by vertically stacking memory arrays in different layers, respectively).

[0033] Figure 1An exemplary neural network 100 according to various embodiments is depicted. As shown, the inner layers of the neural network can be generally viewed as neuron layers, where each neuron layer receives weighted outputs from neurons of other (e.g., previous) neuron layers in a mesh interconnection structure between layers. The weight of the connection from the output of a particular previous neuron to the input of another subsequent neuron is set according to the influence or effect that the previous neuron will have on the subsequent neuron (for simplicity, only one neuron 101 and the weight of the input connection are labeled). Here, the output value of the previous neuron is multiplied by the weight of its connection with the subsequent neuron to determine the specific stimulus presented by the previous neuron to the subsequent neuron.

[0034] The total input stimulus of a neuron corresponds to the combined stimulus of all its weighted input connections. According to various implementations, if the total input stimulus of a neuron exceeds a certain threshold, the neuron is triggered to perform some (e.g.) linear or nonlinear mathematical function on its input stimulus. The output of the mathematical function corresponds to the output of the neuron, which is then multiplied by the corresponding weight of the output connection of the neuron to its subsequent neuron.

[0035] In general, the more connections between neurons, the more neurons per layer, and / or the more layers of neurons, the greater the intelligence that the network is able to achieve. Therefore, neural networks in practical, real-world artificial intelligence applications are usually characterized by a large number of neurons and a large number of connections between neurons. As a result, a large number of computations are required when processing information through a neural network (not only for the neuron output functions, but also for the weighted connections).

[0036] As described above, although neural networks can be implemented entirely in software as program code instructions executed on one or more conventional general-purpose central processing unit (CPU) or graphics processing unit (GPU) processing cores, the read / write activity required to perform all calculations between the CPU / GPU cores and system memory is extremely intensive. The overhead and energy generated by repeatedly moving large amounts of read data from system memory, processing that data by the CPU / GPU cores, and then writing the results back to system memory is not entirely satisfactory in some respects over the millions or billions of calculations required to implement a neural network.

[0037] Figure 2 2 shows a block diagram of an integrated circuit (eg, a CiM system) 200 according to various embodiments, which can effectively output multiple MAC values ​​by simultaneously accessing multiple memory arrays formed in different physical layers. It should be understood that Figure 2The CiM system 200 has been simplified for ease of illustration. Thus, the CiM system 200 may include any of a variety of other components while remaining within the scope of the present disclosure. For example, the CiM system 200 may include one or more precharge circuits configured to precharge the bit lines, and one or more decoder circuits configured to decode address signals.

[0038] As shown, the CiM system 200 includes memory arrays 210A, 210B, 210C, etc. and a control circuit 250. Figure 2 In the illustrated example of FIG. 2 , three memory arrays are shown, but it should be understood that the CiM system 200 may include any number of memory arrays while remaining within the scope of the present disclosure. The memory arrays 210A to 210C may be coupled to the control circuit 250 via a first global access line 262 and a second global access line 264. In various embodiments, the memory arrays 210A to 210C are respectively arranged in different physical layers.

[0039] For example, the memory array 210A may be formed in a first substrate, the memory array 210B may be formed in a second substrate, and the memory array 210C may be formed in a third substrate, wherein the first to third substrates may be vertically bonded to each other. In another example, the memory array 210A may be formed in a first metallization layer of a plurality of metallization layers arranged above the substrate, the memory array 210B may be formed in a second metallization layer of the plurality of metallization layers, and the memory array 210C may be formed in a third metallization layer of the plurality of metallization layers, wherein the first to third metallization layers may be vertically stacked on top of each other.

[0040] Each of the memory arrays 210A to 210C is a hardware component that is configured to store one or more data elements (e.g., one or more weight data elements). Using the memory array 210A as a representative example, the memory array 210A includes a plurality of memory cells (or other storage cells) 201. The memory array 210A includes a plurality of rows R1, R2, R3, ... R M , each row extends in a first direction (eg, X direction), and a plurality of columns C1, C2, C3 . . . C N , each column extends in a second direction (e.g., Y direction). Each row and each column may include one or more conductive (e.g., metal) structures used as local access lines, such as local bit lines (BL), local WLs (WLs), local source / select lines (SL), etc. Each memory cell 201 is arranged at the intersection of a corresponding row and a corresponding column, and may be operated according to a voltage or current through a corresponding local access line of the column and row. For example, each row may include one or more corresponding local WLs, and each column may include one or more corresponding local BLs.

[0041] Each of the memory cells 201 may be implemented in any of a variety of configurations, such as a one-transistor-one-resistor (1T1R) configuration, a one-selector / switch-one-resistor (1S1R) configuration, a one-diode-one-resistor (1D1R) configuration, a single-transistor (1T) configuration, and the like. For example, the memory cell 201 in the 1T1R configuration may include a transistor implemented as a MOSFET or a MESFET and a resistor implemented as a variable resistor, a magnetoresistive stack, a phase change stack, and the like. In another example, the memory cell 201 in the 1S1R configuration may include a selector implemented as a bipolar metal insulator metal structure and a resistor implemented as a variable resistor, a magnetoresistive stack, a phase change stack, and the like. In yet another example, the memory cell 201 in the 1T configuration may include a FeFET with an adjustable threshold voltage, a floating gate flash memory structure, a SONOS memory structure, and the like. In some other embodiments, the memory cells 201 may each be implemented as a 6-transistor (6T) static random access memory (SRAM) cell, an 8-transistor (8T) SRAM cell, or a 10-transistor (10T) SRAM cell. In some other embodiments, memory cells 201 may each represent memory cells based on dynamic random access memory (DRAM) technology.

[0042] In various embodiments, each of the memory arrays 210A to 210C may include at least a row control circuit (e.g., 212A, 212B, 212C) and a column control circuit (e.g., 214A, 214B, 214C). The row control circuit 212A may selectively couple the first global access line 262 to one or more local access lines arranged along the row. For example, the row control circuit 212A may include a plurality of first switches selectively connected to the local access lines (e.g., WL) arranged along the row. In some embodiments, the control circuit 250 may determine whether to activate the first switch of each memory array individually or to activate the first switches of all memory arrays simultaneously based on the operation mode of the CiM system 200. The column control circuit 214A may selectively couple the second global access line 264 to one or more local access lines arranged along the column. For example, the column control circuit 214A may include a plurality of second switches selectively connected to the local access lines (e.g., BLs) arranged along the column. In some embodiments, the control circuit 250 may determine whether to activate the second switch of each memory array individually or to activate the second switches of all memory arrays simultaneously based on the operation mode of the CiM system 200 .

[0043] In addition, since the global access lines 262 and 264 are configured to be accessible to all memory arrays (e.g., 210A to 210C) arranged vertically on top of each other, the global access lines 262 and 264 may extend in a direction perpendicular to the longitudinal or extending direction of the local access lines. For example, the local access lines (WL, BL) of each of the memory arrays 210A to 210C may extend in the X direction and / or the Y direction (or lateral direction), while the global access lines 262 to 264 may extend in the Z direction (or vertical direction). In some embodiments, the memory arrays 210A to 210C may each have their local access lines extending throughout the array in the X direction or the Y direction. Therefore, in a specific operation mode (e.g., generating a MAC value), multiple memory arrays may be accessed simultaneously through the global access lines.

[0044] The control circuit 250 may include an analog processor and an analog-to-digital converter (ADC). In some embodiments, the analog processor of the control circuit 250 may receive at least two inputs (e.g., a weight data element and an input data element) and perform one or more calculations on the inputs. The calculation may be a matrix multiplication, an absolute difference calculation, a dot product multiplication, or other machine learning (ML) operations. In this way, the analog processor of the control circuit 250 may provide one or more MAC values ​​for the received inputs, and the ADC of the control circuit 250 may convert the MAC values ​​into digital bits.

[0045] In one aspect of the present invention, the control circuit 250 may simultaneously receive (e.g., retrieve) different weight data elements (or different bits of weight data elements) from the memory arrays 210A to 210C, respectively, and receive input data elements to calculate one or more MAC values. The control circuit 250 may retrieve weight data elements from different memory arrays through its respective row control circuits (e.g., 212A) and column control circuits (e.g., 214A). For example, the control circuit 250 may multiply the input data element by the first weight data element retrieved from the first memory array in the memory arrays 210A to 210C to generate a first product, and multiply the same input data element by the second weight data element retrieved from the second memory array in the memory arrays 210A to 210C to generate a second product. The control circuit 250 may retrieve weight data elements from each memory array based on a voltage sensing technique (e.g., detecting the voltage level presented on the BL). The control circuit 250 may then add the first product and the second product, for example, as a partial product.

[0046] As a non-limiting example, prior to a read operation, local access lines (e.g., BL) along columns throughout all memory arrays 210A to 210C may be precharged to a supply voltage (e.g., VDD). When control circuit 250 activates multiple rows (e.g., WL) of memory array 210A, at least one of the BLs of memory array 210A may be discharged to a voltage proportional to the value stored in the activated WL. Weighting each row by bit position may result in a column voltage drop (ΔV BL , or the increment / change in bit line voltage), the voltage drop is proportional to the binary stored weight data element. For a 4-bit weight data element W1, W1[3] and W1[0] can represent the MSB and LSB of W1, respectively. The 4-bit W1, W1[3], W1[2], W1[1] and W1[0] are sequentially stored in the 4th, 3rd, 2nd and 1st WL of the memory array 210A, respectively. The voltage drop of BL is proportional to {W0+2×W1+4×W2+8×W3}. The control circuit 250 can multiply the input data element XIN by such a read weight data element W1 to generate a first product. At the same time, the control circuit 250 can read out another weight data element W2 from the memory array 210B based on the same principle, and multiply the input data element XIN by the read weight data element W2 to generate a second product. Then, the control circuit 250 can add the first product and the second product.

[0047] In another aspect of the present invention, the control circuit 250 may simultaneously apply different input data elements (or different bits of the input data elements) to the memory arrays 210A to 210C, respectively, and simultaneously multiply the input data elements by different weight data elements stored in the corresponding memory arrays to calculate one or more MAC values. The control circuit 250 may apply the input data elements to different memory arrays through their respective row control circuits (e.g., 212A), and retrieve the products from the corresponding memory arrays through their column control circuits (e.g., 214A). For example, the control circuit 250 may multiply the first input data element applied to the first memory array of the memory arrays 210A to 210C by the first weight data element stored in the first memory array to generate a first product, and multiply the second input data element applied to the second memory array of the memory arrays 210A to 210C by the second weight data element stored in the second memory array to generate a second product. The control circuit 250 may retrieve the products from the respective memory arrays based on current sensing techniques (e.g., detecting the current level presented on the BL). Control circuit 250 may then add the first product and the second product, eg, as a partial product.

[0048] As another non-limiting example, a first weight data element W1 including a corresponding number of bits may be stored in one or more storage cells coupled to a first local access line (e.g., BL1) along each column of the memory array 210A, and a second weight data element W2 including a corresponding number of bits may be stored in one or more storage cells coupled to a second local access line (e.g., BL2) along each column of the memory array 210B. Upon receiving the input data element XIN, the control circuit 250 may activate one or more local access lines along the corresponding row of the memory array 210A and one or more local access lines along the corresponding row of the memory array 210B. The control circuit 250 may simultaneously determine a first product (of the input data element XIN and the first weight data element W1) based on a first current level detected through BL1, and determine a second product (of the input data element XIN and the second weight data element W2) based on a second current level detected through BL2. Then, the control circuit 250 may add the first product and the second product.

[0049] Figure 3 FIG. 2 is a schematic diagram showing a portion of a CiM system 200 including two memory arrays according to various embodiments of the present invention. Figure 3 In the illustrated example, memory array 210A and memory array 210B are shown. In various embodiments, memory array 210A and memory array 210B are respectively formed in different substrates bonded to each other, or are respectively formed in different (eg, metallization) layers stacked on the same substrate.

[0050] As shown in the figure, the memory array 210A includes a plurality of memory cells 301A arranged along four local access lines (eg, WL A0 , WL A1 , WL A2 , WL A3 ) and four local access lines extending in the Y direction (e.g. BL A0 BL A1 BL A2 BL A3 ). Each memory cell 301A is coupled to (eg, through) a corresponding one of WL0 to WL4 and a corresponding one of BL0 to BL4. Similarly, the memory array 210B includes four local access lines (eg, WL B0 , WL B1 , WL B2 , WL B3 ) and four local access lines extending in the Y direction (e.g., B LB0 BL B1 BL B2 BLB3 ) on a plurality of memory cells 301B. Each memory cell 301B is coupled to (eg, via) WL B0 To WL B4 The corresponding one and BL B0 To BL B4 The corresponding one in .

[0051] In addition, the row control circuit 212A of the memory array 210A includes a plurality of row control circuits 212A and 212B coupled to the WL A0 , WL A1 , WL A2 , WL A3 Multiple switches 312 A0 ,312 A1 ,312 A2 and 312 A3 , and the column control circuit 214A of the memory array 210A includes a column control circuit 214A coupled to BL A0 BL A1 BL A2 BL A3 Multiple switches 314 A0 ,314 A1 ,314 A2 and 314 A3 Similarly, the row control circuit 212B of the memory array 210B includes the row control circuits 212B and 212B coupled to the WL B0 , WL B1 , WL B2 , WL B3 Multiple switches 312 B0 ,312 B1 ,312 B2 and 312 B3 , and the column control circuit 214B of the memory array 210B includes a column control circuit 214B coupled to BL B0 BL B1 BL B2 BL B3 Multiple switches 314 B0 ,314 B1 ,314 B2 and 314 B3 Switches 312 and 314 may each include a transmission gate, an n-type metal oxide semiconductor (NMOS) transistor, a p-type metal oxide semiconductor (PMOS) transistor, or the like.

[0052] In various embodiments, each of the switches 312 and 314 is configured to selectively couple the corresponding local access line to the global access line based on a control signal. Such a control signal may be provided by the control circuit 250 according to at least one of an operation mode of the corresponding memory array, an address of an activated memory cell to be programmed or read, or a conductivity type of the switch. In addition, in various embodiments of the present invention, the global access lines may each have at least a portion extending in a direction perpendicular to the extension direction of the local access lines (e.g., implemented as a through-hole structure, a through-silicon through-hole structure, a through-substrate through-hole structure, etc.). In this way, such a global access line can be shared by memory arrays respectively formed in different layers or substrates.

[0053] For example, in memory array 210A, switch 312 A0 You can optionally A0 Coupled to global WL 3620; switch 312 A1 You can optionally A1 Coupled to global WL 3621; switch 312 A2 Optionally, WL A2 Coupled to global WL3622; switch 312 A3 You can optionally A3 Coupled to global WL 3623; switch 314 A0 Optionally, BL A0 Coupled to global BL 3640; switch 314 A1 BL can be selected A1 Coupled to global BL 3641; switch 314 A2 Optionally, BL A2 Coupled to global BL 3642; switch 314 A3 BL can be selected A3 is coupled to the global BL 3643. Similarly, in the memory array 210B, the switch 312 B0 You can optionally B0 Coupled to global WL3620; switch 312 B1 Optionally, WL B1 Coupled to global WL 3621; switch 312 B2 Optionally, WL B2 Coupled to global WL 3622; switch 312 B3 Optionally, WL B3 Coupled to global WL 3623; switch 314 B0 Optionally, BL B0 Coupled to global BL 3640; switch 314 B1BL can be selected B1 Coupled to global BL 3641; switch 314 B2 Optionally, BL B2 Coupled to global BL 3642; switch 314 B3 BL can be selected B3 Coupled to global BL 3643.

[0054] Figure 4 According to some embodiments Figure 3 An example implementation of a portion of the memory array 210A and a portion of the memory array 210B of the schematic diagram shown. As shown, a connection to WL A3 and BL A0 Memory cell 301A (of memory array 210A) and connected to WL B3 and BL B0 Specifically, WL (of memory arrays 210A and 210B) A3 and WL B3 The switches 312 A3 and 312 B3 Commonly coupled to global WL 3623; and BL (of memory arrays 210A and 210B) A0 and BL B0 The switches 314 can be used to A0 and 314 B0 Commonly coupled to global BL3640. Switch 312 A3 ,312 B3 ,314 A0 and 314 B0 Each can be implemented as a NOMS transistor.

[0055] exist Figure 4 In the implementation of the switch 312 A3 ,312 B3 ,314 A0 and 314 B0 The control lines 401, 403, 405, and 407 may be controlled (e.g., activated or deactivated) by control signals 401, 403, 405, and 407, respectively. Such control signals 401 to 407 may be provided by control circuit 250 to memory arrays 210A to 210B via corresponding control lines. Each of these control lines coupled to control circuit 250 may be arranged in parallel with a corresponding controlled local access line. For example, the control line conducting control signal 401 may be connected to WL A3 The control line conducting the control signal 405 can be connected to BL A0According to various embodiments of the present invention, during a programming operation performed on memory cells 301A-301B, switch 312 may be activated. A3 and switch 312 B3 and during a computation (eg, MAC) operation performed on storage units 301A to 301B, switch 312 A3 and switch 312 B3 Both can be activated.

[0056] Figure 5 Example waveforms of control signals 401, 403, 405, and 407 during various operating modes of memory cells 301A to 301B (or more generally, operating modes of memory arrays 210A to 210B), respectively, are shown according to some embodiments. As shown, during a first time period 510 when one or more memory cells 301A of the memory array 210A are configured to be programmed (e.g., one or more weight data elements are written to these memory cells 301A), control signals 401 and 405 are pulled up to a logic high state, while control signals 403 and 407 remain in (or transition to) a logic low state. During a second time period 520 when one or more memory cells 301B of the memory array 210B are configured to be programmed (e.g., one or more weight data elements are written to these memory cells 301B), control signals 403 and 407 are pulled up to a logic high state, while control signals 401 and 405 remain in (or transition to) a logic low state. During a third time period 530 in which the programmed data elements are configured to be read (eg, for performing a MAC operation), all of the control signals 401 to 407 are pulled up to a logic high state.

[0057] Thus, during time period 510, control circuit 250 may write a first data element to one or more memory cells 301A of memory array 210A; during time period 520, control circuit 250 may write a second data element to one or more memory cells 301B of memory array 210B; and during time period 530, control circuit 250 may multiply a third data element by both the first data element (retrieved from memory array 210A) and the second data element (retrieved from memory array 210B). In some embodiments, time period 530 may occur after either time period 510 or time period 520, but the order of time period 510 and time period 520 may be arbitrarily changed.

[0058] Figure 6 According to some embodiments Figure 3 Another example implementation of a portion of the memory array 210A and a portion of the memory array 210B of the schematic diagram shown. As shown, a connection to WLA3 and BL A0 Memory cell 301A (of memory array 210A) and connected to WL B3 and BL B0 Specifically, WL (of memory arrays 210A and 210B) A3 and WL B3 The switches 312 A3 and 312 B3 Commonly coupled to global WL3623; and BL (of memory arrays 210A and 210B) A0 and BL B0 The switches 314 can be used to A0 and 314 B0 Commonly coupled to global BL 3640. Switch 312 A3 ,312 B3 ,314 A0 and 314 B0 Each can be implemented as a POMS transistor.

[0059] exist Figure 6 In the implementation of the switch 312 A3 ,312 B3 ,314 A0 and 314 B0 The control lines 601, 603, 605, and 607 may be controlled (e.g., activated or deactivated) by control signals 601, 603, 605, and 607, respectively. Such control signals 601 to 607 may be provided to the memory arrays 210A to 210B by the control circuit 250 via corresponding control lines. Each of these control lines coupled to the control circuit 250 may be arranged in parallel with a corresponding controlled local access line. For example, the control line conducting the control signal 601 may be connected to the WL A3 The control line conducting the control signal 605 can be connected to BL A0 According to various embodiments of the present invention, during a programming operation performed on memory cells 301A-301B, one of switches 312A3 and 312B3 may be activated; and during a computation (e.g., MAC) operation performed on memory cells 301A-301B, both switches 312A3 and 312B3 may be activated.

[0060] Figure 7Example waveforms of control signals 601, 603, 605, and 607 during various operating modes of memory cells 301A to 301B (or more generally, operating modes of memory arrays 210A to 210B), respectively, are shown according to some embodiments. As shown, during a first time period 710 when one or more memory cells 301A of the memory array 210A are configured to be programmed (e.g., one or more weight data elements are written to these memory cells 301A), control signals 601 and 605 are pulled down to a logic low state, while control signals 603 and 607 remain in (or transition to) a logic high state. During a second time period 720 when one or more memory cells 301B of the memory array 210B are configured to be programmed (e.g., one or more weight data elements are written to these memory cells 301B), control signals 603 and 607 are pulled down to a logic low state, while control signals 601 and 605 remain in (or transition to) a logic high state. During a third time period 730 in which the programmed data elements are configured to be read out (eg, for performing a MAC operation), all control signals 601 to 607 are pulled down to a logic low state.

[0061] Thus, during time period 710, control circuit 250 may write a first data element to one or more memory cells 301A of memory array 210A; during time period 720, control circuit 250 may write a second data element to one or more memory cells 301B of memory array 210B; and during time period 730, control circuit 250 may multiply a third data element by both the first data element (retrieved from memory array 210A) and the second data element (retrieved from memory array 210B). In some embodiments, time period 730 may occur after either time period 710 or time period 720, but the order of time period 710 and time period 720 may be arbitrarily changed.

[0062] Figure 8 According to some embodiments Figure 3 Another example implementation of a portion of the memory array 210A and a portion of the memory array 210B of the schematic diagram shown. As shown, a portion of the memory array 210A and a portion of the memory array 210B connected to the WL A 3 and BL A0 Memory cell 301A (of memory array 210A) and connected to WL B3 and BL B0 Specifically, WL (of memory arrays 210A and 210B) A3 and WL B3 The switches 312 A3 and 312 B3Commonly coupled to global WL3623; and BL (of memory arrays 210A and 210B) A0 and BL B0 The switches 314 can be used to A0 and 314 B0 They are commonly coupled to the global BL 3640. The switches 312A3 and 312B3 may be implemented as an NMOS transistor and a PMOS transistor, respectively; the switches 314A0 and 314B0 may be implemented as a PMOS transistor and an NMOS transistor, respectively.

[0063] exist Fig. 9 In the implementation of the switch 312 A3 ,312 B3 ,314 A0 and 314 B0 The control lines 801, 803, 805, and 807 may be controlled (e.g., activated or deactivated) by control signals 801, 803, 805, and 807, respectively. Such control signals 801 to 807 may be provided to the memory arrays 210A to 210B by the control circuit 250 via corresponding control lines. Each of these control lines coupled to the control circuit 250 may be arranged in parallel with a corresponding controlled local access line. For example, the control line conducting the control signal 801 may be connected to the WL A3 The control line conducting the control signal 805 can be connected to BL A0 According to various embodiments of the present invention, during a programming operation performed on memory cells 301A-301B, one of switches 312A3 and 312B3 may be activated; and during a computation (e.g., MAC) operation performed on memory cells 301A-301B, both switches 312A3 and 312B3 may be activated.

[0064] Fig. 9Example waveforms of control signals 801, 803, 805, and 807 during various operating modes of memory cells 301A-301B (or more generally, operating modes of memory arrays 210A-210B), respectively, are shown in accordance with some embodiments. As shown, during a first time period 910 in which one or more memory cells 301A of the memory array 210A are configured to be programmed (e.g., one or more weight data elements are written to these memory cells 301A), control signals 801 and 805 are pulled up to a logic high state and pulled down to a logic low state, respectively, while control signals 803 and 807 remain in (or transition to) a logic high state and a logic low state, respectively. During a second time period 920 in which one or more memory cells 301B of the memory array 210B are configured to be programmed (e.g., one or more weight data elements are written to these memory cells 301B), the control signals 803 and 807 are pulled down to a logic low state and pulled up to a high logic state, respectively, while the control signals 801 and 805 remain (or switch to) a logic high state and a logic low state, respectively. During a third time period 930 in which the programmed data elements are configured to be read out (e.g., for performing a MAC operation), the control signals 801, 803, 805, and 807 are set to a logic high state, a logic low state, a logic low state, and a logic high state, respectively.

[0065] Thus, during time period 910, control circuit 250 may write a first data element to one or more memory cells 301A of memory array 210A; during time period 920, control circuit 250 may write a second data element to one or more memory cells 301B of memory array 210B; and during time period 930, control circuit 250 may multiply a third data element by both the first data element (retrieved from memory array 210A) and the second data element (retrieved from memory array 210B). In some embodiments, time period 930 may occur after either time period 910 or time period 920, but the order of time period 910 and time period 920 may be arbitrarily changed.

[0066] Fig.10 1 is a flowchart of an example method 1000 for operating a CiM system including a plurality of memory arrays stacked on top of each other according to various embodiments of the present disclosure. The operations of the method 1000 may be performed by the above components (e.g., Figures 2 to 9 ) is performed, and therefore, some of the reference numerals used above may be reused in the following discussion of method 1000. In addition, it should be understood that method 1000 has been simplified, and therefore, may be described in detail in detail. Fig.10 Additional operations are provided before, during, and after method 1000, and some other operations are only briefly described here.

[0067] Method 1000 begins with operation 1010, where a first data element is programmed into a first memory array. In some embodiments, the first data element may correspond to a first weight data element or a first bit subset of a weight data element. The first memory array may include a plurality of first memory cells formed in a first physical layer or a first memory array layer, which will be in the following Figure 11 to Figure 1 5. To program a first data element to a first memory array (e.g., 210A), a control circuit (e.g., 250) may activate row control circuits (e.g., 212A) and column control circuits (e.g., 214A) of the first memory array to couple one or more of its local access lines (e.g., BL, WL) to corresponding global access lines (e.g., 262, 264), which allows the control circuit 250 to write the first data element to the first memory array 210A through its row control circuits 212A and column control circuits 214A.

[0068] The method 1000 continues with operation 1020 of programming a second data element into a second memory array. In some embodiments, the second data element may correspond to a second weight data element or a second bit subset of the weight data element. The second memory array may include a plurality of second memory cells formed in a second physical layer or a second memory array layer, which will be in the following Figure 11 to Figure 1 5. To program the second data element into the second memory array (e.g., 210B), the control circuit 250 may activate the row control circuit (e.g., 212B) and the column control circuit (e.g., 214B) of the second memory array to couple one or more of its local access lines (e.g., BL, WL) to the corresponding global access lines (e.g., 262, 264), which allows the control circuit 250 to write the second data element into the second memory array 210B through its row control circuit 212B and column control circuit 214B. In various embodiments of the present invention, the control circuit 250 may program the data elements into the corresponding memory arrays separately. For example, the first and second data elements may be written to the corresponding memory arrays at different time periods.

[0069] The method 1000 continues with operation 1030 by simultaneously accessing the first memory array and the second memory array to retrieve the first data element and the second data element. Continuing with the above example, after writing the first data element and the second data element to the respective memory arrays 210A and 210B, the control circuit 250 may simultaneously access the memory arrays 210A and 210B to retrieve the first data element and the second data element, respectively. For example, the control circuit 250 may activate the row control circuit 212A and the column control circuit 214A to couple the global access lines 262 and 264 to the memory array 210A; at the same time, the control circuit 250 may activate the row control circuit 212B and the column control circuit 214B to couple the global access lines 262 and 264 to the memory array 210B.

[0070] The method 1000 continues with operation 1040, multiplying the third data element by the first data element, and multiplying the third data element by the second data element. Continuing with the above example, the control circuit 250 may multiply the third data element (e.g., the input data element) by the retrieved first data element, and multiply the third data element by the retrieved second data element. In one aspect, the control circuit 250 may receive the third data element and perform the multiplication (e.g., within the control circuit 250). In another aspect, the control circuit 250 may receive the third data element, apply the third data element to both the first and second memory arrays, and perform the multiplication (e.g., within the respective memory arrays). After performing the multiplication, the control circuit 250 may sum the first product (of the first and third data elements) and the second product (of the second and third data elements).

[0071] Fig.11 A perspective view of an exemplary memory system 1100 is shown, according to various embodiments, the exemplary memory system 1100 including a plurality of first interconnect structures 1112 and a plurality of second interconnect structures 1114, which are configured to operably couple a physical layer to one or more other physical layers that are integrated (e.g., stacked) with each other in the Z direction. The memory system 1100 may include components substantially similar to the memory systems discussed above, such as Figure 2 200. It should be understood that Fig.11 The configuration of the memory system 1100 shown in has been simplified for illustrative purposes, and thus, the memory system 1100 may include any of a variety of other layers and / or have a different configuration (e.g., different layers coupled to each other, depending on the desired design, etc.) while remaining within the scope of the present disclosure.

[0072] As shown, the memory system 1100 includes at least one peripheral layer 1102, and a plurality of memory array layers 1104 and 1106 arranged above the peripheral layer 1102. According to some embodiments of the present invention, the peripheral layer 1102 and the memory array layers 1104 to 1106 may be formed in different substrates (e.g., wafers), respectively. In this way, the peripheral layer 1102 may be operatively coupled to one or more of the memory array layers 1104 to 1106 through one or more of the first interconnect structure 1112 and the second interconnect structure 1114, wherein at least a portion of the first interconnect structure 1112 and the second interconnect structure 1114 may be implemented as a through silicon / substrate through via (TSV) structure. Alternatively, each of the first interconnect structure 1112 and the second interconnect structure 1114 may be selectively coupled to one or more of the memory array layers 1104 to 1106 based on its configured operation mode. According to some other embodiments of the present invention, the peripheral layer 1102 and the memory array layers 1104 to 1106 may be formed in different layers, but are arranged on the same substrate (e.g., wafer). In this way, the peripheral layer 1102 may be operatively coupled to one or more of the memory array layers 1104 to 1106 through one or more of the first interconnect structure 1112 and the second interconnect structure 1114, wherein at least a portion of the first interconnect structure 1112 and the second interconnect structure 1114 may be implemented as a through-hole structure.

[0073] For example, the peripheral layer 1102 may include a plurality of components operatively used as a control circuit (e.g., 250), which may include an input / output (I / O) circuit, a logic control circuit, a command register, an address register, and a sequencer, while the memory array layers 1104 to 1106 may each include a memory array (e.g., 210A to 210C) and a corresponding row / column control circuit (e.g., 212A to 212C, 214A to 214C). Typically, the input / output circuit is configured to communicate various input / output signals (e.g., a command (CMD) signal, an address information (ADD) signal, and a data (DAT) signal) with a memory controller, which may also be formed in the peripheral layer 1102. When receiving an input / output signal from the memory controller, the input / output circuit may distribute the input / output signal to a CMD signal, an ADD signal, and a DAT signal based on information received from the logic control circuit, which may also be formed in the peripheral layer 1102. The input / output circuit provides the CMD signal and the ADD signal to the command register and the address register, respectively. In addition, the input / output circuit communicates the DAT signal with the sense amplifier. The logic control circuit is configured to receive CLE, ALE, Wen and Ren signals from the memory controller. The logic control circuit can send the above information to the I / O circuit for identifying the CMD signal, ADD signal and DAT signal in the I / O signal. In addition, the logic control circuit provides the RBn signal to the memory controller to notify the state of the corresponding memory device. The command register is configured to store the CMD received from the input / output circuit. CMD includes, for example, instructions for causing the sorter to perform a read operation, a write operation, an erase operation, etc. The address register is configured to store address information ADD received from the input / output circuit. ADD includes, for example, at least a row address (RAd) and a column address (CAd). The row address RAd and the column address CAd can be used to select a word line and a bit line, respectively. The sorter is configured to control the operation of the entire memory device. For example, the sorter can control the row control circuit (for example, based on the CMD stored in the command register) Figure 2 212A to 212C shown in FIG. 2 ), column control circuits (eg, Figure 2 214A to 214C) etc. shown in the figure, and performs a read operation, a write operation, an erase operation, a calculation operation, etc.

[0074] Fig.12 A flow chart of a method 1200 for forming a memory system including different layers operably coupled to each other through TSVs according to one or more embodiments of the present invention is shown. For example, at least some operations (or steps) of the method 1200 may be used to form the memory system discussed above. Note that the method 1200 is merely an example and is not intended to limit the present disclosure. Therefore, it should be understood that the method 1200 may be used to form a memory system including different layers operably coupled to each other through TSVs according to one or more embodiments of the present invention ... Fig.12Additional operations are provided before, during, and / or after method 1200, and only some other operations are briefly described here. In some embodiments, the operations of method 1200 may be different from those of Fig.13A , Fig. 13B , Fig. 13C , Fig.13D and Fig.13E The illustrated cross-sectional views of an example semiconductor device at various stages of fabrication are associated with the invention, which are discussed in further detail below.

[0075] Corresponds to Fig.12 Operation 1202, Fig.13A A cross-sectional view of a portion of a semiconductor device 1300 including a first substrate (or chip) 1302 is shown with a plurality of TSVs 1304 formed over a front surface of the first substrate 1302 in one of a plurality of manufacturing stages according to various embodiments.

[0076] The first substrate 1302 may be a doped (e.g., using a p-type or n-type dopant) or undoped semiconductor substrate, such as a bulk semiconductor, a semiconductor on insulator (SOI) substrate, etc. The first substrate 1302 may be a wafer, such as a silicon wafer. Typically, an SOI substrate includes a semiconductor material layer formed on an insulator layer. The insulator layer may be, for example, a buried oxide (BOX) layer, a silicon oxide layer, etc. The insulator layer is disposed on a substrate, typically a silicon or glass substrate. Other substrates, such as multilayer or gradient substrates, may also be used. In some embodiments, the semiconductor material of the substrate 1302 may include silicon; germanium; compound semiconductors, including silicon carbide, gallium arsenide, gallium phosphide, indium phosphide, indium arsenide, and / or indium antimonide; alloy semiconductors, including SiGe, GaAsP, AllnAs, AlGaAs, GainAs, GainP, and / or GainAsP; or combinations thereof.

[0077] TSV 1304 is formed of a conductive material. The conductive material may include copper, although other suitable materials such as aluminum, alloys, doped polysilicon, combinations thereof, etc. may be used instead. At this stage of manufacturing, TSV 1304 may not extend completely through first substrate 1302, i.e., not extend from the front surface of first substrate 1302 to the back surface. TSV 1304 may be formed by performing at least some of the following processes: forming an opening through the front surface of first substrate 1302; lining the opening with a pad barrier layer (not shown); filling the opening with the conductive material described above; and polishing first substrate 1302. Although not shown, it should be noted that the same process for forming TSV 1304 (and the like) may be performed simultaneously on a second substrate (chip) of semiconductor device 1300. Fig.12 The following operations except operation 1210).

[0078] Corresponds to Fig.12 Operation 1204, Fig. 13B A cross-sectional view of a portion of a semiconductor device 1300 at one of different stages of fabrication according to various embodiments is shown, the semiconductor device 1300 including a plurality of components 1306 , 1308 , and 1310 formed on a front surface of a substrate 1302 .

[0079] exist Fig. 13B In the illustrated example of (and the following figure), component 1306 can represent many devices, such as transistors, memory cells, etc.; component 1308 can represent multiple through-hole structures electrically coupled to TSV 1304 (and component 1306), respectively; and component 1310 can represent multiple interconnect structures electrically coupled to through-hole structures 1308, respectively. Such components 1306 to 1310 can be covered by a dielectric layer 1312, which is generally referred to as an interlayer dielectric (ILD) or a metal dielectric (IMD). When forming such components, according to some embodiments, one of the above-mentioned memory array layers or peripheral layers may have been formed. For example, for the memory array layer, component 1306 can represent: (i) a plurality of memory cells commonly used as one or more memory arrays (e.g., one of 210A to 210C); and (ii) a plurality of transistors commonly used as one or more corresponding basic circuits (e.g., one of 212A to 212C, and one of 214A to 214C). Components 1308 and 1310 may represent: (i) a plurality of local access lines (eg, bit lines, word lines, etc.) of a memory array; and (ii) a plurality of interconnect structures coupled to the memory array.

[0080] Corresponds to Fig.12 Operation 1206, Fig. 13C A cross-sectional view of a portion of a semiconductor device 1300 at one stage at different stages of fabrication according to various embodiments is shown, wherein a first substrate 1302 is thinned from its back side. As shown, the first substrate 1302 is thinned from its back side until the bottom surface of the TSV 1304 is exposed. In some embodiments, the first substrate 1302 can be thinned using a polishing process (e.g., a chemical mechanical polishing (CMP) process) while having its front surface coupled to a carrier wafer 1316.

[0081] Corresponds to Fig.12 Operation 1208, Fig.13DA cross-sectional view of a portion of a semiconductor device 1300 at one stage at different stages of fabrication according to various embodiments is shown, the semiconductor device 1300 including a plurality of pads 1320 respectively coupled to TSVs 1304. After the bottom surfaces of the TSVs 1304 are exposed, the pads 1320 are formed to electrically couple to the TSVs 1304, thereby allowing the TSVs 1304 to be electrically coupled to other components, as described below. The bonding pads 1320 are formed of a conductive material. The conductive material may include copper, although other suitable materials such as aluminum, alloys, doped polysilicon, combinations thereof, etc. may be used alternatively.

[0082] Corresponds to Fig.12 Operation 1210, Fig.13E A cross-sectional view of a portion of a semiconductor device 1300 is shown at one of various manufacturing stages according to various embodiments, the semiconductor device 1300 including a first layer and a second layer bonded to each other. As described above, operations 1202 to 1208 may be performed simultaneously on a second substrate (chip), which results in the formation of similar layers. Fig.13E As shown, after forming the bonding pad 1320, the first layer (which may be one of the above-mentioned memory array layer or peripheral layer) is bonded to the second layer (which may be one of the above-mentioned memory array layer or peripheral layer). Similar to the first layer, the second layer includes a thinned substrate 1322, one or more TSVs 1324 extending through the thinned substrate 1322, components 1326, 1328 and 1330, ILD / IMD 1332 and one or more pads 1330. Fig.13E In the illustrated example of , the first layer is joined (eg, operatively coupled) to the second layer via TSV 1304. It should be understood that each of the first layer and the second layer may be coupled to one or more other layers via their respective TSVs to form one of the memory systems described above.

[0083] Fig.14 A flow chart of a method 1400 for forming a memory system including different layers operably coupled to each other through TSVs according to one or more embodiments of the present invention is shown. For example, at least some operations (or steps) of the method 1400 may be used to form the memory system discussed above. Note that the method 1400 is merely an example and is not intended to limit the present disclosure. Therefore, it should be understood that the method 1400 may be used in the following embodiments. Fig.14 Additional operations are provided before, during, and / or after method 1400, and only some other operations are briefly described here. In some embodiments, the operations of method 1400 may be different from those of Fig.15A , Fig. 15B , Fig. 15C , Fig.15D and Fig.15EFIG. 1 is an exploded view of a semiconductor device at various stages of fabrication shown in FIG. 1 , which are discussed in further detail below.

[0084] Corresponds to Fig.14 Operation 1402, Fig.15A A cross-sectional view of a portion of a semiconductor device 1500 is shown according to various embodiments, the semiconductor device 1500 including a first substrate (or chip) 1502 having a plurality of components 1504 , 1506 , and 1508 formed on a front surface of the first substrate 1502 .

[0085] The first substrate 1502 may be a doped (e.g., using a p-type or n-type dopant) or undoped semiconductor substrate, such as a bulk semiconductor, a semiconductor on insulator (SOI) substrate, etc. The first substrate 1502 may be a wafer, such as a silicon wafer. Typically, an SOI substrate includes a semiconductor material layer formed on an insulator layer. The insulator layer may be, for example, a buried oxide (BOX) layer, a silicon oxide layer, etc. The insulator layer is disposed on a substrate, typically a silicon or glass substrate. Other substrates, such as multilayer or gradient substrates, may also be used. In some embodiments, the semiconductor material of the substrate 1502 may include silicon; germanium; compound semiconductors, including silicon carbide, gallium arsenide, gallium phosphide, indium phosphide, indium arsenide, and / or indium antimonide; alloy semiconductors, including SiGe, GaAsP, AllnAs, AlGaAs, GainAs, GainP, and / or GainAsP; or combinations thereof.

[0086] exist Fig.15AIn the illustrated example of (and the following figure), component 1504 can represent multiple devices, such as transistors, memory cells, etc.; component 1506 can represent multiple via structures electrically coupled to component 1504; component 1508 can represent multiple interconnect structures electrically coupled to via structures 1506, respectively. Such components 1504 to 1508 can be covered by a dielectric layer 1510, which is generally referred to as an interlayer dielectric (ILD) or an intermetallic dielectric (IMD). When forming such components, according to some embodiments, one of the above-mentioned memory array layers or peripheral layers may have been formed. For example, for the memory array layer, component 1504 can represent: (i) multiple memory cells that are collectively used as one or more memory arrays (e.g., one of 210A to 210C); and (ii) multiple transistors that are collectively used as one or more basic circuits (e.g., one of 212A to 212C and one of 214A to 214C). Components 1306 and 1308 may represent: (i) a plurality of local access lines (e.g., bit lines, word lines, etc.) of a memory array; and (ii) a plurality of interconnect structures coupled to the memory array. It should be noted that the same process of forming components 1504 to 1508 may be performed simultaneously on a second substrate (chip) of semiconductor device 1500, as will be shown below.

[0087] Corresponds to Fig.14 Operation 1404, Fig. 15B 1 shows a cross-sectional view of a portion of a semiconductor device 1500 at one stage of different manufacturing stages according to various embodiments, the semiconductor device 1500 including a first layer and a second layer bonded to each other. Fig. 15B As shown, a first layer (which may be one of the above-mentioned memory array layers or peripheral layers) including a first substrate 1502 and components 1504 to 1508 is bonded to a second layer (which may be one of the above-mentioned memory array layers or peripheral layers). Similar to the first layer, the second layer includes a (second) substrate 1522, components 1524, 1526 and 1528, and ILD / IMD 1530. Fig. 15B In the example shown, the second layer is joined to the first layer by being turned upside down.

[0088] Corresponds to Fig.14 Operation 1406, Fig. 15C A cross-sectional view of a portion of a semiconductor device 1500 at one stage at different stages of fabrication according to various embodiments is shown, wherein a second substrate 1522 is thinned from its back side. As shown, the second substrate 1522 is thinned from its back side. In some embodiments, the second substrate 1522 can be thinned using a polishing process (e.g., a chemical mechanical polishing (CMP) process) while having its front surface coupled to the first substrate 1502.

[0089] Corresponds to Fig.14 Operation 1408, Fig.15D A cross-sectional view of a portion of a semiconductor device 1500 at one stage at different stages of fabrication according to various embodiments is shown, the semiconductor device 1500 including one or more TSVs 1534. As shown, the TSVs 1534 may extend from the back side of the thinned substrate 1522, through the thinned substrate 1522 and the IMD / ILD 1530, and to the components 1508 of the first layer. Thus, the second layer may be operably coupled to the first layer via the TSVs 1534. The TSVs 1534 may be formed by the same process as the TSVs 1504 and have the same material as the TSVs 1504. Therefore, the description is not repeated.

[0090] Corresponds to Fig.14 Operation 1410, Fig.15E A cross-sectional view of a portion of a semiconductor device 1500 at one stage at different stages of manufacture according to various embodiments is shown, the semiconductor device 1500 including a plurality of bonding pads 1540 coupled to TSVs 1534, respectively. The bonding pads 1540 may allow the TSVs 1534 to be electrically coupled to other components, such as one or more other layers, to form one of the memory systems described above. The bonding pads 1540 are formed of a conductive material. The conductive material may include copper, although other suitable materials such as aluminum, alloys, doped polysilicon, combinations thereof, etc. may be used instead.

[0091] Fig.16 A flow chart of a method 1600 for forming a memory system including different layers operatively coupled to each other through a through-hole structure according to one or more embodiments of the present invention is shown. For example, at least some operations (or steps) of the method 1600 may be used to form the memory system discussed above. Note that the method 1600 is merely an example and is not intended to limit the present disclosure. Therefore, it should be understood that the method 1600 may be used in the following embodiments. Fig.16 Additional operations are provided before, during, and / or after method 1600, and only some of the other operations are briefly described herein.

[0092] Method 1600 begins with operation 1602, where a plurality of peripheral transistors are formed along a major surface of a substrate. Such transistors formed along the major surface are generally referred to as part of a front-end-of-line (FEOL) network or process. In some embodiments, these peripheral transistors may correspond to Fig.11 The peripheral layer 1102 is shown in .

[0093] The substrate may be a doped (e.g., using a p-type or n-type dopant) or undoped semiconductor substrate, such as a bulk semiconductor, a semiconductor on insulator (SOI) substrate, etc. The substrate may be a wafer, such as a silicon wafer. Typically, an SOI substrate includes a semiconductor material layer formed on an insulator layer. The insulator layer may be, for example, a buried oxide (BOX) layer, a silicon oxide layer, etc. The insulator layer is disposed on a substrate, typically a silicon or glass substrate. Other substrates, such as multilayer or gradient substrates, may also be used. In some embodiments, the semiconductor material of the substrate may include silicon; germanium; compound semiconductors, including silicon carbide, gallium arsenide, gallium phosphide, indium phosphide, indium arsenide, and / or indium antimonide; alloy semiconductors, including SiGe, GaAsP, AllnAs, AlGaAs, GainAs, GainP, and / or GainAsP; or combinations thereof.

[0094] The method 1600 proceeds to operation 1604 where a plurality of metallization layers are formed on the primary surface, wherein at least a first metallization layer includes a first memory array, a first row control circuit, and a first column control circuit, and at least a second metallization layer includes a second memory array, a second row control circuit, and a second column control circuit. Such metallization layers formed over the primary surface are generally referred to as part of a back-end of line (BEOL) network or process. In some embodiments, the first metallization layer and the second metallization layer may correspond to Fig.11 Memory array layer 1104 and memory array layer 1106 are shown in FIG.

[0095] In various embodiments, the memory array and the row / column control circuit formed in the metallization layer may include a plurality of two-dimensional back-gate transistors (2D transistors) and / or a plurality of three-dimensional back-gate transistors (3D transistors), respectively. Fig.17 and Fig.18 As shown. Fig.17 In the illustrative example of , the 2D transistor includes a bottom gate 1702, a gate dielectric 1704 disposed above the bottom gate 1702, a channel structure 1706 disposed on the gate dielectric 1704, and a pair of source / drain structures 1708 and 1710 disposed above the channel structure 1706. The term "two-dimensional back-gate transistor" may refer to a transistor whose gate is formed as a relatively planar or thin structure and whose channel structure is in contact with the top surface of its gate. In some embodiments, the bottom gate 1702 includes TiN, the gate dielectric 1704 includes a high-k dielectric material (e.g., HfO2), the channel structure 1706 includes InGaZnO (IGZO), and the source / drain structures 1708 and 1710 include TiN. Fig.18In the illustrative example of , the 3D transistor includes a bottom gate 1802, a gate dielectric 1804 disposed above the bottom gate 1802, a channel structure 1806 disposed above the gate dielectric 1804, and a pair of source / drain structures 1808 and 1810 disposed on the channel structure 1806. The term "three-dimensional back-gate transistor" may refer to a transistor whose gate is formed as a relatively protruding structure and whose channel structure contacts multiple surfaces of its gate. In some embodiments, the bottom gate 1802 includes TiN, the gate dielectric 1804 includes a high-k dielectric material (e.g., HfO2), the channel structure 1806 includes InGaZnO (IGZO), and the source / drain structures 1808 and 1810 include TiN.

[0096] Method 1600 proceeds to operation 1606, coupling the peripheral transistor to each of the first and second metallization layers through a via structure. Such a via structure may be formed between any adjacent metallization layers of the metallization layer. In some embodiments, the via structure may correspond to Fig.11 The interconnect structure 1112 / 1114 or at least a portion thereof shown in FIG.

[0097] In one aspect of the present invention, a memory circuit is disclosed. The memory circuit includes a first memory array, the first memory array including a plurality of first memory cells, the plurality of first memory cells configured to store a first data element; a second memory array vertically spaced from the first memory array, including a plurality of second memory cells, the plurality of second memory cells configured to store a second data element; and a control circuit operatively coupled to both the first memory array and the second memory array, the control circuit configured to provide a multiply-accumulate (MAC) value based at least on simultaneously multiplying the third data element by the first data element and multiplying the third data element by the second data element.

[0098] In some embodiments, the control circuitry is further configured to provide the MAC value by adding a first product of the third data element and the first data element and a second product of the second data element and the third data element.

[0099] In some embodiments, the control circuit is further configured to access the first memory array and the second memory array simultaneously to retrieve the first data element and the second data element.

[0100] In some embodiments, the first memory array includes a plurality of first local access lines and a plurality of second local access lines, each of the first memory cells being operatively coupled to a corresponding one of the plurality of first local access lines and a corresponding one of the second plurality of local access lines; and wherein the second memory array includes a plurality of third local access lines and a plurality of fourth local access lines, each of the second memory cells being operatively coupled to a corresponding one of the third local access lines and a corresponding one of the fourth local access lines.

[0101] In some embodiments, the first plurality of local access lines and the third plurality of local access lines are connected in parallel with each other, and the second plurality of local access lines and the fourth plurality of local access lines are connected in parallel with each other.

[0102] In some embodiments, the memory circuit further comprises: a plurality of first local access lines, a plurality of third local access lines, a first global access line connected to the plurality of first local access lines and the plurality of third local access lines; and a second global access line connected to the plurality of second local access lines and the plurality of fourth local access lines.

[0103] In some embodiments, the memory circuit further includes: a first switch connected between the first global access line and a corresponding one of the plurality of first local access lines; a second switch connected between the first global access line and a corresponding one of the plurality of third local access lines; a third switch connected between the second global access line and a corresponding one of the plurality of second local access lines; and a fourth switch connected between the second global access line and a corresponding one of the plurality of fourth local access lines.

[0104] In some embodiments, when providing the MAC value, the control circuit is further configured to activate the first to fourth switches simultaneously.

[0105] In some embodiments, when programming the first data element into the first memory array, the control circuit is further configured to: enable the first switch and the third switch; and disable the second switch and the fourth switch.

[0106] In some embodiments, when programming the second data element into the second memory array, the control circuit is further configured to: disable the first switch and the third switch; and enable the second switch and the fourth switch.

[0107] In some embodiments, the first data element and the second data element are both weight data elements, and the third data element is an input data element.

[0108] In some embodiments, the first memory array is formed in a first physical layer, the second memory array is formed in a second physical layer, and the first physical layer and the second physical layer are vertically spaced apart from each other. In another aspect of the present invention, a memory circuit is disclosed. The memory circuit includes a first memory array formed in a first physical layer, the first memory array including a plurality of first storage cells, the plurality of first storage cells being configured to store a first data element; a second memory array formed in a second physical layer, the plurality of second storage cells being configured to store a second data element, wherein the first physical layer and the second physical layer are vertically spaced apart from each other; and a control circuit operatively coupled to both the first memory array and the second memory array. The control circuit is configured to receive a third data element; and based on simultaneously accessing the first memory array and the second memory array, provide a multiply-accumulate (MAC) value of the third data element multiplied by each of the first and second data elements.

[0109] In some embodiments, the first data element and the second data element are both weight data elements, and the third data element is an input data element.

[0110] In some embodiments, the control circuit is further configured to provide the MAC value by adding a first product of the third data element and the first data element and a second product of the third data element and the second data element.

[0111] In some embodiments, wherein the first memory array includes a plurality of first local access lines and a plurality of second local access lines, each of the first memory cells is operatively coupled to a corresponding one of the plurality of first local access lines and a corresponding one of the plurality of second local access lines; wherein the second memory array includes a plurality of third local access lines and a plurality of fourth local access lines, each of the second memory cells is operatively coupled to a corresponding one of the plurality of third local access lines and a corresponding one of the plurality of fourth local access lines; and wherein the memory circuit further includes a first global access line connected to the plurality of first local access lines and the plurality of third local access lines, and a second global access line connected to the plurality of second local access lines and the plurality of fourth local access lines.

[0112] In some embodiments, the memory circuit further includes: a first switch connected between the first global access line and a corresponding one of the plurality of first local access lines; a second switch connected between the first global access line and a corresponding one of the plurality of third local access lines; a third switch connected between the second global access line and a corresponding one of the plurality of second local access lines; and a fourth switch connected between the second global access line and a corresponding one of the plurality of fourth local access lines.

[0113] In some embodiments, when providing the MAC value, the control circuit is further configured to activate the first switch to the fourth switch simultaneously; when programming the first data element into the first memory array, the control circuit is further configured to activate the first switch and the third switch, and deactivate the second switch and the fourth switch; and when programming the second data element into the second memory array, the control circuit is further configured to deactivate the first switch and the third switch, and activate the second switch and the fourth switch.

[0114] In another aspect of the present invention, a method for manufacturing a memory circuit is disclosed. The method includes forming a first memory array in a first physical layer. The method includes forming a second memory array in a second physical layer. The method includes coupling the first physical layer to the second physical layer, the first and second physical layers being vertically spaced apart from each other. The first memory array and the second memory array are accessed simultaneously to retrieve a first data element from the first memory array and a second data element from the second memory array, and the retrieved first and second data elements are multiplied by a third data element.

[0115] In some embodiments, the method further comprises: coupling the first global access line and the second global access line to the first memory array while disconnecting the first global access line and the second global access line from the second memory array to program the first data element into the first memory array; coupling the first global access line and the second global access line to the second memory array while disconnecting the first global access line and the second global access line from the first memory array to program the second data element into the second memory array; and coupling the first global access line and the second global access line to both the first memory array and the second memory array to access the first memory array and the second memory array simultaneously.

[0116] As used herein, the terms "about" and "approximately" generally refer to a value for a given quantity that may vary based on a particular technology node associated with the subject semiconductor device. Based on a particular technology node, the term "about" may indicate a value for a given quantity that varies within, for example, 10-30% of a value (e.g., +10%, ±20%, or ±30% of a value).

[0117] The foregoing summarizes the features of several embodiments so that those skilled in the art can better understand aspects of the present disclosure. Those skilled in the art should understand that they can easily use the present disclosure as a basis for designing or modifying other processes and structures to achieve the same purposes and / or achieve the same advantages as the embodiments introduced herein. Those skilled in the art should also recognize that such equivalent constructions do not depart from the spirit and scope of the present disclosure, and that they can make various changes, substitutions and modifications without departing from the spirit and scope of the present disclosure.

Claims

1. A memory circuit, comprising: a first memory array comprising a first plurality of memory cells configured to store a first data element; a second memory array vertically spaced from the first memory array and comprising a second plurality of memory cells configured to store a second data element; as well as A control circuit is operatively coupled to both the first memory array and the second memory array and is configured to provide a multiply-accumulate (MAC) value based at least on simultaneously multiplying a third data element by the first data element and multiplying the third data element by the second data element.

2. The memory circuit according to claim 1, wherein: The control circuit is further configured to provide the multiplication-accumulated value by adding a first product of the third data element and the first data element and a second product of the second data element and the third data element.

3. The memory circuit according to claim 1, wherein: The control circuit is further configured to simultaneously access the first memory array and the second memory array to retrieve the first data element and the second data element.

4. The memory circuit according to claim 1, in, the first memory array comprising a first plurality of local access lines and a second plurality of local access lines, each of the first memory cells being operatively coupled to a respective one of the first plurality of local access lines and a respective one of the second plurality of local access lines; as well as The second memory array includes a plurality of third local access lines and a plurality of fourth local access lines, and each of the second memory cells is operatively coupled to a corresponding one of the third local access lines and a corresponding one of the fourth local access lines.

5. The memory circuit according to claim 4, wherein: The plurality of first local access lines and the plurality of third local access lines are connected in parallel to each other, and the plurality of second local access lines and the plurality of fourth local access lines are connected in parallel to each other.

6. A memory circuit comprising: a first memory array formed in a first physical layer and comprising a first plurality of memory cells configured to store a first data element; a second memory array formed in a second physical layer and comprising a plurality of second memory cells configured to store second data elements, wherein the first physical layer and the second physical layer are vertically spaced apart from each other; and a control circuit operatively coupled to both the first memory array and the second memory array, wherein the control circuit is configured to: receiving a third data element; and A multiply-accumulate (MAC) value of the third data element multiplied by each of the first data element and the second data element is provided based on concurrently accessing the first memory array and the second memory array.

7. The memory circuit according to claim 6, wherein: The first data element and the second data element are both weight data elements, and the third data element is an input data element.

8. The memory circuit according to claim 6, wherein: The control circuit is further configured to provide the multiplication-accumulated value by adding a first product of the third data element and the first data element and a second product of the third data element and the second data element.

9. A method of operating a memory circuit, comprising: forming a first memory array in a first physical layer; forming a second memory array in a second physical layer; as well as coupling the first physical layer to the second physical layer, the first physical layer and the second physical layer being vertically spaced apart from each other; wherein the first memory array and the second memory array are accessed simultaneously to retrieve a first data element from the first memory array and a second data element from the second memory array, and the retrieved first data element and the second data element are multiplied by a third data element.

10. The method according to claim 9, further comprising: coupling a first global access line and a second global access line to a first memory array while disconnecting the first global access line and the second global access line from the second memory array to program the first data element into the first memory array; coupling the first global access line and the second global access line to the second memory array while disconnecting the first global access line and the second global access line from the first memory array to program the second data element into the second memory array; as well as The first global access line and the second global access line are coupled to both the first memory array and the second memory array to access the first memory array and the second memory array simultaneously.