Integrated circuit having ganged microvault memories

WO2026178137A1PCT designated stage Publication Date: 2026-08-27VERSUM MATERIALS US LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/015686
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-19
Filing Date
2026-02-18
Publication Date
2026-08-27

Smart Images

  • Figure US2026015686_27082026_PF_FP_ABST
    Figure US2026015686_27082026_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure relates to a memory system comprising multiple modules, including a first and second module with overlapping read address spaces. The system features a ganged read-address bus that facilitates coordinated reading from the first and second modules through dedicated read peripherals. Each read peripheral is equipped with specific read-address, read-clock, and output-enable ports, allowing for efficient data retrieval. The system can utilize ferroelectric field-effect transistors (FeFETs) in its bit-cells, and may incorporate local buffers and multiplexers.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. P25-059-SEC-W001INTEGRATED CIRCUIT HAVING GANGED MICROVAULT MEMORIESCROSS-REFERNCE TO RELATED APPLICATION

[0001] This application claims the benefit of U.S. Provisional Application No.63 / 760,228, filed February 19, 2025, the entire disclosure of which is hereby incorporated in its entirety.BACKGROUNDRelevant Field

[0002] The present disclosure relates to integrated circuits. More particularly, the present disclosure relates to integrated circuits having multiple modules including a microvault memory.Description of Related Art

[0003] Chiplets refer to miniature chips that are designed to work as a single entity while using advanced packaging technology. These miniaturized chips are created by dividing the larger chip into several smaller chips, each with its own function or capability. The concept originated from the semiconductor industry's need to overcome the physical restrictions of traditional monolithic chip designs and achieve higher levels of integration. The idea behind chiplets is to create a modular system of interconnected and interchangeable chips that can be combined in different configurations to create advanced computing systems with improved performance, power efficiency, and functionality.

[0004] Chiplets can be based on different architectures, such as CPU, GPU, memory, or IO, and can be assembled and stacked in a variety of ways, depending on the specific application requirements. One of the advantages of the chiplet approach is the ability to mix and match different chiplets from different manufacturers to create custom solutions that meet specific computing needs. This approach also allows for faster time to market, reduced development costs, and increased flexibility, as chiplets can be upgraded or replaced without the need for a complete system redesign.

[0005] The use of chiplets may be used in various industries, including consumer electronics, cloud computing, and data centers, where the demand for high-performance computing and energy efficiency is high. Chiplets are expected to play a significant role in the future of computing and are likely to unlock new possibilities for creating more powerful and / or sophisticated electronic devices.Attorney Docket No. P25-059-SEC-W001SUMMARY

[0006] The present disclosure relates to a memory system that may include a group of modules comprising multiple modules, specifically a first module and a second module. The first module may possess a first read address space and a first write address space, while the second module may have a second read address space and a second write address space. Notably, the first read address space may at least partially overlap with the second read address space. The memory system may also feature a write peripheral that is configured to receive a write address, write data, and a write clock signal. This write peripheral may be capable of writing the write data to one or more of the modules based on the provided write address.

[0007] Additionally, the memory system may include a plurality of read peripherals, which may consist of a first read peripheral and a second read peripheral. The first read peripheral may be coupled to the first module, while the second read peripheral may be coupled to the second module. Each read peripheral may have its own read-address port, read-clock port, and output-enable port. A ganged read-address bus may be coupled to both read peripherals, allowing it to provide a read address that includes a first portion and a second portion. The first portion may be connected to the read-address ports of both read peripherals, while the second portion may be linked to their respective output-enable ports. This configuration may enable the memory system to perform read operations from both read peripherals using the ganged read-address bus.

[0008] The memory system may further allow access to a first memory space of the first module and a second memory space of the second module via the ganged read- address bus. These memory spaces may be contiguous when accessed through the read address provided by the ganged read-address bus. The second portion of the read address may consist of a predetermined bit, which could be the most significant bit of the read address. The ganged read-address bus may be constructed from a plurality of conductive paths, and the first portion of the read address may include the least significant bits corresponding to addresses within both modules. The second portion may include a most significant bit that is configured to select between the first and second read peripherals through their output-enable ports.

[0009] Each read peripheral may be designed to receive the first portion of the read address at its respective read-address port, with this portion corresponding to addresses within its associated module. The write peripheral may also be capable ofAttorney Docket No. P25-059-SEC-W001writing data to multiple modules simultaneously based on the write address. Each module may comprise a plurality of bit-cells, with each bit-cell potentially including a ferroelectric field-effect transistor (FeFET).

[0010] Moreover, the memory system may include a multiplexer that is configured to select one of the modules based on the most significant bit of the read address. The write peripheral may be designed to perform write operations at a slower speed compared to the read operations conducted by the plurality of read peripherals. The modules may be arranged in a three-dimensional configuration, and a third read peripheral may be included to facilitate parallel read operations alongside those performed via the ganged read- address bus.

[0011] The write peripheral may incorporate a shared write logic system that utilizes a shift register-based design. An interlock mechanism may be integrated into the memory system to prevent simultaneous read and write operations to the same module. The number of modules whose read peripherals are ganged via the ganged read-address bus may be configurable, and the memory system may be capable of dynamically adjusting this number based on power consumption requirements.

[0012] Each read peripheral may further include a local buffer that is configured to temporarily store read data before it is output. The memory system may also be designed to adjust read latency based on the number of modules whose read peripherals are ganged via the ganged read-address bus. Additionally, a multiplexer may be included to receive the second portion of the read address and generate an enable signal, which may be provided to a selected output-enable port of a respective read peripheral to facilitate data reading from the corresponding module.

[0013] Each read peripheral may be configured to receive a read clock signal, which may be synchronized across all read peripherals to coordinate read operations. The read clock signal may be coupled to both the first and second read-clock ports. Furthermore, a logic circuit may be included to invert the most significant bit of the read address before it is sent to the second output-enable port of the second read peripheral. Finally, the read data paths from the first and second read peripherals may be configured to converge into a common data bus, enhancing the efficiency of data retrieval within the memory system.

[0014] In each and every embodiment, each of the plurality of modules may include a respective write port and write peripheral when appropriate, feasible, and / or suitable.Attorney Docket No. P25-059-SEC-W001BRIEF DESCRIPTION OF THE DRAWINGS

[0015] These and other aspects will become more apparent from the following detailed description of the various embodiments of the present disclosure with reference to the drawings wherein:

[0016] Fig. 1 is a block diagram of an integrated circuit that may be part of a semiconductor device such as a chiplet in accordance with an embodiment of the present disclosure;

[0017] Fig. 2 shows a perspective view of an assembly having the integrated circuit of Fig. 1 implemented on a semiconductor device that is electrically connected to another device to form the assembly in accordance with an embodiment of the present disclosure;

[0018] Fig. 3 shows a block diagram illustrating the memory address space of the integrated circuit of Fig. 1 in accordance with an embodiment of the present disclosure;

[0019] Fig. 4 shows a block diagram illustrating the memory address space with the signal interfaces of the integrated circuit of Fig. 1 in accordance with an embodiment of the present disclosure;

[0020] Fig. 5 shows an illustration of an integrated circuit that may be part of a semiconductor device such as a chiplet in accordance with an embodiment of the present disclosure;

[0021] Fig. 6 shows a perspective of an assembly having the integrated circuit of Fig. 1 implemented on a semiconductor device that is electrically connected to a system-on-a-chip in accordance with an embodiment of the present disclosure;

[0022] Fig. 7 shows a perspective view of an assembly having a semiconductor device with an array of processing elements and a second semiconductor device having an array of microvaults;

[0023] Fig. 8 shows an assembly of a semiconductor devices including several memory types in accordance with an embodiment of the present disclosure;

[0024] Fig. 9 shows an assembly of semiconductor devices including a semiconductor device with a system-on-chip and another semiconductor with microvaults disposed on top in accordance with an embodiment of the present disclosure;Attorney Docket No. P25-059-SEC-W001

[0025] Fig. 10 shows a semiconductor assembly incorporating a daisy-chained configuration of microvaults operatively connected to a multiplexer and managed by a counter for coordinated data selection and retrieval, in accordance with an embodiment of the present disclosure.

[0026] Fig. 11 shows a semiconductor assembly incorporating a daisy-chained configuration of microvaults in multiple semiconductor devices that are operatively connected to multiplexers and managed by counters for coordinated data selection and retrieval, in accordance with an embodiment of the present disclosure;

[0027] Fig. 12 illustrates a three-dimensional (3D) memory column configured as a 3D-NOR or 3D-AND structure, featuring a series of ferroelectric field-effect transistors (FeFETs) with interconnected drain terminals linked to a common select line and individual gate terminals connected to respective read / write enable lines, all coupled to a common bit line, in accordance with an embodiment of the present disclosure;

[0028] Fig. 13 depicts a three-dimensional (3D) memory column configured as a 3D-NAND structure, consisting of a vertical stack of ferroelectric field-effect transistors (FeFETs), in accordance with an embodiment of the present disclosure;

[0029] Fig. 14 depicts a three-dimensional (3D) memory column configured as a 3D-NAND with an integrated pass gate, in accordance with an embodiment of the present disclosure;

[0030] Fig. 15 illustrates a three-dimensional (3D) memory column 1500, which may be configured as either a 3D-NOR or a 3D- AND structure with independent Read / Write enable capabilities, in accordance with an embodiment of the present disclosure;

[0031] Fig. 16 shows a cross-sectional view of a 3D memory structure configured as a single-port 3D NAND, in accordance with an embodiment of the present disclosure;

[0032] Fig. 17 shows a cross-sectional view of a 3D memory structure that is a dual-port 3D NAND arrangement, in accordance with an embodiment of the present disclosure;

[0033] Fig. 18 illustrates a 3D memory structure that can be configured as a 3D NOR Vertical Transistor memory array, in accordance with an embodiment of the present disclosure;Attorney Docket No. P25-059-SEC-W001

[0034] Fig. 19 shows a planar FeFET in accordance with an embodiment of the present disclosure;

[0035] Fig. 20 shows electrical characteristics of an embodiment of a FeFET in accordance with an embodiment of the present disclosure; and

[0036] Fig. 21 shows a block diagram illustrating ganging of modules in accordance with an embodiment of the present disclosure.fDET AILED DESCRIPTION

[0037] Fig. 1 shows a block diagram of an integrated circuit 100 that may be packaged as a bondable chiplet (e.g., face-to-face chiplet bondable) in accordance with an embodiment of the present disclosure. The integrated circuit (IC) 100 includes a modules group 106 consisting of modules 108, 110, 112, and 114. The IC 100 also features a shared write port 102, configured to write to the modules group 106 using a write peripheral 104. Additionally, it includes read peripherals 116, 118, 120, and 122 and read ports 124, 126, 128, and 130, configured to read from the modules 108, 110, 112, 114.

[0038] The write port 102 may be configured to provide a single write address space for all of the modules group 106 where each of the modules 108, 110, 112, 114 has a dedicated read port 124, 126, 128, 130 respectively. The integrated circuit 100 may be packaged as part of a chiplet configured to be electrically connected to another integrated circuit device (e.g., another chiplet, or IC package, with or without electrical contacts, electrical bumps, etc.). The chiplet may be electrically connected to another device including, for example, by bonding, soldering, wafer-to-wafer bonding, face-to-face chiplet bonding, chiplet-to-wafer bonding, chiplet-to-interposer bonding, and / or may be connected together with an interposer or other interfacing technology. None, one, or more of interposers may be used or other interfacing technologies that are common to heterogeneous 3D system-in-package solutions may be utilized in electrically connecting a chiplet to another device.

[0039] Each read port (124, 126, 128, 130) in the chiplet may feature electrical contacts on a side of the chiplet or on multiple sides of the chiplet. The read ports 124, 126, 128, 130 may use multi -cycle pipelined circuitry. Upon bonding to another device (e.g., wafer, chiplet, chip, SOC, package, FPGA, etc.), the electrical contacts may line up in a manner that provides dedicated access to specific modules of the modules 108, 110, 112, 114. For instance, a processing / computing element may have exclusive access to module 108 via the read port 124, which may contain the neural networkAttorney Docket No. P25-059-SEC-W001weights in a registry file. Similarly, a different processing / computing element may have exclusive read access to module 110 via the read port 126, which includes a different registry file. In this specific embodiment, this arrangement of the electrical contacts ensures that each computing / processing element has the dedicated access it needs to cany out its specific computation efficiently thereby providing a compact, modular, and scalable system that allows different processing elements to maintain dedicated access to specific modules 108, 110, 112, 114. Without dedicated access, different processing elements might have to queue up to use the same resource which would slow down overall processing speed. By providing dedicated access, the proposed chiplet ensures that each processing element can operate at its maximum capability without interference from other computing elements in this specific embodiment.

[0040] The write peripheral 104 is a peripheral circuitry responsible for processing and writing data into the memory cells found within the modules 108, 110, 112, 114. The write peripheral 104 may include dedicated contacts so that a chip electrically connected (e.g., bonded) to a chiplet of the integrated circuit, such that the write port 102 is accessible via a shared write logic system that involves utilizing a shift register-based, different voltage design, preferably high voltage design, that has a shared write address and data components. This shared write logic system is designed to be accessed via a bonded chip, another bonded chiplet, and / or via other circuitry in the same package as the integrated circuit 100. A shift register could allow the system to move data through a series of stages, with each subsequent stage receiving the data from the previous stage. By utilizing a shift register, the system can increase the data throughput while maintaining a low rate of data transfers. The shared write address space refers to the location where data is written in the chiplet.

[0041] In another embodiment, an interlock 132 may disable the read ports 124, 126, 128, 130 while data is being written to the modules group 106 via the write port 102. Likewise, the interlock 132 may disable the write port 102 when read operations are being carried out on the read ports 124, 126, 128, 130. The written data can later be accessed concurrently by all processing elements that need to read the data via a respective one of the read ports 124, 126, 128, 130. This ensures that all processing elements have the most commonly used data available to them without regard to other reads being concurrently carried out by other processing elements.

[0042] The write peripheral 104 circuit includes a write driver. This unit receives the data to be written and converts it into suitable signals that can change theAttorney Docket No. P25-059-SEC-W001state of the memory cells. Depending on the type of memory technology used, these signals could involve voltage levels, current pulses, or other types of energy. The shared write logic system may be high voltage due to the specific voltage requirements of the chiplet. The write driver must provide enough power to reliably change the state of the memory cells, but it must also operate within suitable parameters to avoid causing damage or unnecessary wear.

[0043] The write peripheral 104 circuit may also feature a data buffer or write buffer. This component temporarily stores the data to be written, allowing the write operation to be performed at an optimal pace. By balancing the speed of incoming data with the speed at which the memory cells can be written, the write buffer helps prevent data loss and optimizes system performance.

[0044] The write peripheral 104 may also include, in some embodiment, a write control unit that orchestrates the sequence of operations in the write process. It generates control signals to activate the write driver at the appropriate times, controls the flow of data from the write buffer, and coordinates the timing of the write operations. By synchronizing these various activities, the write control unit ensures efficient and reliable write operations.

[0045] The write peripheral 104 may also include data encoding mechanisms to improve reliability and data integrity. For example, before the data is written to the memory cells, these mechanisms encode it in a way that allows potential errors to be detected, and in some cases, corrected when the data is later read. This can be helpful in systems where data integrity has a higher priority, such as in servers or scientific research devices.

[0046] The write peripheral 104 may also include a timing unit that serves as the system's heartbeat, supplying clock signals that synchronize the operation of the system’s various components. In some systems, it may include components like oscillators, clock generators, or phase-locked loops. The timing unit may ensure that all operations occur at the suitable time relative to each other.

[0047] The IC 100 may be implemented as a face-to-face bonded chiplet, with modules 108, 110, 112, and 114 formed from a non-volatile memory. In some specific embodiments, the IC 100 may also feature a dynamic allocation circuitry to allocate memory blocks to the modules group 106 based on the usage of the modules group 106 (e.g., each module 108 may include dynamic allocation circuitry for dynamically allocating a range of read locations for a respective processing element).Attorney Docket No. P25-059-SEC-W001

[0048] The IC 100 features a plurality of clocks, with each clock of the plurality of clocks feeding a respective module of the plurality of modules, providing each respective module with decoupled timing relative to the other modules of the plurality of modules. The modules group 106 may be arranged in any topology known to one of ordinary skill in the relevant art. Bit-cell density can be up to 10 times more dense than embedded SRAM cells in the modules group 106.

[0049] The IC 100 may be formed on a chiplet that includes a first side and a second side, with the second side configured for bonding to a second semiconductor device. The IC 100 may include a high voltage write logic adjacent to the first side of the chiplet. A decoder circuitry, a driver circuitry, and a register circuitry may be formed on the silicon substrate portion of the chiplet, while the modules group 106 is formed on a second layer portion of the chiplet. The second semiconductor device may comprise a plurality of processing elements. Each processing element includes a respective interface to communicate with a respective module of the plurality of modules on the modules group 106 when the second semiconductor device is bonded to the chiplet.

[0050] The silicon substrate traditionally serves as the initial stage of IC fabrication, focusing on the creation of active components, particularly transistors. Techniques like diffusion, ion implantation, oxidation, and material deposition are employed to fashion the intricate structures of transistors. These processes operate at small scales. The application of photolithography, etching, and implantation techniques enables the definition of transistor structures with precision. The silicon substrate’s significance lies in its ability to establish the fundamental building blocks used for signal processing, amplification, and control within the IC. This layer is sometimes called Front-End-Of-The-Line (“FEOL”).

[0051] Next in the manufacturing process, a second layer may be added that traditionally takes on the role of interconnect fabrication, facilitating the electrical connections between various IC components. This phase traditionally focused on the creation of passive components, including interconnects, vias, and metal-insulator-metal (MIM) capacitors. The second layer processes typically differ from the processes used on the silicon substrate in terms of precision and scale. The interconnects are formed by depositing and patterning metal layers, typically aluminum or copper, to construct the wiring network. Dielectric layers, such as silicon dioxide or low-k dielectrics, are introduced to insulate the interconnects and prevent signal interferenceAttorney Docket No. P25-059-SEC-W001between different wiring layers. The second layer’s traditional function is to establish the interconnections that enable the routing and distribution of electrical signals throughout the IC. However, as described herein, circuity may be utilized within this second layer (sometimes referred to as Back-End-Of-The-Line (“BEOL”)).

[0052] Alternate embodiments of the IC 100 may be implemented as a stacked die, a monolithic design, TSVs, or silicon through vias. In a stacked die design, several dies may be stacked on top of each other, with each die performing different functions, such as memory and processing. The stacked die may communicate through wire bonds, microbumps, or bump-less bonds. In a monolithic design, the various functions and modules of the IC 100 may be integrated onto a single die, forming a more compact and power-efficient design.

[0053] Additionally, the IC 100 may include one or more interlocks 132 to prevent conflicts in reading and writing data. The modules group 106 may be formed from a variety of non-volatile or semi-volatile (e.g., very long refresh periods) memory technologies, such as Static Random-Access Memory (SRAM), Ferroelectric Field Effect Transistor (FeFET), Ferroelectric Random Access Memory (FeRAM), Resistive Random Access Memory (ReRAM), Spin-Orbit Torque (SOT) Memory, Spin Transfer Torque (STT) Memory, charge trap, floating gate memories, and / or Schottky diodes.

[0054] The modules group 106 may utilize a Static Random-Access Memory (SRAM) Topology. The SRAM topology may employ a cross-coupled flip-flop structure (e.g., latching flip-flops), ensuring the stored data remains intact as long as power is supplied. Thus, in some specific embodiments, the modules group 106 may utilize heterogeneous types of memory including volatile and non-volatile memory types.

[0055] The modules group 106 may utilize a Flash Memory Topology. The Flash memory is a non-volatile memory technology used in applications where data persistence is needed, such as solid-state drives (SSDs) and USB flash drives. The flash memory topology disclosed herein features a matrix of memory cells, each consisting of a floating-gate transistor or charge trap device. The modules group 106 may also use wear-leveling techniques to prolong the lifespan of the memory cells.

[0056] The modules group 106 may utilize a Ferroelectric Random- Access Memory (FeRAM) Topology. The FeRAM topology utilizes a ferroelectric material capable of retaining polarization states. One such memory topology may, in specific embodiments, utilize a FeFET to retain state information and program the ferroelectricAttorney Docket No. P25-059-SEC-W001material. These ferroelectric materials may be used to retain state information and act as a memory bit cell.

[0057] The modules group 106 may utilize a Phase Change Memory (PCM) Topology, which is a non-volatile memory technology that utilizes reversible phase changes in materials to store data. The PCM topology may include any phase change material, for example a chalcogenide alloy or a chalcogenide glass housed within a memory cell.

[0058] The modules group 106 may utilize a Resistive Random-Access Memory (ReRAM) Topology, which is a non-volatile memory technology based on resistive switching phenomena. The ReRAM topology may utilize a thin-film material that exhibits reversible changes in resistance upon the application of electrical stimuli.

[0059] The modules group 106 may utilize a Spin-Orbit Torque (SOT) Magnetic Random-Access Memory Topology. SOT-MRAM is a type of non-volatile memory that utilizes spin-orbit torque to switch the magnetic state of a storage element. The SOT-MRAM topology may incorporate a magnetic tunnel junction (MT J) structure and leverages the spin-orbit coupling effect to write and read data. The magnetic tunnel junction may have a dielectric layer between a magnetic fixed layer and a magnetic free layer. Writing may be done by switching magnetization of the free magnetic layer by injecting an in-plane current in an adjacent SOT layer. Reading may be done by putting current into the magnetic tunnel junction. The SOT-MRAM can optimize the spinorbit materials by using current-driven switching schemes while minimizing write energy consumption, in some specific embodiments.

[0060] The modules group 106 may utilize a Spin Transfer Torque (STT) Magnetic Random-Access Memory Topology. The STT-MRAM is another type of non-volatile memory that relies on spin transfer torque to manipulate the magnetic state of a storage element. The STT-MRAM topology can use a magnetic tunnel junction (MTJ) structure, where the magnetization orientation determines the stored data. Additionally, the orientation of a magnetic layer in a magnetic tunnel junction or spin valve can be changed using a spin-polarized current, for example.

[0061] The IC 100 may include a single write peripheral 104 with a dedicated clock, or each module 108, 110, 112, 114 may have its own dedicated write peripheral utilizing a shared clock (not shown in Fig. 1). Additionally, the modules group 106 may be organized into separate partitions, each with a dedicated read peripheral 116, 118, 120, 122 having an independent clock.Attorney Docket No. P25-059-SEC-W001

[0062] Another possible embodiment of the IC 100 includes an interface (e.g., the same, different, higher or lower voltage) to enable data transfer external to the packaging of the IC 100. The IC 100 may also include an integrated microcontroller unit (MCU) or a digital signal processor (DSP) for processing data within the IC in yet additional specific embodiments.

[0063] Fig. 2 illustrates a perspective view of assembly 200, which features an integrated circuit 212 implemented on a chiplet 230. This chiplet 230 comprises a plurality of modules, specifically microvaults, each disposed in spaced relation to one another and adjacent to a first surface 228. Each respective microvault, as a type of module, has a dedicated read periphery 202a and a write periphery 202b, allowing for efficient data operations. The chiplet 230 is bonded to a second device 226, which may take various forms, such as another chiplet, a semiconductor wafer, or an Al accelerator. The integrated circuit 212 encompasses the circuitry within chiplet 230, while the second device 226 allows each processing unit to access specific modules within modules group 236. For instance, the second device 226 may function as a network controller, utilizing an offload circuit to manage data packets.

[0064] Although the second device 226 may use the shared write port 222 via an address & data bus with a clock and a enable signal to write data to any modules within the modules group 236, other ways of writing data may be considered. For example, serial connections, parallel connections, various buses, or ports, may be used, such as a DDR (Double data Rate) Interface, a SRAM (Static Random-Access Memory) Interface, a NAND Flash Memory Interface, a NOR Flash Memory Interface, a HBM (High Bandwidth Memory) Interface, a GDDR (Graphics Double data Rate) Interface, a NVMe (Non-Volatile Memory Express) Interface, SPI, IC2, etc. Each of the modules has a read port with a read address 218 (to send an address to a module 234) and read data 214 (which is the data read from the module 232.

[0065] The modules group 236 is formed on a chiplet 230 having two sides including a surface 228 that can be bonded to and complement a second device 226. The chiplet 230 may be formed by forming circuitry on a silicon substrate 204 and then by adding a second layer 206. In other embodiments, these layers may be reversed and / or other layers may be added, removed, etc. The read address 218 and read data 220 are used for reading the module 232.

[0066] Although the second device 226 may use an address & data bus with a clock and a enable signal to read data from the module 232, other ways of reading dataAttorney Docket No. P25-059-SEC-W001may be considered. For example, serial connections, parallel connections, various buses, or ports, may be used, such as a DDR (Double data Rate) Interface, a SRAM (Static Random-Access Memory) Interface, a NAND Flash Memory Interface, a NOR Flash Memory Interface, a HBM (High Bandwidth Memory) Interface, a GDDR (Graphics Double data Rate) Interface, a NVMe (Non-Volatile Memory Express) Interface, SPI, IC2, etc.

[0067] All of the read ports (e.g., 218 and 222) are configured to be inactive when a write operation is applied to the shared write port 222. The read ports may also be configured to process reads concurrently with each other. The shared write port 222 is configured to write to an address space, where the shared write port 222 is configured to write to the first module 232 via a first portion of the address space and write to the second module 234 via a second portion of the address space. Each module of the plurality of modules 236 includes an independent read port for concurrent reading via a respective independent read port of any of the plurality of modules.

[0068] Each read port for a respective module may include contacts for circuitry found within the second device 226 to interface via metallic contacts. Thus, there may be metallic contacts on the top layer 208 that are configured to interface with metallic contacts on the surface 228 of the chiplet 230 such that the metallic contacts allow for a read space that is coextensive with a read space of a module of the modules 236. The read spaces of the modules group 236 may all be coextensive with each other.

[0069] In one embodiment, the read peripheral for the first module 232 is implemented on a silicon substrate 204 (sometimes referred to as a Front-end-of-the-line). The second layer 206 (sometimes call the Back-end-of-the-line) may be built next in the manufacturing process on top of the silicon substrate 204 (and any circuitry) and may contain the respective memory bit cells. In an alternative embodiment, the read peripheral for the first module 232 is implemented in the second layer 206 and is disposed between the modules group 236 and the surface 228 of the chiplet 230.

[0070] The modules group 236 may be configured to process write commands only during reset. The write commands may be “slow write” commands. That is, the modules group 236 may have very low write speeds relative to its read speed. The write logic may be frozen (or disabled) when the modules group 236 are used for reading data. In some specific embodiments, the integrated circuit 212 provides functionality to allocate memory blocks to the modules group 236 based on the usage of the modules group 236. In other embodiments, the memory addresses are fixed along with theAttorney Docket No. P25-059-SEC-W001allocation. The integrated circuit 212 may be implemented as a face-to-face bonded chiplet 230. The face-to face bonding may be bump-less wafter bonding.

[0071] The modules group 236 can have a single write peripheral 202. In other embodiments, each module of the modules group 236 may have a dedicated write peripheral that utilizes a shared clock. In yet other embodiments, the modules group 236 may also be organized into separate partitions each with partition having a dedicated read peripheral, where each dedicated read peripheral has an independent clock. The partitions may be one, two, or more modules of the modules group 236.

[0072] The write peripheral 202 circuitry's overall architecture may include a series of different components, including write driver, address decoders, sense amplifiers, data input latches, data bus, etc. and / or some combination thereof. Write drivers or write buffers, may be tasked with transferring data onto the memory cell. They may enhance the input signal to achieve a level appropriate for the memory cell. Address decoders may be used to interpret the memory address that is fed as an input where the data needs to be written. By activating the specific row and column of the memory array linked to that address, they may be used to select the target memory cell. Sense amplifiers may be used to identify and boost the signal from the memory cells during reading operations, also participate in refreshing the memory cell post data write in write operations. The write operation is instigated by a write enable signal. When a write command is initiated, this signal propels the write drivers and decoders into the writing process, data input latches may be used as temporary storage units, retaining the data set to be written into the memory until the write operation is implemented. A data bus with a transmission route, can be used to facilitate the movement of data from the data input latches to the memory cells.

[0073] A write operation to the modules group may be performed through a priority arbitration circuit that facilitates the modules to be accessed in a predetermined order, and the shared write port 222 may be configured to write to a virtual address space that is mapped onto a physical memory space. The integrated circuit 212 may include a high voltage write logic used within the write peripheral 202, and the second semiconductor device 226 may comprise a plurality of processing elements, whereby each processing element includes a respective interface to communicate with a respective module of the modules group 236. Furthermore, the chiplet 230 may include an interface to the shared write port 222 on the second side to thereby interface with a complementary interface on the second semiconductor device 226.Attorney Docket No. P25-059-SEC-W001

[0074] The integrated circuit 212 may also include a power gating circuitry that selectively powers down a module of the modules 236 when not in use. Additionally, the integrated circuit 212 may have a write peripheral 202 of the modules group 236 connected to a dedicated I / O pad to enable data transfer external to the package of the integrated circuit.

[0075] The integrated circuit 212 may utilize multiple modules of the modules group 234 grouped together. These modules may be synchronized with one another in specific embodiments. In some cases, all the modules are synchronized, while in other instances, only specific modules are to be synchronized. For instance, the circuit on a second device 226 may need to synchronize with a specific module when reading data from one of the modules in the module group 236.

[0076] To synchronize the modules, the integrated circuit 212 may use various timing technologies. In some cases, a plurality of clocks may feed each respective module of the modules group 236, thereby allowing each module to have decoupled timing relative to the other modules in the group. This decoupling ensures that any delay in one module will not affect the functioning of other modules. It is worth noting that the clocks used may or may not need to be synchronized. In some cases, a common clock can be used to synchronize the modules. In yet other embodiments, the clock signal or signals may be provided by the second device 226.

[0077] In alternative embodiments, other synchronization techniques can be used, such as phase comparison of the clock signals or a phase-locked loop (PLL) synchronization method. Another embodiment for synchronizing the modules in the IC could use delay-locked loop (DLL) synchronization. In this method, a delay element is added to the clock signal path, and the output is compared to the input clock signal. The feedback loop adjusts the delay element until the output of the DLL matches the input, resulting in synchronization of the clock signals.

[0078] In another embodiment, the integrated circuit 212 could use a combination of different synchronization techniques to achieve synchronization between the modules. For example, some modules may use PLL synchronization while others use clock delay lines or DLL synchronization, depending on their specific requirements. Additionally, the integrated circuit 212 can also use redundant synchronization techniques to ensure reliability and redundancy in case one method fails. For example, the integrated circuit 212 could use both PLL synchronization andAttorney Docket No. P25-059-SEC-W001DLL synchronization simultaneously, so that if one method fails, the other can still maintain synchronization.

[0079] Fig. 3 shows a block diagram 300 illustrating the memory address space of the integrated circuit of Fig. 1 in accordance with an embodiment of the present disclosure. The memory address space includes a write address space 316 and read data address spaces 310, 312, 314.

[0080] The write address space 316 consists of various units where data, e.g., weights, and / or instractions can be stored. These units are referred to as memory addresses. The module group 302 includes multiple memory modules 304, 306, 308. The write address space 316 may be distributed among the memory modules 304. 306, 308 such that the write address space 316 spans from 0 to N*M-1. As shown in Fig. 3, the modules group 302 has N memory modules 304, 306, 308, where N is a positive integer, and each module has a memory size of M. The total number of unique write memory addresses in the write address space will be N*M, which can be referenced by an integer from 0 to N*M-1.

[0081] Starting at 0, memory addresses of the write address space 316 are ordered sequentially up to N*M-1. In other words, the first address is 0 and the final address is N*M-1, encompassing a total of N*M addresses. This ordering can be linear (each address increases by one) or some other specified pattern depending.

[0082] The write memory addressing can be implemented in a variety of ways based on the system architecture. One method used in a specific embodiment is to use the base and limit registers. The base register holds the smallest legal physical write memory address, and the limit register specifies the size of the range. Therefore, to generate a logical address, you would add the base to the relative address. In other embodiments, a memory addressing scheme may be used where the base used is set to be 0. Yet additional write addressing techniques will be appreciated by one or ordinary skill in the relevant art.

[0083] For any device that writes to the modules group 302, each memory module can possess a unique set of write memory addresses such all memory addresses within the modules group 302 is unique with respect to writing data, e.g., the first module starting at 0 and the last one ending at N*M-1. This allocation, in some embodiments, may be dependent on the memory management system of the device writing data to the modules 304, 306, 308, which could range from simple fixed partitioning schemes to more complex dynamic partitioning models.Attorney Docket No. P25-059-SEC-W001

[0084] For instance, in a straightforward linear model where each module (304, 306, or 308) has an equal size of M addresses, the first module 304 would possess write addresses 0 to M-l, the second module would have write addresses M to 2M-1, the third module would have write addresses 2M to 3*M-1, and so forth. The Nth module 308, therefore, would possess write addresses from (N-1)*M to N*M-1.

[0085] It is contemplated that one of ordinary skill in the relevant art may use other implementations of write memory addresses from 0 to N*M-1 that depends on various factors such as the hardware architecture, operating system, memory management schemes, and the nature of the programs being run on the system, etc.

[0086] The modules group 302 has different read data address spaces 310, 312, 314. These read address spaces 310, 312, 314 may have overlapping addresses spaces, may have contiguous address spaces, or may have coextensive address spaces. The read address spaces 310, 312, 314 may be independent relative to each other. The system includes three independent read address spaces, labeled as read address spaces 310, 312, and 314. Each of these read address spaces is distinct from the others, meaning that reads can be performed in each space without affecting the others.

[0087] The read address spaces 310, 312, 314 may be defined as contiguous blocks of memory addresses, each with its own starting address and ending address. In modules group 302, each read address space 310, 312, 314 may have a range of addresses that corresponds to values from 0 to M-l, where M is a maximum value determined by the size of the modules 304, 306, 308 being used.

[0088] In one embodiment, allowing one processing unit to interface with each read address space 310, 312, 314, the concurrent reads may be implemented as described herein. The independence of the read address spaces 310, 312, 314 ensures that each processing unit can access its desired data without causing any interference or conflict with other processing units.

[0089] Fig. 4 shows a block diagram illustrating the memory address space with the signal interfaces of the integrated circuit of Fig. 1 in accordance with an embodiment of the present disclosure. The signals used in Fig. 4 may be used with any embodiment described herein. However, one of ordinary skill in the relevant art will appreciate that different signaling schemes may be used.

[0090] The modules group 402 includes modules 404, 406, 408 that share a common write peripheral 410. The write peripheral 410 includes a write address bus that includes the address of the data being written, a write data bus that includes theAttorney Docket No. P25-059-SEC-W001data, a write clock cause the writes to occur (e.g., either on a leading or trailing edge of the clock signal, etc.). The writes only occur if the write enable signal indicates a write should occur. Any logic may be used, e.g., high voltage may correspond to 1 and a low voltage may correspond to 0, or vice versa. In some embodiments, the write peripheral 410 may be on the chiplet 230 and in other embodiments, the write peripheral 410 is on the second device 226.

[0091] The modules group 402 has modules 404, 406, 408 where each has a respective read peripheral 410, 412, 414. Each of the read peripheral 410, 412, 414 has a read address bus to send an address for reading, a read data bus to receive the data, a read clock which is the clock used to control the timing of the output of the digital data, and an output enable that is a precondition to outputting data. Any logic may be used, e.g., high voltage may correspond to 1 and a low voltage may correspond to 0, or vice versa. In yet additional embodiments, multi-bit or analog data storage may be used. In some embodiments, one or more of the read peripherals 410, 412, 414 may be on the chiplet 230 and in other embodiments, one or more of the read peripherals 410, 412, 414 are on the second device 226.

[0092] Fig. 5 shows an illustration of an integrated circuit 500 that may be part of a semiconductor device such as a chiplet in accordance with an embodiment of the present disclosure. The integrated circuit 500 may be disposed on a semiconductor device, such as a chiplet, that has a silicon substrate 506 and a second layer portion 508. Within the integrated circuit 500, there may be an array section forming the module 502 where a three-dimensional column array of memory bit cells 522 has the components to store memory in non-voltage, semi-volatile memory, or a memory format as described herein.

[0093] The integrated circuit 500 may include a modules group having a plurality of modules including a first module and a second module, etc. even though only a single module 502 is shown. The memory bit cells 522 are written to by the shared write port 512, 516, which includes both a write address bus line 512 and a write data bus 516. These buses run through the second layer 508 and can be connected to a second semiconductor device via an interposer. The second device has electrical contacts that complement those on the surface 518, allowing it to be electrically coupled to the write address and data buses. The memory bit cells 522 can be read from via the read port 524, 526, which includes a read address bus line 524 and a read data bus 526. Both of these buses can also run through the second layer 508 to the secondAttorney Docket No. P25-059-SEC-W001semiconductor device coupled to the surface 518, which also has complementary electrical contacts to allow it to be electrically coupled to the read address and data buses.

[0094] Various kinds of memory technologies may be used for the memory bit cells 522, such as a vertical connectivity fabric structure formed from non-volatile memory unit cells arranged in a three-dimensional column array 522. The memory bit cells 522 may utilize one or more of a cross-point, 3D NANDs, 3D NORs, 3D ANDs, and / or a stacked planar layer.

[0095] In some embodiments, the integrated circuit 500 is electrically connected to a second semiconductor device (not shown in Fig. 5) comprising another integrated circuit, which may be a system-on-chip or a Field-Programmable-Gate-Array. In some embodiments, the memory bit cells 522 may be formed from various non-volatile memory types, such as FeFET, FeRAM, ReRAM, SOT, or STT. Additionally, alternatively, or optionally, the memory bit cells may be formed from non-volatile memory unit cells having 2-terminal devices, 3 -terminal devices, or 4-terminal devices.

[0096] For example, the memory unit bit cells 522 may be formed from ferroelectric materials, such as a ferroelectric tunnel junction, a diode, a capacitor, a single -gate transistor, or a dual-gate transistor. Alternatively, the memory unit bit cells 522 may be formed from memristive materials, such as at least one ReRAM, or magnetic materials, such as at least one spin-orbit-torque device or at least one spintransfer-torque device. Moreover, the non-volatile memory unit cells 522 may also be formed from phase-change materials or anti -ferroelectric materials.

[0097] In some alternative embodiments, the non-volatile memory unit cells522 can be formed from other types of materials, such as phase change materials, antiferroelectric materials, or multi-bit PCM materials. The non-volatile unit cells can be formed utilizing different structures, such as resistive random-access memory (RRAM) technology, magnetic random-access memory (MRAM) technology, or ferroelectric random-access memory (FRAM) technology.

[0098] Moreover, in some implementations, 3D NAND technology may be utilized to form the memory unit bit cells 522. For example, the memory unit bit cells 522 may be formed from stacked memory layers where each layer includes a plurality of memory cells that can be accessed using shared bit lines. In such a case, the read portAttorney Docket No. P25-059-SEC-W001524, 526 may be coupled to the bit lines, and the write port 512, 516 may be coupled to the word lines that control the access to each layer.

[0099] In another embodiment, the 3D connectivity fabric structure can be built with stacked layers of either NAND gates, NOR gates, or AND gates, and in some cases, different types of logic gates may be combined to optimize the structure's functionality. In addition, the 3D connectivity fabric structure may be formed utilizing through-silicon-via (TSV) technology, which allows the vertical interconnection of the different layers of the structure.

[0100] Additionally, the non-volatile memory unit cells may include 2-terminal devices, such as a capacitive or a memristive device with or without an additional selector device such as a diode in series, 3 -terminal devices, such as a floating-gate transistor, a transistor with an access gate, or 4-terminal devices, such as a transistor with two access gates. The type and configuration of the non-volatile memory unit cells 522 may depend on the specific application requirements, including the speed, power consumption, and reliability of the circuit. The memory unit cell may include or be a single ferroelectric transistor or 6T SRAM cell. The memory unit cell may be a combination of many different devices, including, but not limited to, one or more of a transistor, a memristor, a capacitor, etc.

[0101] In some embodiments, a ferroelectric material can be utilized to form the non-volatile memory unit cells 522. The ferroelectric material may be implemented as any kind of device, including, but not limited to, a thin-film device, such as a ferroelectric tunnel junction, a capacitor, a single-gate transistor, or dual-gate transistors, etc.

[0102] In another embodiment, the non-volatile memory unit cells522 may be formed from a memristive material, such as a Metal Oxide Memristor (MOM), Conductive-Bridging RAM (CBRAM), or valence change memory (VCM), each of which provides different benefits regarding power consumption, speed, endurance, etc.

[0103] Moreover, in some embodiments, the non-volatile memory unit cells 522 may be formed from a magnetic material, such as spin-orbit-torque (SOT) devices, spin-transfer-torque (STT) devices, or perpendicular magnetic tunnel junctions (p-MTJ).

[0104] In one embodiment, the modules group may include many modules where each of which can be accessed through dedicated read ports 524, 526 with a dedicate read peripheral 520 while sharing the same write port 512, 516 and sharedAttorney Docket No. P25-059-SEC-W001write peripheral 510. The shared write port 512, 516 can be configured to selectively write to one or more of the plurality of modules within the modules group including the memory bit cells 522. Each of the modules may have the same or different sizes, and different module sizes may be configured to optimize the utilization of the memory array with different operating scenarios, etc.

[0105] Furthermore, the integrated circuit 500 may be formed utilizing different manufacturing processes and techniques, which include but not limited to, a CMOS or Bipolar-CMOS-DMOS (BCD) process, a silicon-on-insulator (SOI) process, a FinFET process, a silicon germanium (SiGe) process, a gallium arsenide (GaAs) process, etc.

[0106] In some embodiments, the where a three-dimensional column array of memory bit cells 522 forms is configured as a microvault. Additionally or alternatively, each of the micvrovaults will have a dedicated read periphery 520 and a write periphery 510 that is also a dedicated write periphery rather than a shared write periphery. That is, in some embodiments, each mircovault includes a dedicate write connection and a dedicated read connection, predetermined number of microvaults (e.g., 2 or 4), may have a dedicated write connection and a dedicated read connection, with or without dedicated respective peripheries, etc.

[0107] Fig. 6 shows a perspective of an assembly 600 having the integrated circuit of Fig. 1 implemented on a semiconductor device, such as the chiplet 230, that is electrically connected to a system-on-a-chip (“SOC”) 610 in accordance with an embodiment of the present disclosure. The semiconductor device, in this embodiment, is the chiplet 230 that is electrically connected to a system-on-a-chip (“SOC”) 610.

[0108] Referring to Fig. 6, the SOC 610 includes a silicon substrate 602 on which a plurality of processing elements is formed, including a processing element 606. The processing elements can communicate with each other through a Network-on-Chip (“NOC”) 604, which is a communication fabric that directs data transfer between the processing elements. The communication fabric can take various forms, including buses, switches, NOCs, etc. The NOC 604 in the SOC 610 directs data traffic between the various nodes (e.g., the processing element 606) and links, which provide the communication paths between the nodes.

[0109] The plurality of processing elements including the processing element 606 processing elements can be any suitable type of processors capable of executing instructions, including microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), or application-specific integrated circuits (ASICs).Attorney Docket No. P25-059-SEC-W001

[0110] Additionally, the SOC 610 may comprise various modules, such as module 232, which are grouped together to provide memory functionality to the assembly 600 as described here. The modules in modules group 236 can be coupled to a respective processing element provide it readable memory. In some embodiments, the coupling between the module (e.g., module 232) and the processing element (e.g., 606) can be achieved through interconnects on the silicon substrate 602.

[0111] After the circuitry is formed on the silicon substrate 602, a second layer 608 can be disposed on top of the substrate. The second layer 608 can be any suitable material, such as an insulating material, a metal, a dielectric, or interconnect layer, and it may be bonded to the chiplet 230. The bonding can be done using any suitable technique, including but not limited to, adhesives, soldering, or welding, etc.

[0112] In general, the assembly 600 provides a means of integrating the chiplet 230, which can include the integrated circuit of Fig. 1, with the SOC 610. Integrating the chiplet 230 provides various advantages, such as enhanced functionality, higher performance, and lower power consumption. Moreover, the integration of the chiplet 230 with the SOC 610 can be accomplished in various ways, depending on the particular application and design objectives of the system.

[0113] The assembly 600 can incorporate various variations and modifications, depending on the specific requirements of the system. For example, the processing elements formed on the silicon substrate 602 can vary in their number, type, and arrangement. Similarly, the modules in modules group 236 can vary in their number, type, and function.

[0114] Furthermore, the second layer 608 can be modified to include additional functionality. For instance, the second layer 608 can include passive components, such as resistors, capacitors, and inductors, or active components, such as transistors or diodes. Incorporating these components in the second layer 608 can further enhance the functionality and performance of the system.

[0115] In another variation, the assembly 600 can incorporate a heterogeneous integration approach, where the chiplet 230 is fabricated using a different technology than that used for the SOC 610. This approach allows for the optimal use of different fabrication technologies for different parts of the system, resulting in improved performance and reduced power consumption.

[0116] Fig. 7 shows a perspective view of an assembly 700 having a semiconductor device 707 with an array of processing elements 706 (on a grid ofAttorney Docket No. P25-059-SEC-W001processing element 706a, a to 706n,n, where the first subscribe is the columns and the second subscript is the row), and a second semiconductor device 709 having an array of microvaults 708. The array of microvaults 708 are on a grid of microvaults 708a, a to 708n,n, where the first subscript is the columns and the second subscript is the row). These subscripts may line up such that a respective subscript of a processing element 706 corresponds to a respective subscript of a microvault 708. The microvaults 708 are a type of module described herein where it is positioned in a vertical direction, e.g., above a respective processing element 706. The semiconductor device 707 may be a chiplet. Also, the semiconductor device 709 may also be a chiplet. The chiplets 707,709 may be bonded together. The different layers 710 (e.g., 710a through 710d can correspond) may be allocated for separate Al models (e.g., parameters in a neural network, such as a CNN or transformer model).

[0117] The assembly 700 is the overarching structure that houses the various components shown in Fig. 7. It provides mechanical support and integration for the other elements, allowing them to function as a unified system.

[0118] The assembly 700 includes two semiconductor devices - the semiconductor device 707 and the semiconductor device 709. The semiconductor device 707 contains an array of processing elements labeled 706a, a to 706n,n. Similarly, the semiconductor device 709 contains an array of microvaults labeled 708a, a to 708n,n.

[0119] The subscripts a, a to n,n indicate that the processing elements 706 and microvaults 708 are arranged in a grid pattern, with the first subscript referring to the column and the second subscript referring to the row. This grid arrangement allows each processing element 706 to have a corresponding microvault 708 positioned vertically above it. For example, processing element 706a,a has microvault 708a, a above it, processing element 706b, b has microvault 708b, b above it, and so on. The alignment of the grid allows tight integration between the processing and storage components.

[0120] In some embodiments, the semiconductor devices 707 and 709 are potentially separate chiplets that are integrated using packaging techniques into the unified assembly 700. The chiplet form factor allows greater flexibility and customization in assembling the system. This arrangement is such that each processing element 706 can access its respective microvault 708 located above it to retrieveAttorney Docket No. P25-059-SEC-W001relevant data, such as weights for neural networks or Al models. This may provide high bandwidth and low latency access to the data needed for efficient processing.

[0121] Input data enters the system via the input DRAM memories 702. This data flows into the processing elements 706, where it is operated on locally using weights or parameters from the vertically integrated microvaults 708. The processing results output via the output DRAM memories 704. The input DRAM memories 702 consist of multiple individual DRAM modules labeled 702a, 702b, and 702c. The DRAM memories 702 can be any type of dynamic random access memory, including but not limited to DDR SDRAM, LPDDR SDRAM, GDDR SDRAM, and HBM. The DRAM memories provide high-bandwidth data input capabilities to feed data, such as inference inputs or training data, into the processing pipeline.

[0122] In some embodiments, each individual DRAM module 702a, 702b, and 702c has a dedicated interface and data path to each processing element 706. For example, DRAM module 702a may feed data only to processing element 706a, a, while DRAM module 702b feeds data only to processing element 706b, b. This provides modular scalability, as additional DRAM modules can be added to feed more processing elements.

[0123] The number of input DRAM memories 702 and individual modules 702a-702c may vary depending on the application requirements. For instance, there could be 4, 8, 16, or more input DRAM modules. The capacity of each module can range from gigabytes to terabytes depending on factors such as access speed, power, and cost budget.

[0124] High-speed interfaces like DDR5, GDDR6, or HBM3 may be used to maximize data transfer bandwidth between the input DRAM memories 702 and the processing elements 706 across the semiconductor device 707. Shared data buses, crossbar switches, or on-chip networks may interconnect groups of DRAM modules 702 and processing elements 706.

[0125] In some implementations, the input DRAM modules 702 may be stacked or arranged in a multi-dimensional configuration to increase overall memory capacity and bandwidth while reducing latency and power consumption. Specialized memory controllers and schedulers may manage parallel data access across multiple input DRAM modules 702.

[0126] The input DRAM memories 702 supply the high-bandwidth data needs of the parallel processing elements 706, enabling fast and efficient data-intensiveAttorney Docket No. P25-059-SEC-W001computations such as neural network inferencing. Each processing element 706 can directly access the required input data from its dedicated DRAM module 702 without contending with other processors for data access.

[0127] In alternative embodiments, instead of, or in additional to, the DRAM memories 702, 706, adjacent accelerator chiplets may be in communication with the semiconductor device 707. That is, there may be a grid-like arrangement of assemblies 700 in communication with each other to perform Al inference and / or Al training (e.g., transformer inference, CNN interference, ANN interference, etc.). In some embodiments, there may be clusters of semiconductor devices 707 that share a bank or portion of DRAM memories 702 and / or 706. In some embodiments, the input DRAM memories 702 and the output DRAM memories 704 may be combined into the same DRAM memory.

[0128] The assembly 700 includes a semiconductor device 707 that comprises an array of processing elements labeled from 706a, a to 706n,n. Each processing element in the array may be configured to execute specialized computations and data processing operations. For instance, in some embodiments, the processing elements could be optimized for artificial intelligence workloads like neural network inference. In other cases, the processing elements may focus more on general-purpose capabilities. Ultimately, the capabilities of each processing element depend on its specific microarchitecture which can be tailored for certain applications if desired.

[0129] The processing elements 706 can access nearby memory storage to retrieve data that feeds into their computations. This memory may be physically separate from the processing element arrays 706 as is the case with the microvaults 708 shown in Figure 7. The processing elements 706 and microvaults 708 are aligned so that each microvault is positioned directly above its corresponding processing element in a vertical configuration. This tight coupling provides fast data transfer speeds between each vault-element pair.

[0130] In terms of physical implementation, the array of processing elements 706 resides within the semiconductor device 707. The semiconductor device 707 could potentially be manufactured as a standalone chiplet using advanced packaging techniques. This modular chiplet can then be integrated with other components like the microvault chiplet 709 through high-density interconnections. Some options include bumpless hybrid bonding, interposers, or even monolithic 3D integration. Ultimately,Attorney Docket No. P25-059-SEC-W001combining chiplets allows creating powerful heterogenous systems with optimized dies.

[0131] The specific number, design, and interconnect scheme of processing elements 706 present can vary between implementations of assembly 700. For instance, simpler systems may need only a 2x2 grid of elements whereas a sophisticated Al accelerator could feature a 32x32 array. The processing elements 706 themselves can also have different memory access routes across the assemblies. Point-to-point links, crossbar switches, or shared buses are possible connection structures. Such architectural decisions depend on the performance and area constraints trying to be met.

[0132] The second semiconductor device 709 is a separate device from the first semiconductor device 707. Like the first semiconductor device 707, the second semiconductor device 709 may also be implemented as a chiplet. The second semiconductor device 709 includes an array of microvaults 708, arranged on a grid spanning from microvault 708a, a to 708n,n.

[0133] As mentioned, the microvaults 708 on the second semiconductor device 709 are positioned vertically above the processing elements 706 on the first semiconductor device 707. Each microvault 708 lines up with and corresponds to the processing element 706 underneath it, based on the subscripts identifying their position in the grid. For example, microvault 708a, a is vertically aligned with and corresponds to processing element 706a, a. This allows each processing element 706 to access the microvault 708 above it.

[0134] The microvaults 708 act as a memory structure, storing things like Al model weights that can be accessed by the processing elements 706 underneath during operations like neural network inference. The microvaults 708 may be optimized for very fast read times but slower write times. This allows the processing elements 706 to rapidly access the weights and data needed for their computations, while less frequently updated data can still be written at a slower pace.

[0135] In some embodiments, the second semiconductor device 709 containing the array of microvaults 708 is directly bonded to the first semiconductor device 707 with the processing elements 706. This bonding aligns each microvault 708 with its corresponding processing element 706 underneath. Electrically conductive interconnects between the devices allow each processing element 706 to communicate directly upwards with its respective overlying microvault 708. This provides a compact, modular, and efficient system architecture.Attorney Docket No. P25-059-SEC-W001

[0136] The microvaults 708 may contain multiple memory layers, labeled 710a to 710d, each storing weights or data for a different Al model. For example, layer 710a contains the weights for model A, layer 710b contains the weights for model B, and so on. Stacking these layers vertically contributes to the high density and fast access times of the microvault design.

[0137] The microvaults 708 can be implemented using various memory technologies, including but not limited to SRAM, FeFET, ReRAM, SOT, and STT, optimized for fast readout times to supply data to the processing elements 706 with minimal latency. Specific embodiments may configure the microvaults 708 to have much faster read speeds compared to their write speeds. The microvaults 708 may be implimented using any FeFET or memory structure described herein.

[0138] In some embodiments, each microvault 708 may have a capacity between 4 kilobytes to 128 kilobytes for storing parameters for machine learning models or other data. The bit density per layer may exceed 0.4 gigabits per square millimeter. The microvaults' 708 compact size, between less than 100 micrometers on each side in one embodiment or 12 micrometers by 12 micrometers in another, allows high-density integration of the memory modules.

[0139] The array -based arrangement of the microvaults 708 may enables concurrent parallel data access by the processing elements 706, supporting high-throughput data processing by the assembly 700. The one-to-one alignment of the microvaults 708 and processing elements 706 also ensures that each processing element 706 has dedicated access to its required data without contention.

[0140] The microvaults 708 share the semiconductor device interface provided by the second semiconductor device 709, which facilitates writing data to the microvaults 708 from the input DRAM memories 702. Reading data from the microvaults 708 to the processing elements 706 and output DRAM memories 704 is handled through dedicated pathways between each vertically aligned microvault 708 and processing element pair.

[0141] The micro vault memory layers 710 refer to multiple layers of micro vault memories stacked vertically within the second semiconductor device 709. As illustrated in Fig. 7, there are four separate microvault memory layers labeled as 710a, 710b, 710c, and 710d. Each layer contains an array of microvaults, such as the array of microvaults 708 shown in the diagram.Attorney Docket No. P25-059-SEC-W001

[0142] In one embodiment, the microvaults 708 may utilize a stacked 3D NAND architecture built from multiple layers of NAND memory arrays using charge trap flash technology. Each microvault 708 may contain a dedicated set of wordline drivers on the bottom layer to facilitate access to the 3D NAND cell arrays above, spaced by alternating dielectric layers. The 3D NAND implementation may be used to maximize density and throughput by leveraging vertical scaling.

[0143] In an alternative embodiment, the microvaults 708 employ a 3D NOR architecture constructed from multiple tiers of NOR flash memory arrays. Each plane features NOR strings with a source line and bit line architecture, stacked on top of each other using vias. The 3D NOR arrangement optimizes random read access times to stored data.

[0144] The microvaults 708 may also adopt a hybrid configuration with different types of volatile and / or non-volatile memory, such as combining FeRAM and ReRAM cells, organized into vertical sub-arrays. This heterogeneous 3D integration allows optimizing for speed, endurance, and retention within the same vault structure.

[0145] In certain embodiments, the microvaults 708 integrate processing logic like analog computing directly into the memory array stack itself. This processing-in-memory approach places basic computational operators within the memory periphery or bit cells, enabling highly parallel and efficient in-situ data processing.

[0146] Some implementations may utilize 2.5D or 3D stacking to integrate the microvaults 708 with other components like logic, CPUs, GPUs or application-specific accelerators. This tight packaging integration via techniques like high-bandwidth memory cube architectures reduces data transfer latency and power consumption.

[0147] The microvaults 708 may also employ a virtualized architecture, with an external memory controller handling translation between the physical array organization and dynamically allocated virtual memory domains. These virtual domains mapped onto the physical array effectively creates separate virtual vaults with flexible capacities tailored to application needs.

[0148] In certain embodiments, the microvaults 708 are designed as Computational RAM (CRAM) with integrated processing capabilities within the bit cell periphery to enable highly parallel in-memory computing architectures. Gateless transistor structures integrated into the CRAM arrays facilitate efficient execution of bulk bitwise operations.Attorney Docket No. P25-059-SEC-W001

[0149] Some implementations arrange the microvaults 708 into modular Memory Processing Unit (MPU) structures containing dedicated processing logic tailored for workloads like Al inferencing. The MPU architecture couples vault arrays to vector processors via high-speed interfaces like HBM2 enabling low-latency data transfers.

[0150] The microvaults 708 may also implement content-addressable capabilities by integrating comparison logic into the memory periphery. This facilitates searching or accessing data based on content rather than explicit addresses, enabling powerful pattern matching capabilities.

[0151] Certain embodiments may stack multiple microvault dies on top of base logic dies featuring things like GPUs or Al accelerators. This creates dense, high-bandwidth heterogeneous systems optimized for data-centric workloads while minimizing data movement.

[0152] The microvault memory layers 710 may be fabricated utilizing three-dimensional integrated circuit manufacturing processes to stack multiple dies or wafers containing microvault 708 arrays on top of each other. Through-silicon vias (TSVs) or other vertical interconnect technologies can be employed to enable communication between the layers.

[0153] In some embodiments, each microvault memory layer 710 corresponds to a different artificial intelligence (Al) model or application. For example, layer 710a could store the weights and parameters for Al model A, layer 710b could store the weights and parameters for Al model B, and so on. This allows multiple Al models to be stored efficiently within the same microvault memory 708 structure.

[0154] The microvaults 708 may possess capacities ranging from 4 kilobytes to 128 kilobytes in some embodiments. In other cases, the capacity could be between 4 kilobytes to 16 kilobytes. Each microvault could have lateral dimensions less than 100 micrometers by less than 100 micrometers, while extending vertically to incorporate potentially over 200 memory cell layers in some implementations.

[0155] In some embodiments, the bit density of per square millimeter per layer within the microvault memory layers 710 may facilitate high-capacity storage with a small footprint. The layers may utilize non-volatile memory technologies, such as FeFET, STT-MRAM, or ReRAM, to retain data when power is removed.Attorney Docket No. P25-059-SEC-W001

[0156] In operation, the processing elements 706 may access weights or parameters from the microvault memory layers 710 to perform neural network inferencing or other machine learning computations.

[0157] The inference results of the various Ais may be sent to the output DRAM memories 704. which comprise individual DRAM memory modules labeled 704a, 704b, and 704c. The output DRAM memories 704 are positioned adjacent to the array of microvaults 708 and the second semiconductor device 709. The output DRAM memories 704 may serve as temporary data storage that can buffer output data retrieved from the microvaults 708 before it is transmitted externally.

[0158] Each DRAM memory module 704a, 704b, and 704c may have similar or different storage capacities, depending on the design requirements. For example, in one embodiment, each module contains 16 megabits of storage. The DRAM storage cells utilize a capacitor to retain data bits in the form of electrical charges. Due to charge leakage, the DRAM memories require periodic refresh cycles to maintain the stored data integrity. To enable concurrent reads and writes across multiple modules, each DRAM module 704a, 704b, and 704c can have dedicated internal control circuitry and I / O ports.

[0159] The data outputs from the individual microvaults 708 may get aggregated and buffered in the output DRAM memories 704 before being transmitted to external components via peripheral circuitry. Buffering the data allows the transmission rate to be regulated to match the requirements of the external interfaces. It also enables data processing operations like formatting, encoding, or encryption to be performed by the second semiconductor device 709 prior to output.

[0160] In some embodiments, the output DRAM modules 704a, 704b, and 704c are designed to provide high-density, low-cost temporary data storage to support the high-bandwidth parallel reads from the array of microvaults 708. Optimizing these performance parameters allows efficient extraction of data from the microvaults to feed the computational workflows hosted on external chips or devices. Specific implementations may utilize various types of DRAM, including asynchronous DRAM, synchronous DRAM, graphics DRM, and low-power DRM tailored to the application. Overall, the output DRAM memories 704 facilitate seamless data movement from the integrated microvaults to external execution pipelines.

[0161] Fig. 8 shows an assembly 800 of a semiconductor devices 802, 804, 806, 808, 810, 812, 814 including several memory types in accordance with an embodimentAttorney Docket No. P25-059-SEC-W001of the present disclosure. Specifically, Fig. 8 shows an exemplary assembly 800 of semiconductor devices 802, 804, 806, 808, 810, 812, 814 configured to provide a hierarchical memory structure. The assembly 800 is modular and scalable, allowing for various combinations and numbers of semiconductor devices, which may be implemented as chiplets in certain embodiments, to be stacked to meet specific performance and density requirements.

[0162] The assembly 800 includes an application semiconductor 802, that may be a plurality of processing elements as described herein. On top of the semiconductor device 802 is semiconductor device 814 that includes an array of microvaults. One top of the semiconductor device 814 is semiconductor device 812, which may also include an array of microvaults. On top of the semiconductor device 812, are semiconductor devices 810, 808, which may be SRAM vaulted dies. One top of the semiconductor device 808, is semiconductor devices 806, 804 may be DRAM vaulted dies. These vaults 816 may be arranged in a grid-like fashion such that 816a, a to 816n,n subscripts the vaults. Each of these vaults may include a respective microvault from semiconductor devices 804, 812, respective SRAM vaults from semiconductor devices 810, 808, and respective DRAM vaults from semiconductor devices 806, 804.

[0163] At the base of assembly 800 lies the semiconductor device 802, which comprises a plurality of processing elements. These processing elements execute computational tasks and facilitating data flow within the system.

[0164] Directly above semiconductor device 802 is semiconductor device 814, which includes an array of microvaults. These microvaults utilize Field-Effect Transistors (FeFETs) known for their non-volatile characteristics and suitability for high-density memory applications. The FeFET-based microvaults may be designed to enable high-speed read operations essential for rapid data retrieval during processing tasks, such as Al inferencing, while supporting slower write operations that are more tolerant to latency. Stacked on top of semiconductor device 814 is semiconductor device 812, which similarly includes an array of microvaults. The presence of multiple layers of microvaults in semiconductor devices 814 and 812 exemplifies the scalable nature of the assembly, where additional memory capacities and functionalities can be integrated through additional layers.

[0165] Further contributing to the memory hierarchy, semiconductor devices 810 and 808, positioned above semiconductor device 812, are depicted as SRAMAttorney Docket No. P25-059-SEC-W001vaulted dies. SRAM provides fast access memory that can serve as a cache or buffer to the slower, but denser, FeFET microvault memory layers beneath.

[0166] At the top of the assembly 800 and hence the memory structure are semiconductor devices 806 and 804, illustrated as DRAM vaulted dies. DRAM is typically used for main memory due to its relatively high speed and low cost per bit compared to SRAM, offering a balance between performance and economy.

[0167] The arrangement of vaults 816 in a grid-like fashion, subscripted from 816a, a to 816n,n, indicates that each processing element at a given location (a,b) in semiconductor device 802 has dedicated access to the corresponding vertically aligned vaults of microvaults and memory cells in the layers above. This vertical stacking and alignment ensure that data and control signals can be directly routed between processing elements and their respective memory stacks, facilitated by interconnect technologies such as through-silicon vias (TSVs) and micro-bumps, which are sued in the assembly's 8003D integrated circuit architecture.

[0168] The modular and scalable design of assembly 800 allows for various combinations of semiconductor devices or chiplets to be integrated into more extensive systems. The flexibility in the number and combination of stacks provides the adaptability to tailor the assembly to the requirements of different applications and performance demands. Each vault within the vaults 816 within the assembly 800 presents a multi-die structure that contributes to the overall capacity and performance of the system.

[0169] Fig. 9 shows an assembly of semiconductor devices including a semiconductor device with a system-on-chip 914 and another semiconductor 906 with microvaults disposed on top in accordance with an embodiment of the present disclosure. The assembly 900 integrates a semiconductor device 906 and a semiconductor device 914, which may be implemented as separate chiplets bonded together. The semiconductor device 914 includes various components to facilitate reading data from the microvaults on the semiconductor device 906, such as a read address register input interconnect 924, read data register 922, and read data register output interconnect 950. These components pass the read address to the microvaults on semiconductor device 906 and return the read data back to the semiconductor device 914.

[0170] Specifically, the read address enters via interconnect 924 into the read address register 926. The output of this register 962 connects through interconnects andAttorney Docket No. P25-059-SEC-W001bumpless bonds to another read address register 938 on the semiconductor device 906, which then addresses the target microvault 936. The microvault 936 outputs read data via interconnect 940 to a read data register 942, which passes the data back through bumpless bonds 910, 918 to read data register 922 on semiconductor device 914. This data can then be accessed externally via the read data register output interconnect 950. Additionally, the semiconductor device 914 and 906 have interconnected Through-Silicon Vias 916 and 944 to allow communication with devices potentially stacked above semiconductor device 906.

[0171] The semiconductor device 906 features various memory structures to provide data storage capabilities. This includes a microvault 936, which offers high-density, low-latency data storage, along with other peripheral memory components like read data register 942 and read address register 938 to facilitate data reads. The microvault 936 resides on the BEOL portion of the chiplet, allowing dense 3D integration of memory layers. In some implementations, the microvault utilizes nonvolatile memory technologies like FeFET or STT-MRAM for data retention without power.

[0172] The semiconductor device 914 comprises processing elements and data routing circuitry to retrieve and manipulate data stored in semiconductor device 906. Components like read address register 926 and read data register 922 handle sending read addresses and receiving data from the microvault 936 respectively. The device 914 also includes interconnects 924, 950 and Through-Silicon Via 916 to communicate externally.

[0173] The two devices 906 and 914 integrate via fine -pitch interconnects like bumpless hybrid bonds 908, 910, 918, 920, 930 and 932. This allows direct data transfer pathways between processing components in device 914 and memory structures in device 906. Alignment during bonding ensures dedicated access - for instance, read data register output interconnect 950 on 914 links directly to read data register 922 to receive requested data.

[0174] The pathways facilitating data flow during reads can be summarized as follows: A read address enters through interconnect 924 into read address register 926 on device 914. This gets communicated via interconnects and bumpless bonds to read address register 938 on device 906, which then addresses microvault 936. Requested data gets passed via interconnect 940 to read data register 942, then transfers throughAttorney Docket No. P25-059-SEC-W001bonds back to read data register 922 on 914, where it becomes available externally via interconnect 950.

[0175] The assembly 900 exemplifies a modular, high-density architecture optimized for data-centric applications like Al inferencing. Tight integration of processing and storage dies via advanced packaging techniques allows localized data access with minimal latency and power. Scalability is also enabled by incorporating multiple chiplets, in this case devices 906 and 914. The assembly 900 illustrates a potential configuration suited for space-constrained, high-performance computing systems.

[0176] The Through-Silicon Via (TSV) 916 is an electrical connection that passes vertically through the semiconductor device 914. Its purpose is to provide a pathway for signals to travel between the top and to a processing element within the semiconductor device 914. This allows the device to be stacked and interconnected with other components in a vertical configuration. The TSV 916, along with other TSVs on the device, facilitates high-density 3D integration and heterogeneous stacking of multiple devices like chiplets.

[0177] The TSV 916 interacts with several other components within the system. On the top side of semiconductor device 914, it connects to interconnect 912, which couples it to bumpless bonds 918. These bonds interface with complementary bumpless bonds 910 on the bottom side of semiconductor device 906 when the two devices are stacked. This allows signals to travel from device 914 to device 906 through the TSV 916. The route continues as signals go through interconnect 902 to TSV 944 on device 906. TSV 944 provides a vertical signal pathway to the top surface of device 906 where additional devices could be stacked. In the reverse direction, signals can travel from TSV 944 down through device 906, back up TSV 916, and down into device 914. So the TSV 916 provides bidirectional vertical communication across device boundaries.

[0178] There are a few possible variations for the TSV 916 implementation, first, multiple TSVs arranged in an array could be used instead of a single via to increase throughput and redundancy. Second, the dimensions and materials of the TSV could be optimized- for example, smaller TSV diameters using denser materials like tungsten could be advantageous. Additionally, the interface circuitry driving signals into the TSV, like interconnects 912 and 902, could employ variable line drivers to support different voltage levels or signal integrity enhancements. Further embodiments may include integrated monitoring circuitry within TSV 916 to track metrics likeAttorney Docket No. P25-059-SEC-W001temperature and link utilization. And alternative signaling schemes besides electrical signals could be employed in future cases. For instance, integrated silicon photonics utilizing modulated light to convey data through the TSVs could enable very high bandwidth and low latency connectivity. There are multiple avenues to further develop the capabilities of TSV-based vertical links like TSV 916 within these complex 3D integrated architectures.

[0179] There are several variations and alternatives for the interconnect 912 implementation. For example, different conductive materials such as copper or aluminum may be utilized to fabricate the pathways forming interconnect 912 and optimize for conductivity or thermal dissipation. Additionally, interconnect 912 may feature redundant signal paths or self-repair capabilities using spare interconnect Ones to improve reliability and resilience. The bumpless bonds 918 and 910 connecting devices 906 and 914 could also be replaced with other high-density bonding approaches like hybrid bonding or Through-Silicon Vias. Furthermore, alternate signaling schemes besides simple digital logic could be employed on interconnect 912, such as analog signaling or multi-level digital waveforms to enhance data transmission capabilities. The routing and dimensions of interconnect 912 can also be adapted according to bandwidth requirements or circuit layout considerations. Overall, many structural and functional alternatives exist for crafting interconnect 912 to meet application needs.

[0180] The bumpless bonds 918 are electrical connections located on the semiconductor device 914 between an interconnect 912 and bumpless bonds 910 of the semiconductor device 906. The bumpless bonds 918 provide an electrical pathway for signals to travel between the semiconductor device 914 and any additional semiconductor devices, such as the semiconductor device 906, stacked on top of the assembly 900. The signals communicated over the bumpless bonds 918 can include data signals, control signals, address signals, or any other signals needed to coordinate operations between the multiple semiconductor devices.

[0181] There are several possible variations for the bumpless bonds 918. The number of individual bond sites can range from just a few to hundreds, depending on signal bandwidth requirements. The bonding method can utilize techniques like direct bonding, plasma-activated bonding, adhesive bonding, or compression bonding. Hybrid bonding approaches are also possible, combining direct wafer bonds with intermediate metal bonds. The size and pitch of each bond site can vary and may use pitches under 10 micrometers to enable high-density connections. Redundant bonds can provideAttorney Docket No. P25-059-SEC-W001backup pathways. Shielding structures may surround bonds for noise immunity. Overall, many embodiments of bumpless bonds 918 are possible to meet cost, reliability, and performance needs.

[0182] The bumpless bonds 910 provide an interface for communicating signals between the semiconductor device 906 and semiconductor device 914. Specifically, the bumpless bonds 910 of the semiconductor device 906 are electrically coupled to the complementary bumpless bonds 918 of the semiconductor device 914. This allows signals like read / write data and addresses to be transmitted between the two devices. The bumpless nature of the bonds allows for a low-profile, high-density interconnection.

[0183] The bumpless bonds 910 interact with other components in the system to facilitate data transfer operations. For writes, data enters the semiconductor device 914 via the Through-Silicon Via 916, passes through interconnect 912 and bumpless bonds 918 before reaching bumpless bonds 910 of device 906. For reads, addresses flow from the read address register 926 of device 914 through interconnects 928, 930 and bumpless bonds 932 into the read address register 938 on device 906. Read data then returns through bumpless bonds 908 and 920 back to device 914. So the bumpless bonds 910 provide key data and address routing between the devices.

[0184] Possible variations of the bumpless bonds 910 include using different bond densities, materials, or electrical contact configurations to optimize performance. The bonds can use alloying or doping techniques to improve conductivity. Additionally, the routing of signals can be changed, for example by using separate ports for input and output instead of shared ports. More bumpless bonds can be added to increase bandwidth between devices. Shielding may be added around the bonds to reduce interference. Overall, many modifications to the bumpless bonds 910 are possible within the scope of electrically interconnecting multiple devices.

[0185] The Through-Silicon Via (TSV) 944 is an electrical connection that passes vertically through the semiconductor device 906 from the top surface to the bottom surface. Its purpose is to facilitate communication of signals and data between the semiconductor device 906 and any additional semiconductor devices potentially stacked on top of it in a 3D integrated circuit configuration. The TSV 944 enables high-density interconnections between multiple stacked semiconductor layers, providing an efficient means for data routing and signaling.Attorney Docket No. P25-059-SEC-W001

[0186] The TSV 944 interfaces with surrounding circuitry within the semiconductor device 906, allowing signals to be transmitted upwards or downwards depending on the system configuration. On one end, the TSV 944 couples to the read data register 942 via interconnect 946. The read data register 942 can use the TSV 944 path to transfer read data from the microvault 936 to external semiconductor devices. This enables efficient data offloading from the on-chip memory. On the other end, the TSV 944 continues through to the top surface of semiconductor device 906, where it may interface with complementary contacts or interconnects on the bonded semiconductor above it. This facilitates the vertical transfer of signals and data along the assembly 900.

[0187] There can be many variations in the specific implementation of the TSV 944. Its dimensions can range from a few microns to tens of microns to match pitch requirements. The TSV 944 can be tapered, straight, or have non-uniform crosssections. It may utilize different conductive materials as liners and fills, including metals like copper, tungsten or alloys. Insulating liners made of materials like silicon dioxide can separate the conductive fill from the substrate. The contacts and interconnects coupling into the TSV 944 can also have diverse layouts. Multiple TSVs can be placed adjacent to each other in a high-density array configuration if desired. Overall, many architectural optimizations in the design and fabrication process of the TSV 944 are possible within the scope of the present disclosure.

[0188] Fig. 10 shows a semiconductor assembly 1000 incorporating a daisy-chained configuration of microvaults 1036, 1058, operatively connected to a multiplexer 1060 and managed by a counter 1062 for coordinated data selection and retrieval, in accordance with an embodiment of the present disclosure. This assembly 1000 is designed to carry out data processing tasks, potentially for applications such as artificial intelligence (Al) and machine learning, where high-speed data access and processing are utilized.

[0189] The assembly 1000 comprises two primary semiconductor devices: semiconductor device 1006 and semiconductor device 1014. Semiconductor device 1014 is depicted as containing several interfaces and registers for data communication, including a read address register input interconnect 1024. This interconnect 1024 facilitates the delivery of read addresses to a read address register 1026, which temporarily holds these addresses before they are transmitted to corresponding microvaults 1036, 1058 in semiconductor device 1006 for data retrieval operations.Attorney Docket No. P25-059-SEC-W001

[0190] In semiconductor device 1014, interconnect 1028 serves as a pathway for read addresses from the read address register 1026 to transition to bumpless bonds 1030. Bumpless bonds 1030 and 1032 represent high-density, low-profile electrical connections between semiconductor device 1014 and semiconductor device 1006, ensuring the transmission of read addresses with minimal signal loss and physical space requirements.

[0191] Through interconnect 1034, the received addresses reach the read address register 1038 in semiconductor device 1006, which then directs the microvault 1036 to output the requested read data. Microvault 1036, an memory storage unit, may encompass a variety of memory technologies, such as FeFETs and / or 3D-NAND structures as described herein, to facilitate the storage and rapid retrieval of data.

[0192] The multiplexer 1060 selects the appropriate data stream from multiple microvault 1036, 1058 outputs. Controlled by a counter 1062, which may operate according to a predefined sequence or be driven by external control signals, the multiplexer 1060 arbitrates between the outputs of micro vault 1036 and another microvault, denoted as microvault 1058. Microvault 1058, similar in function and potential memory technology to microvault 1036, provides an additional source of data for the multiplexer 1060 to select from.

[0193] Once the desired data is selected by the multiplexer 1060, it is temporarily stored in a read data register 1042, also located within semiconductor device 1006. This register 1042 acts as a buffer, holding the data for subsequent processing or transmission. The read data is then routed via interconnect 1004 to bumpless bonds 1008, which facilitate the transfer of data to the semiconductor device 1014.

[0194] Bumpless bonds 1010, 1018 facilitate the continued data's journey through the assembly 1000, ensuring data transfer from semiconductor device 1006 to semiconductor device 1014. Once the read data arrives at semiconductor device 1014, it is channeled via interconnect 1048 to a read data register, specifically read data register 1022, where it can be accessed by external systems, such as an applicationspecific integrated circuit (ASIC) or a system-on-chip (SoC), via the read data register output interconnect 1050.

[0195] Additionally, the assembly 1000 encompasses Through-Silicon Vias (TSVs) 1016 and 1044, providing vertical electrical connections through the semiconductor devices 1014 and 1006, respectively. These TSVs enable the stackingAttorney Docket No. P25-059-SEC-W001of additional semiconductor devices or chiplets atop the assembly 1000, thus allowing for vertical expansion of the system's capabilities. Interconnects 1012 and 1046 serve as horizontal pathways for signals to travel to and from the TSVs 1016 and 1044, respectively.

[0196] Although the description presents a specific configuration, the assembly 1000 may be subject to various modifications and alternative embodiments. For example, the number and arrangement of microvaults, the specific types of memory technologies employed within the microvaults, and the configuration of interconnects and bonding areas may be tailored to meet the requirements of different applications.

[0197] In some embodiments, the semiconductor devices 1006 and 1014 may be designed to accommodate additional functionality, such as thermal management layers for heat dissipation, hardware-based encryption modules for data security, or power management circuits to optimize energy consumption. The detailed structure of Fig. 10, therefore, serves as a foundation upon which a variety of sophisticated semiconductor systems can be constructed, each tailored to the specific needs of its intended application.

[0198] In the configuration of assembly 1000 as depicted in Fig.. 10, the microvaults, exemplified by microvaults 1036 and 1058, present a daisy-chaining configuration that allows for an expandable and flexible memory architecture within the semiconductor device 1006. This daisy-chaining is facilitated through a series of interconnected pathways and controlled by the multiplexer 1060 in coordination with the counter 1062.

[0199] Each microvault, such as 1036 and 1058, is designed to hold and provide rapid access to data, which may be in the form of stored charge, magnetic states, ferroelectric material states, or other physical embodiments of binary information. The microvaults are interconnected such that the output of one microvault can be routed to the input of another, creating a chain of memory elements. This is achieved through a series of interconnects, such as interconnect 1034 for microvault 1036 and interconnect 1056 for microvault 1058, which serve as conduits for the read data signals emanating from the micro vaults.

[0200] The multiplexer 1060 manages the flow of data from this daisy chain of microvaults. It is designed with multiple inputs, each connected to the output of a micro vault via respective interconnects. In the example provided, interconnect 1040 carries the read data from microvault 1036, and interconnect 1056 carries the read dataAttorney Docket No. P25-059-SEC-W001from microvault 1058 to the multiplexer 1060. The multiplexer 1060 is capable of selecting which input to connect to its output at any given time, thus controlling which microvault's data is forwarded to the read data register 1042.

[0201] The counter 1062 orchestrates the operation of the multiplexer 1060. It may be a binary counter or any form of sequential logic circuit that produces a series of output states in response to a clock signal. The counter 1062 progresses through its states with each tick of the clock, which may be provided by an external clock source or generated internally within the semiconductor device 1006. As the counter 1062 advances, it outputs a control signal that instructs the multiplexer 1060 on which input to select.

[0202] For instance, at the first clock pulse, the counter 1062 may instruct the multiplexer 1060 to connect the output from microvault 1036 to the read data register 1042. At the next clock pulse, the counter may switch the connection to microvault 1058's output, and so on, cycling through the available microvaults in a predefined order. The counter's sequence and timing can be configured based on the desired data access patterns and the specific requirements of the processing tasks at hand.

[0203] This clock-driven coordination allows for an efficient and organized retrieval of data from a potentially large array of microvaults. It ensures that each microvault has an equal opportunity to present its data for processing, and it simplifies the control scheme by reducing it to a predictable, rhythmic progression of states. This is particularly advantageous in systems where a large volume of data must be processed in parallel, as it provides a systematic method for accessing and utilizing the stored information.

[0204] It should be noted that while Fig. 10 illustrates only two microvaults, the described daisy-chaining mechanism can be extended to accommodate any arbitrary number of microvaults. Additional microvaults can be added to the chain, with each new microvault connected to the multiplexer via an additional input line. The multiplexer 1060 and counter 1062 would be scaled accordingly to manage the increased number of inputs, maintaining the same clock-driven, sequential data retrieval process across the expanded memory architecture.

[0205] This daisy-chaining of microvaults, in conjunction with the multiplexing and counter-driven control system, exemplifies a modular and scalable approach to memory design in semiconductor devices. It allows for the customization of memory arrays to match the capacity and performance needs of a wide range of applications,Attorney Docket No. P25-059-SEC-W001from embedded systems to large-scale data centers, providing a versatile solution for modem computing challenges.

[0206] Fig. 11 shows a semiconductor assembly 1100 incorporating a daisy-chained configuration of microvaults in multiple semiconductor devices 1106, 1116 that are operatively connected to multiplexers and managed by counters for coordinated data selection and retrieval, in accordance with an embodiment of the present disclosure.

[0207] The present detailed description relates to Fig. 1 1 of the accompanying drawings, which illustrates an embodiment of an assembly 1100 as part of an integrated circuit. The assembly 1100 may be seen as a hierarchical structure that includes a bottom semiconductor device 1122, a middle semiconductor device 1116, and a top semiconductor device 1106 (each of these may be chiplets). Each semiconductor device is configured to interface with the others through a series of bumpless bonds, such as bumpless bonds 1128, 1129, 1158, 1159, 1160, 1161, 1162, and 1163, which facilitate electrical connectivity without the added profile of traditional bonding methods, thus enabling a compact and dense stacking of semiconductor layers.

[0208] The bottom semiconductor device 1122 includes a read address register 1126, which may be configured to store and communicate read addresses to microvaults located across the assembly 1100. The read address register 1126 communicates via an interconnect 1174, which serves as a conduit for signals directed to bumpless bonds 1128. These bonds in turn engage with bumpless bonds 1129 of the middle semiconductor device 1116, thus transferring the read addresses into the middle semiconductor device. The bottom semiconductor device 1122 also comprises a read data register 1124, which may serve as a repository for read data received. Read data is received through an interconnect 1176 that connects to bumpless bonds 1158, which are in communication with bumpless bonds 1159 of the middle semiconductor device 1116.

[0209] The middle semiconductor device 1116 serves as an intermediary layer within the assembly 1100, housing microvaults such as microvault 1132 and micro vault 1118, each of which may be designed to store and rapidly provide access to data. These microvaults are linked to other components within the device via interconnects, such as interconnect 1130 and interconnect 1134, which guide the flow of read addresses and read data, respectively. The middle semiconductor device 1116 also features a readAttorney Docket No. P25-059-SEC-W001address register 1146, which receives read addresses from interconnect 1130, and a read data register 1154, which collects read data from TSV 1152.

[0210] The multiplexer 1114, within the middle semiconductor device 1116 selects between various data streams. This multiplexer is controlled by a phase counter 1110, which determines the sequence of data selection based on the input received from the top semiconductor device 1106 via the TSV 1152. The read data register 1112 serves as a holding area for the selected data stream from the multiplexer 1114.

[0211] The top semiconductor device 1106 features a phase counter 1102 and a read data register 1104, which are used for coordinating data flow within the assembly 1100. The phase counter 1102, in conjunction with multiplexer 1150, dictates the output of read data from microvaults 1138 and 1164 based on the selected phase. The read data register 1104 captures the output from the multiplexer 1150, which is then relayed through interconnect 1108 to bumpless bonds 1162, facilitating communication with the middle semiconductor device 1116.

[0212] The assembly 1100 illustrates the integration of multiple microvaults across different semiconductor devices. For example, the microvault 1132 on the middle semiconductor device 1116 may receive read addresses from read address register 1146 via an interconnect 1134, path "1". Similarly, microvault 1118 may receive read addresses from the same register via an interconnect 1156, path "2". Microvaults 1138 and 1164 on the top semiconductor device 1106 receive read addresses via interconnect 1140 and paths "3" and "4", respectively, from read address register 1142, which is in communication with the middle semiconductor device 1116 via TSV 1136 and bumpless bonds 1160 and 1161.

[0213] Each microvault, such as 1132, 1118, 1138, and 1164, can potentially output read data to the multiplexer 1114 or 1150, where the data is then selected based on the configuration of the respective phase counter, 1110 or 1102. The selected data is temporarily stored in a read data register, either 1112 or 1104, before being transmitted down the assembly 1100 through the respective bumpless bonds and interconnects, ultimately reaching the read data register 1124 of the bottom semiconductor device 1122. This arrangement allows for a synchronized read-out of data from all microvaults, which may be essential in applications requiring parallel processing and high-speed data access.

[0214] The TSVs, such as 1136 and 1152, provide vertical connectivity across the semiconductor devices, enabling the integration of additional layers orAttorney Docket No. P25-059-SEC-W001functionalities atop the existing assembly 1100. These TSVs are coupled to various interconnects and bumpless bonds that establish the pathways for signal transmission both within and between the semiconductor devices.

[0215] In some embodiments, the microvaults within the assembly 1100 may include various memory technologies, such as 3D-NAND or 3D-NOR structures, and are arranged to facilitate parallel processing and efficient data retrieval. Each micro vault may include additional features, such as thermal management layers for heat dissipation, hardware -based encryption modules for data security, or power management circuits to optimize energy consumption.

[0216] datapath 1 within the assembly 1100 exemplifies a route through which read addresses and corresponding read data are transmitted across the assembly, specifically directing operations from the read address register 1126 located on the bottom semiconductor device 1122 to the read data register 1124 within the same device.

[0217] The process begins with the read address register 1126 holding a specific read address. This address is sent through interconnect 1174, which acts as a channel for the signal. The read address is then transmitted to bumpless bonds 1128, which are meticulously designed to create a reliable electrical connection without the physical protrusion associated with traditional bonding methods. These bonds ensure a low-profile interface that preserves the compactness of the semiconductor stack.

[0218] The signal continues from bumpless bonds 1128 to engage with bumpless bonds 1129 of the middle semiconductor device 1116. The read address is earned forward by interconnect 1130, which delivers the address to the read address register 1146 of the middle semiconductor device. Read address register 1146, in turn, propagates the read address through interconnect 1134, designated as path "1" guiding the signal to the microvault 1132.

[0219] Upon receiving the read address, microvault 1132 accesses the requested data. This data is then outputted through interconnect 1170 and directed to the multiplexer 1114. In some embodiments, multiplexer 1114 functions as a selective switch that chooses between data streams based on the configuration determined by phase counter 1110. This phase counter may be designed to cycle through a sequence that dictates the timing and selection of data streams, ensuring that each microvault is read in a coordinated manner.Attorney Docket No. P25-059-SEC-W001

[0220] The selected data from multiplexer 1114 is then captured by read data register 1112, which holds the data momentarily. The data is subsequently sent via interconnect 1120, which carries the signal to bumpless bonds 1159. These bonds are part of a sophisticated electrical interconnection system that, along with bumpless bonds 1158 on the bottom semiconductor device 1122, enables vertical and horizontal integration within the semiconductor stack.

[0221] The signal, now in the form of read data, traverses from bumpless bonds 1159 to bumpless bonds 1158 and is finally introduced into interconnect 1176. This interconnect completes the connection to read data register 1124, which is configured to receive and hold the read data. The read data register 1124 may be equipped to retain the data for subsequent processing or external communication.

[0222] datapath 2 within the assembly 1100 delineates a route specifically designed for the transmission of read addresses from the read address register 1126 on the bottom semiconductor device 1122 to the microvault 1118 located on the middle semiconductor device 1116 and the subsequent transfer of read data back to the read data register 1124 on the bottom device.

[0223] The journey commences at the read address register 1126, where a read address is held in preparation for dispatch. This register is an part of the semiconductor device's control mechanism, orchestrating the retrieval of data by issuing specific addresses to the memory units. From the read address register 1126, the read address is sent through the interconnect 1174, which provides a secure and reliable pathway for electrical signals within the integrated circuit.

[0224] The read address continues from the interconnect 1174 to bumpless bonds 1128, which offer a seamless and low-profile connection to the middle semiconductor device 1116 via the corresponding bumpless bonds 1129. These bonds maintain the signal's integrity during inter-layer communication and are designed to accommodate the requirements of modern semiconductor architectures.

[0225] The signal is then channeled via interconnect 1130 to the read address register 1146 within the middle semiconductor device 1116. The read address register 1146 acts as a secondary store and hold register From there, the read address is directed down interconnect 1156, labeled as path "2," which terminates at the microvault 1118.

[0226] Upon receipt of the read address, microvault 1118 accesses the corresponding data. This data retrieval process is facilitated by the microvault's internal architecture, which may comprise an array of memory cells optimized for rapid accessAttorney Docket No. P25-059-SEC-W001and data stability. The read data is outputted from microvault 1118 and travels through interconnect 1172, which leads to the multiplexer 1114.

[0227] Multiplexer 1114 determines which data stream to forward based on input from the phase counter 1110. The phase counter 1110 operates in synchronization with the system clock or an external control signal, cycling through various states to control the selection process of the multiplexer 1114 in a precise and predictable manner.

[0228] The output of the multiplexer 1114, now carrying the selected read data, is conveyed to read data register 1112. This register temporarily stores the read data, acting as a buffer. The read data is then dispatched via interconnect 1120 towards bumpless bonds 1159.

[0229] Bumpless bonds 1159 form the interface with bumpless bonds 1158 on the bottom semiconductor device 1122, where the signal is transmitted downward through the assembly. The read data then traverses the interconnect 1176 to reach its final destination, the read data register 1124. The read data register 1124 captures the read data, holding it in readiness for further processing or transmission to external circuits.

[0230] datapath 3 within the assembly 1100 is another communication route that illustrates the data transfer sequence from the read address register 1126 on the bottom semiconductor device 1122, through various components, ultimately to the microvault 1138 on the top semiconductor device 1106, and then back to the read data register 1124 on the bottom device.

[0231] The sequence initiates at the read address register 1126, which serves as the origin point for read addresses. The register 1126 securely holds the address before it is dispatched through the interconnect 1174. Interconnect 1174 acts as a dedicated channel, ensuring that the read address is conveyed with precision to the bumpless bonds 1128. These bumpless bonds 1128 facilitate a streamlined connection to the bumpless bonds 1129 of the middle semiconductor device 1116, preserving the integrity and compactness of the signal pathway.

[0232] Upon reaching the middle semiconductor device 1116, the read address is relayed through interconnect 1130 to the read address register 1146. The read address register 1146 acts as a juncture that further propagates the address signal through the Through-Silicon Via (TSV) 1136. The TSV 1136 is a vertical interconnect that pierces through the semiconductor substrate, providing a direct link from the middleAttorney Docket No. P25-059-SEC-W001semiconductor device 1116 to the top semiconductor device 1106, thus exemplifying the 3D integration capabilities of semiconductor design.

[0233] The read address ascends from TSV 1136 and emerges onto bumpless bonds 1160 on the middle semiconductor device 1116. The bumpless bonds 1160 are connected to the bumpless bonds 1161 of the top semiconductor device 1106. The address signal is conducted through interconnect 1140 to the read address register 1142 on the top device 1106.

[0234] The read address register 1142, upon receiving the read address, directs the signal along interconnect 1144. This path, denoted as "3," leads the address to the microvault 1138. Microvault 1138, designed for data storage, retrieves the requested information in response to the read address. The read data is outputted through interconnect 1166, which feeds the data into the multiplexer 1150.

[0235] The multiplexer 1150 in the top semiconductor device 1106 is governed by the phase counter 1102, which dictates the selection of the data stream to be channeled to the read data register 1104. The chosen data stream is temporarily housed in the read data register 1104, where it awaits downstream transmission.

[0236] The read data departs from the read data register 1104 via interconnect 1108, which connects to bumpless bonds 1162. These bonds 1162 engage with the corresponding bumpless bonds 1163 on the middle semiconductor device 1116, transferring the read data to TSV 1152.

[0237] TSV 1152 operates as a vertical conduit, allowing the read data to traverse down into the middle semiconductor device 1116, where it is received by the read data register 1154. The read data register 1154 holds the read data momentarily before it is directed to the multiplexer 1114 through interconnect 1156.

[0238] The multiplexer 1114 in the middle semiconductor device 1116, coordinated by the phase counter 1110, selects the appropriate data for output. The read data is then channeled to the read data register 1112, where it is briefly stored. Following this, the read data travels via interconnect 1120 to bumpless bonds 1159.

[0239] The bumpless bonds 1159 form an interface with bumpless bonds 1158 on the bottom semiconductor device 1122. The read data signal is then carried through interconnect 1176, culminating its journey at the read data register 1124 on the bottom device.

[0240] datapath 4 within the assembly 1100 is a path that establishes the flow of read addresses from the read address register 1126 on the bottom semiconductorAttorney Docket No. P25-059-SEC-W001device 1122 to the microvault 1164 located on the top semiconductor device 1106, and subsequently facilitates the movement of read data back down to the read data register 1124 on the bottom device.

[0241] This datapath begins at the read address register 1126, which is responsible for holding and issuing the read addresses needed for data retrieval from the microvaults. The read address is sent from the register 1126 through interconnect 1174, a pathway that maintains signal integrity and facilitates electrical communication.

[0242] From interconnect 1174, the read address is directed to bumpless bonds 1128. These bonding areas create an interconnection between the bottom semiconductor device 1122 and the middle semiconductor device 1116 through bumpless bonds 1129. The design of these bumpless bonds facilitates data transmission.

[0243] Once the read address reaches the middle semiconductor device 1116, it is carried forward by interconnect 1130 to the read address register 1146. This register acts as an intermediary, preparing the address for its vertical ascent through the device stack. The address is then transmitted via the Through-Silicon Via (TSV) 1136 that facilitates vertical integration by providing a direct electrical link through the semiconductor substrate.

[0244] After ascending through TSV 1136, the read address emerges onto bumpless bonds 1160, which are aligned to connect with bumpless bonds 1161 on the top semiconductor device 1106. The read address then proceeds along interconnect 1140 to the read address register 1142 located on the top device.

[0245] The read address register 1142 serves to forward the read address to its final destination, the microvault 1164, through interconnect 1148, labeled as path "4". Microvault 1164, upon receiving the read address, retrieves the requested data, which is then outputted through interconnect 1168. This data is directed to the multiplexer 1150, which is under the control of phase counter 1102.

[0246] The phase counter 1102 determines which data stream is selected by the multiplexer 1150, which then sends the read data to the read data register 1104. The read data register 1104 acts as a temporary repository, holding the data until it can be sent downwards through the device stack.

[0247] The data leaves the read data register 1104 and travels via interconnect 1108 to bumpless bonds 1162. These bonds maintain a connection to bumpless bonds 1163 on the middle semiconductor device 1116. The read data is then transferred toAttorney Docket No. P25-059-SEC-W001TSV 1152, which carries the data vertically down to the read data register 1154 on the middle semiconductor device 1116.

[0248] The read data register 1154 temporarily holds the read data before it is fed into the multiplexer 1114 via interconnect 1156. The multiplexer 1114, coordinated by the phase counter 1110, channels the appropriate data stream to the read data register 1112. This register serves as a staging area for the read data, which is then sent through interconnect 1120 to bumpless bonds 1159.

[0249] Bumpless bonds 1 159 interface with bumpless bonds 1158 on the bottom semiconductor device 1122 to provide a downward transmission of the read data. Finally, the signal is routed through interconnect 1176 and arrives at read data register 1124, where the data is made available for subsequent processing or external communication.

[0250] Fig. 12 illustrates a three-dimensional (3D) memory column 1200 configured as a 3D-NOR or 3D-AND structure, featuring a series of ferroelectric fieldeffect transistors (FeFETs) 1202 with interconnected drain terminals 1204 linked to a common select line 1212 and individual gate terminals 1206 connected to respective read / write enable lines 1214 (e.g., 1214a for FeFET 1202a, all coupled to a common bit line, in accordance with an embodiment of the present disclosure.

[0251] Thus, Fig. 12 depicts a three-dimensional (3D) memory column, designated as element 1200, which can be configured in various embodiments as either a 3D-NOR or 3D-AND structure, providing flexibility in the application and use of the integrated circuit. This memory column is an assembly of multiple ferroelectric fieldeffect transistors (FeFETs), collectively referred to as fefets 1202, where each FeFET is indicated by elements such as 1202a, 1202b, 1202c, and 1202d, among others potentially present in the array.

[0252] Within each FeFET, such as 1202a, there is a drain terminal 1204a. This drain terminal is part of the memory cell's output path and is connected to a common select line 1212. In some embodiments, the common select line 1212 serves as acontrol mechanism that enables the selection of a particular FeFET for data read or write operations.

[0253] The gate terminal of each FeFET, exemplified by 1206a for FeFET 1202a, is individually connected to a respective read / write enable line, such as 1214a. This enables control of the FeFET's state, allowing it to be in a conductive (on) state for reading or writing data or in a non-conductive (off) state to prevent data flow. TheAttorney Docket No. P25-059-SEC-W001presence of individual read / write lines for each FeFET may allow for precise control and operation of each memory cell.

[0254] Moreover, each FeFET, such as 1202a, comprises a source terminal, such as 1208a, which is coupled to a common bit line 1210. The bit line 1210 provides a conduit for data being written to or read from the FeFETs. In some embodiments, this bit line can be shared across multiple memory columns, which can facilitate parallel processing and increased data throughput.

[0255] In accordance with various embodiments of the present disclosure, the 3D memory column 1200 may incorporate additional elements and configurations to enhance performance and functionality. For instance, the 3D memory column 1200 may include insulating materials, conductive pathways, and other structural components not explicitly shown in Fig. 12 but which are inherent to the implementation of such 3D memory structures. The FeFETs 1202 may also exhibit variations in terms of material composition, structural dimensions, and electrical properties, contributing to a range of performance characteristics suitable for different applications.

[0256] Furthermore, the memory column 1200 may be incorporated into larger memory arrays, forming part of a memory module or system. These arrays can be arranged in various configurations, such as rows and columns, to create a matrix that efficiently addresses the demands of high-density data storage. The memory column 1200 can also be interfaced with other circuit elements and control logic, which may govern the operation of the memory array, including data management protocols, error correction algorithms, and power optimization strategies.

[0257] In some embodiments, the memory column 1200 may be fabricated using advanced semiconductor manufacturing techniques, such as photolithography, etching, deposition, and planarization processes. The choice of materials for the FeFETs, including the ferroelectric material, the semiconductor channel, and the conductive elements, can be selected based on desired electrical characteristics, such as charge retention, switching speed, and energy efficiency.

[0258] The 3D memory column 1200, as illustrated in Fig. 12, may include FeFETs, such as 1202, fabricated from a variety of materials that provide the electrical and physical properties to achieve the desired functionality. For instance, in some embodiments, the channel layer of each FeFET in the FeFETs 1202 could be constructed from materials such as Indium Gallium Zinc Oxide (IGZO) or other Amorphous Oxide Semiconductors (AOS) like Zinc Tin Oxide or Indium TungstenAttorney Docket No. P25-059-SEC-W001Oxide (IWO). These materials are selected fortheir electronic properties, such as carrier mobility and stability.

[0259] The ferroelectric material rigidly coupled to the channel layer in each FeFET may comprise hafnium zirconium oxide (HfZrO2) or other transition metal oxides, perovskites, etc. These ferroelectric materials are chosen for their ability to maintain a polarization state when an electric field is applied, which is used for the nonvolatile memory characteristics of the FeFETs. The thickness, crystalline structure, and stoichiometry of the ferroelectric layer can be controlled to achieve the desired coercive voltage, remanent polarization, and other electrical parameters for reliable data storage and retrieval.

[0260] The drain 1204 and source 1208 terminals of the FeFETs 1202 are connected to the common select line 1212 and common bit line 1210, respectively. These common lines may be formed from conductive materials such as tungsten, titanium nitride, or other metals and metal alloys that provide low-resistance pathways for electrical signals. The configuration of these terminals and their respective common lines ensures that the FeFETs can be accessed and controlled effectively during operation.

[0261] Each gate terminal, such as the gate 1206 of the FeFET 1202a, is connected to its respective read / write enable line, such as 1214a. The gate terminals are help control the state of the FeFET, and the materials chosen for these terminals may include various conductive materials that can provide a reliable electrical interface with the ferroelectric material. The read / write enable 1214 lines are designed to deliver suitable voltage levels to the gates 1206 of the FeFETs 1202 for switching between states.

[0262] The memory column 1200 as a whole is designed to support a range of operating parameters. In some embodiments, these parameters may include, but are not limited to, an off-state current of less than 10A-8 amps per centimeter cubed, an on-state current greater than 10 -7 amps per centimeter cubed, and a channel mobility that is maintained despite the presence of the ferroelectric layer. The channel layer's thickness can be less than 30 nm to ensure high device density, while the ferroelectric layer's characteristics, such as coercive voltage and remanent polarization, are optimized to provide the memory functionality.

[0263] In some embodiments, the FeFETs 1202 may include additional materials or dopants to enhance their electrical properties. For instance, dopants suchAttorney Docket No. P25-059-SEC-W001as gallium (Ga), indium (In), or zinc (Zn) may be introduced into the channel layer to modulate the carrier concentration or to adjust the threshold voltage of the FeFETs. Similarly, the ferroelectric layer may include dopants like lanthanum (La) or niobium (Nb) to adjust its ferroelectric properties.

[0264] In other embodiments, the 3D memory column 1200 may be integrated with additional semiconductor devices and structures to form complex memory systems. These systems can provide storage capabilities and support various memory architectures.

[0265] Fig. 13 depicts a three-dimensional (3D) memory column 1300 configured as a 3D-NAND structure, consisting of a vertical stack of ferroelectric fieldeffect transistors (FeFETs) 1302, each with source 1304 and drain 1308 terminals. The source 1304 of each FeFET, such as 1304a for FeFET 1302a, is coupled to the start of a bit line 1310 or connected to the drain of the preceding FeFET, exemplified by source 1304b of FeFET 1302b coupled to drain 1308a of FeFET 1302a. Each FeFET includes a gate 1306, such as 1306a for FeFET 1302a, connected to a respective read / write enable line, illustrated by 1314a for FeFET 1302a, in accordance with an embodiment of the present disclosure.

[0266] The 3D memory column 1300 is composed of a series of vertically stacked Field-Effect Transistors (FeFETs), identified collectively as FeFETs 1302. These transistors, which include FeFETs 1302a, 1302b, 1302c, 1302d, and so on, are characterized by their incorporation of ferroelectric materials within their gate structure. Each FeFET in the series is of the memory column contributes to the memory storage capabilities of the device.

[0267] In the depicted embodiment, each FeFET, such as FeFET 1302a, includes a source terminal 1304, for instance, source 1304a, which is coupled to a bit line 1310. The bit line 1310 serves as a conduit for electrical signals that are used to read from and write to the memory cell associated with FeFET 1302a. In scenarios where FeFET 1302a is not the bottom-most transistor in the column, its source 1304b may be connected to the drain 1308a of the immediately preceding FeFET, such as FeFET 1302a, facilitating a serial connection that defines the vertical NAND architecture.

[0268] Each FeFET within the FeFETs 1302 is further equipped with a gate terminal 1306, exemplified by gate 1306a for FeFET 1302a. This gate terminal 1306 is coupled to a respective read / write enable line, exemplified by 1314a for FeFET 1302a.Attorney Docket No. P25-059-SEC-W001The read / write enable line 1314a is responsible for controlling the state of the FeFET, allowing it to either conduct or prevent the flow of current through the device, thereby enabling the writing or reading of data.

[0269] Moreover, each FeFET of the FeFETs 1302 also includes a drain terminal 1308, such as drain 1308a for FeFET 1302a, which is typically connected to the source of the subsequent FeFET in the vertical stack. This arrangement ensures that the charge stored in the ferroelectric material of the gate can modulate the cunent flowing from the source to the drain, allowing for the storage and retrieval of data.

[0270] The memory column 1300 within Fig. 13 is indicative of a memory architecture that can be utilized in various applications, from portable electronics to enterprise-level data storage systems. In some embodiments, the ferroelectric material used in the FeFETs may include various compositions, such as hafnium oxide, zirconium oxide, or any combination thereof, which can be doped with elements such as lanthanum or yttrium to adjust the ferroelectric properties as required.

[0271] In some variations, the 3D memory column 1300 can incorporate additional features that enhance performance, reliability, or manufacturability. For instance, the FeFETs 1302 may include protective layers to shield the ferroelectric material from environmental factors or process-induced damage. The column 1300 may also be integrated with other circuit elements, such as capacitors or diodes, to facilitate operations like charge pumping or to provide additional functionality within the memory array.

[0272] The 3D memory column 1300 may be fabricated from a variety of materials that confer specific electrical properties to enhance device performance. In some embodiments, the channel layer of each FeFET may be formed from materials such as Indium Gallium Zinc Oxide (IGZO). Other materials for the channel layer could include Amorphous Oxide Semiconductors (AOS) like Zinc Tin Oxide or Aluminum Zinc Oxide.

[0273] The ferroelectric layer within the FeFETs 1302 may comprise materials such as Hafnium Zirconium Oxide (HfZrO2). The ferroelectric layer's thickness and material composition can be controlled through methods like Atomic Layer Deposition (ALD) to achieve the desired coercive voltages, remanent polarizations, and endurance characteristics. In some implementations, the coercive voltage of the ferroelectric layer may be tuned to be between -3 Volts to +3 Volts, facilitating low-voltage operation of the memory devices.Attorney Docket No. P25-059-SEC-W001

[0274] The source and drain terminals of the FeFETs 1302 may be composed of conductive materials such as Tungsten or Titanium Nitride. These materials may also be selected to optimize the contact resistance with the channel layer, reducing overall power consumption and improving the lon / Ioff ratio of the device.

[0275] Additionally, the FeFETs 1302 may be engineered to exhibit specific electrical parameters. For instance, the channel layer's thickness may be less than 30 nm in some embodiments. In some embodiments, the channel layer may demonstrate a carrier concentration of 10A17 to 10A20 per centimeter-cubed, which can be adjusted through doping with elements such as Gallium, Indium, or Zinc to modulate the electrical properties.

[0276] The memory cells formed by the FeFETs 1302 within the 3D memory column 1300 may also target operational parameters such as read and write latencies, endurance, and energy consumption. For example, read and write operations may be executed with energies less than 10 picojoules and within timeframes less than 20 nanoseconds, contributing to the low power and high-speed attributes of the memory column.

[0277] Furthermore, the 3D-NAND configuration of the memory column 1300 may be designed to achieve a high off-state resistance to on-state resistance ratio (Roff / Ron), which may be used for distinguishing between different data states and ensuring reliable data retention. This ratio may be about 10A3 or greater, which helps to maintain a high signal-to-noise ratio during memory operations.

[0278] The FeFETs 1302 in the memory column 1300 may also be designed to sustain a high degree of reliability, with endurance ratings greater than or equal to 10 ll cycles, ensuring the longevity and durability of the memory device. This endurance is complemented by the ferroelectric layer's ability to maintain data retention for at least 1 minute at room temperature, which is 25°C.

[0279] Fig. 14 depicts a three-dimensional (3D) memory column configured as a 3D-NAND with an integrated pass gate, in accordance with an embodiment of the present disclosure. This figure illustrates a series of ferroelectric field-effect transistors (FeFETs) 1402, each including source 1404 and drain 1408 terminals, gated by respective gate terminals 1406 and coupled to read / write enable lines 1414. The FeFETs are interconnected, forming a vertical memory structure with pass gates 1418 linked to a pass gate line 1416.Attorney Docket No. P25-059-SEC-W001

[0280] Thus, Fig. 14 illustrates a three-dimensional (3D) memory column, designated as element 1400, which can be configured as a 3D-NAND structure with an integrated pass gate. This configuration enables enhanced control over individual memory cells within the 3D structure, potentially improving read / write operations and facilitating efficient memory management.

[0281] In detail, the 3D memory column 1400 comprises multiple ferroelectric field-effect transistors (FeFETs), collectively referred to as FeFETs 1402. Each FeFET within the 1402 series, such as 1402a, 1402b, 1402c, 1402d, etc., is a constituent memory cell of the 3D memory column 1400. These FeFETs are utilized for their ability to retain data in a non-volatile manner due to the ferroelectric properties of their gate material, which allows for data retention without continuous power supply for a time.

[0282] Each FeFET in the series 1402 includes a source, exemplified by source 1404a for FeFET 1402a. The source 1404 for each FeFET is either coupled to a bit line, illustrated as bit line 1410 for FeFET 1402a, or is connected to the drain of the preceding FeFET in the series. For instance, source 1404b of FeFET 1402b is electrically coupled to drain 1408a of FeFET 1402a. This serial connection forms the basis for the daisy-chain configuration, which is used in NAND architectures, allowing for sequential access to the array of FeFETs.

[0283] Furthermore, each FeFET within the FeFETs 1402 is equipped with a gate terminal, such as gate 1406a for FeFET 1402a. The gates of the FeFETs are connected to their respective read / write enable lines, which are depicted as element 1414 in the figure. For example, gate 1406a of FeFET 1402a is influenced by read / write enable line 1414a. These enable lines control the application of appropriate voltages for the reading and writing of data.

[0284] Additionally, each FeFET in the FeFETs 1402 series includes a drain, such as drain 1408a for FeFET 1402a. This drain is connected to the source of the subsequent FeFET in the series, thus establishing the continuity for the columnar structure of the 3D memory stack.

[0285] In some embodiments, each FeFET of the FeFETs 1402 incorporates a pass gate, for example, pass gate 1418a, which is connected to a pass gate line, represented by 1416 in the figure. The pass gate line 1416 is a conductive pathway that provides electrical signals to control the pass gates 1418 of the FeFETs. The inclusion of pass gates in the FeFETs may allow for improved isolation between memory cellsAttorney Docket No. P25-059-SEC-W001during operation, thereby reducing interference and potentially enhancing the reliability of data storage and retrieval.

[0286] The 3D memory column 1400, as depicted in Fig. 14 also encompasses a diverse range of materials and parameters that could be utilized to optimize its performance in various embodiments. Each FeFET 1402 within the column could be fabricated using a variety of semiconductor materials. For instance, the channel layer of the FeFETs could be formed from materials such as Indium Gallium Zinc Oxide (IGZO).

[0287] The ferroelectric layer, which is a defining characteristic of the FeFETs, may be composed of materials like Hafnium Zirconium Oxide (HfZrO2) or other perovskite materials, which are known for their remanent polarization. This property determines the data retention capabilities of the FeFETs. The coercive voltage of this layer, which affects the energy required to switch the polarization state, is another parameter that can be adjusted according to the requirements of the specific application, with a range in one embodiment being between -3 Volts to 3 Volts.

[0288] The source and drain terminals of the FeFETs, which include elements 1404 and 1408, respectively, could be composed of conductive materials such as Tungsten or Titanium Nitride. These materials provide pathways for current flow, which is for switching. The read / write enable lines 1414, which control the gates 1406 of the FeFETs, could also be fabricated from similar materials, ensuring consistent electrical characteristics throughout the device.

[0289] In terms of the physical parameters, the channel layer thickness may be less than 30 nm in thickness. The electron mobility within these channel layers may be maintained at a predetermined level even when the layer is less than 30nm in thickness

[0290] The pass gates 1418 may be manufactured using low-resistance materials to enable quick switching times, which is beneficial when the memory column is accessed frequently during operation.

[0291] Fig. 15 illustrates a three-dimensional (3D) memory column 1500, which may be configured as either a 3D-NOR or a 3D-AND structure with independent Read / Write enable capabilities, in accordance with an embodiment of the present disclosure. This memory column encompasses a series of vertically aligned FeFETs 1502, such as FeFETs 1502a, 1502b, 1502c, 1502d, and so forth, each integrated with a source 1504 (e.g., source 1504a for FeFET 1502a) linked to a respective read enable line 1520 (e.g., read enable line 1520a for FeFET 1502a), and a gate 1506 (e.g., gateAttorney Docket No. P25-059-SEC-W0011506a for FeFET 1502a) connected to a corresponding write enable line 1522 (e.g., write enable line 1522a for FeFET 1502a). All FeFETs within the column share a common bit line 1510 connected to their drains 1508, enabling the column to perform coordinated memory operations.

[0292] Fig. 15 presents a detailed depiction of a three-dimensional (3D) memory column 1500, which can be configured as a 3D-NOR or 3D- AND structure with independent Read / Write enable functionalities. This memory column is an assembly of Field-Effect Transistors with ferroelectric gate layers, commonly referred to as FeFETs 1502, which are individually identified, for example, as 1502a, 1502b, 1502c, 1502d, etc., each representing a memory cell within the column.

[0293] In the illustrated embodiment, each FeFET 1502 includes a source 1504, such as source 1504a corresponding to FeFET 1502a. The source 1504 is designed to be electrically coupled to a respective read enable line 1520, such as read enable line 1520a which is dedicated to FeFET 1502a. The read enable line 1520 functions to selectively activate the FeFET 1502 for reading operations, allowing the readout of stored data from the memory cell.

[0294] Additionally, each FeFET 1502 is equipped with a gate 1506, exemplified by gate 1506a for FeFET 1502a. The gate 1506 is connected to a respective write enable line 1522, such as write enable line 1522a, which is specific to FeFET 1502a. The write enable line 1522 serves to selectively activate the FeFET 1502 for writing operations, enabling the storage of data within the memory cell.

[0295] Furthermore, each FeFET 1502 includes a drain 1508, for instance, drain 1508a affiliated with FeFET 1502a. The drain 1508 is connected to a common bit line 1510. The bit line 1510 acts as a conduit for transferring data to and from the memory cells during read and write operations. The commonality of the bit line 1510 across multiple FeFETs 1502 signifies that data from any activated memory cell can be routed through this shared path.

[0296] In some embodiments of the memory column 1500, the configuration of the FeFETs 1502 allows for a high density of memory cells vertically stacked within a compact footprint.

[0297] The ferroelectric material utilized in the gate 1506 of the FeFETs 1502 may comprise various compositions, such as hafnium oxide-based materials, which can be deposited using atomic layer deposition techniques. The ferroelectric property of theAttorney Docket No. P25-059-SEC-W001material allows for data retention, enabling the memory cells to maintain stored information even when power is not supplied.

[0298] The source 1504, gate 1506, and drain 1508 of each FeFET 1502 can be fabricated from materials that provide predetermined electrical performance. These materials may include metals such as tungsten or copper, or metal nitrides such as titanium nitride.

[0299] The read enable lines 1520 and write enable lines 1522 can be designed to minimize crosstalk and interference between adjacent lines, in some embodiments. In some specific embodiments, shielding layers or insulating materials may be included to further isolate the signal paths.

[0300] Furthermore, the described memory column 1500 may be integrated within a larger semiconductor device, such as a processor or a storage module. It may form part of a system-on-chip (SoC) or be included in a multi-chip module (MCM), contributing to a data storage and retrieval system.

[0301] The materials that constitute the FeFETs 1502 within the memory column 1500 are selected to provide specific electrical and physical properties to optimize the performance of the integrated circuit. For instance, the channel layer in each FeFET may be formed from advanced semiconductor materials, such as Indium Gallium Zinc Oxide (IGZO) or other Amorphous Oxide Semiconductors (AOS) like Zinc Tin Oxide or Cadmium Oxide. These materials are chosen for their excellent electron mobility characteristics and stability.

[0302] The ferroelectric layer, integral to the FeFETs 1502, may be fabricated from various ferroelectric materials that exhibit suitable polarization properties. Materials such as Hafnium Zirconium Oxide (HfZrO2) or Lead Zirconate Titanate (PZT) could be utilized. These materials can be doped with elements such as Lanthanum, Yttrium, or other suitable dopants to modify their ferroelectric properties, such as coercive voltage, remanent polarization, and crystallization temperature. The ferroelectric layer's thickness and material composition may be adjusted to achieve desired memory characteristics, such as write endurance and retention time, while ensuring the layer remains compatible with the overall semiconductor manufacturing process, or other considerations, etc.

[0303] The source 1504, gate 1506, and drain 1508 terminals of the FeFETs 1502 may be composed of conductive materials like Tungsten, Titanium Nitride,Attorney Docket No. P25-059-SEC-W001Nickel, or Molybdenum. Connections to the read enable lines 1520 and the write enable lines 1522 may be facilitated through conductive vias or contacts.

[0304] The read enable lines 1520 and write enable lines 1522, along with the common bit line 1510, may be patterned using lithographic techniques to achieve the predetermined precision and alignment for proper functionality. These lines may be insulated from one another using dielectric materials like Silicon Dioxide (SiO2), Silicon Nitride (Si3N4), or low-k dielectrics to reduce parasitic capacitance and crosstalk.

[0305] Each element within the memory column 1500 may consider factors such as line width, spacing, and aspect ratio to ensure manufacturability, functionality, and / or other goals or characteristics. The materials and processes used in the construction of the memory column 1500 are chosen to ensure compatibility with standard semiconductor fabrication techniques, such as photolithography, etching, deposition, and annealing, while also enabling the integration of materials and structures.

[0306] The fabrication of the FeFETs 1502 within the memory column 1500 may involve deposition techniques such as atomic layer deposition (ALD), chemical vapor deposition (CVD), or physical vapor deposition (PVD) to create uniform and / or non-uniform layers.

[0307] Fig. 16 presents a cross-sectional view of a 3D memory structure, designated as 1600, configured as a single-port 3D NAND, in accordance with an embodiment of the present disclosure. The structure includes a first vertical structure 1608a and a second identical vertical structure 1608b, each comprising a dielectric column 1610a, 1610b, a channel column 1612a, 1612b disposed around the dielectric column, and a ferroelectric column 1614a, 1614b disposed around the channel column. A series of horizontal gate-electrode layers 1606a-c are disposed at predetermined distances from each other, adjacent to the ferroelectric column along the length of the vertical structures. The assembly further includes a drain select layer 1602 and a source select layer 1604, with respective end dielectric columns 1618a, 1618b, and 1616a, 1616b positioned at the interfaces with the vertical structures, illustrating a detailed and intricate design for high-density data storage.

[0308] Thus, fig. 16 provides a cross-sectional view of a three-dimensional (3D) memory structure, designated as 1600, which is configured as a single-port 3D NAND architecture. This structure incorporates a pair of vertical structures, 1608a andAttorney Docket No. P25-059-SEC-W0011608b, which may be fabricated to be substantially identical, as indicated by their respective subscripts a and b, suggesting the potential for a modular and scalable memory array design.

[0309] Each vertical structure, exemplified by the first vertical structure 1608a, includes a dielectric column 1610a. The dielectric column may adopt various geometric forms — it can be cylindrical, substantially cylindrical, or feature curves. Additionally, it may present a tapered form, having different diameters at each end, implying a design that narrows towards the top. Both solid and hollow configurations of the dielectric column are contemplated within the scope of the disclosure, offering design flexibility for different electrical and structural requirements.

[0310] Surrounding the dielectric column 1610a is a channel column 1612a, which is the locus for charge carriers during device operation. The channel column is also described as potentially cylindrical, substantially cylindrical, or feature curves and / or and may exhibit similar variations in diameter along its length as the dielectric column.

[0311] Enveloping the channel column 1612a is a ferroelectric column 1614a, which extends along the length of the channel column but may recede at the ends, thereby meaning the ferroelectric column 1614 does not extend the entire length of the channel column 1612a.

[0312] Intersecting with the vertical structures is a series of horizontal gateelectrode layers, 1606a-c, which are positioned at predetermined distances from one another. These layers play a role in controlling the operational states of the device by influencing the electric field within the ferroelectric column.

[0313] Atop the 3D memory structure 1600 sits a drain select layer 1602, parallel to the horizontal gate-electrode layers 1606. Where the drain select layer 1602 meets the vertical structures 1608a, 1608b, end dielectric columns, 1618a and 1618b, are discernible. These end dielectric columns 1618 interface with the channel column 1612 and the drain select layer 1602, contributing to the isolation and control of the charge carriers within the channel column. They may contact the ferroelectric layer 1614, as they envelop the channel column 1612 at different positions along its length.

[0314] Similarly, a source select layer 1604 is situated at the bottom of the structure 1600, again parallel to the horizontal gate-electrode layers 1606. Corresponding end dielectric columns, 1616a and 1616b, are present where the sourceAttorney Docket No. P25-059-SEC-W001select layer 1604 interfaces with the vertical structures, serving analogous functions to the end dielectric columns 1618 near the drain select layer 1602.

[0315] The horizontal gate -electrode 1606 layers could be constructed from a range of conductive materials, including metals and metal compounds, which may offer different work functions, conductivity, and compatibility with other materials in the structure. Similarly, the ferroelectric column 1614 might incorporate a variety of ferroelectric materials each with its unique polarization characteristics, coercive fields, and dielectric constants, affecting the device's memory retention and switching behaviors.

[0316] The channel column 1612 materials can be chosen based on their electronic properties, such as carrier mobility and bandgap, to achieve the desired levels of on-state and off-state current. The dielectric column 1610 provides the electrical insulation to prevent leakage currents and ensure the proper functioning of the device.

[0317] The dielectric column, such as 1610a for the first vertical structure, may be constructed from materials that offer insulating properties to mitigate any potential leakage currents. Choices for the dielectric material may be Hafnium Oxide (HfO2) or Silicon Dioxide (SiO2).

[0318] Surrounding the dielectric column, the channel column (1612a) has channel material can be selected from a wide range of semiconducting materials that offer predetermined carrier mobility. For example, Indium Gallium Zinc Oxide (IGZO) can be used for its electrical properties. The channel layer’s thickness may vary, with some embodiments considering a thickness less than 30 nm. This thickness is chosen to achieve a predetermined electrical performance. The ferroelectric column, like 1614a, may include perovskite structures, such as Lead Zirconate Titanate (PZT).

[0319] The horizontal gate-electrode layers, represented by 1606a-c, are composed of conductive materials that facilitate the application of an electric field to the ferroelectric column, such as Tungsten or Titanium Nitride, which may be chosen for their electrical behavior. The selection of gate-electrode materials also takes into consideration factors such as work function, thermal stability, and ease of integration with the existing semiconductor manufacturing processes.

[0320] The drain and source select layers, 1602 and 1604 respectively, are incorporated to enable the addressing of individual memory cells within the array. The materials used for these layers are chosen for their conductive properties and compatibility with the channel and ferroelectric materials. The design of these layersAttorney Docket No. P25-059-SEC-W001may also incorporate considerations for reducing parasitic capacitance and ensuring swift data access.

[0321] The end dielectric columns, like 1618a and 1616a, provide electrical insulation at the ends of the channel column, where the ferroelectric material does not extend.

[0322] The disclosed embodiments within the 3D memory structure 1600 outline an assembly capable of providing data storage. The design allows for variations in structural dimensions, such as the diameter of the cylindrical columns, which can be uniform or tapered. Additionally, the option for solid or hollow configurations may be used.

[0323] Fig. 17 shows a 3D memory structure that is a dual-port 3D NAND arrangement in accordance with an embodiment of the present disclosure. The three-dimensional (3D) memory structure illustrated in Fig. 17, referred to as 3D memory structure 1700, exemplifies a dual-port 3D NAND arrangement to provide a memory functionality. This structure is characterized by two primary vertical formations, designated as the first vertical structure 1708a and the second vertical structure 1708b, which may be identical or near identical, as evidenced by the designating subscripts ‘a’ and ‘b’.

[0324] The first vertical structure 1708a includes a hollow or solid, tapered pass-gate electrode column 1718a that is substantially cylindrical in shape. The passgate electrode column 1718a may be made of titanium nitride and may have a larger diameter at the bottom end compared to the top end.

[0325] Surrounding the pass-gate electrode column 1718a is a dielectric column 1710a that may be made of hafnium oxide. The dielectric column 1710a is also substantially cylindrical with a slightly tapered shape, having a marginally larger diameter at the top. The dielectric column 1710a provides electrical isolation between the pass-gate electrode and subsequent layers.

[0326] Disposed around the dielectric column 1710a is a cylindrical channel column 1712a that may be made of IGZO semiconductor material. The channel column 1712a features curves along its length and has a uniform diameter throughout. The thickness of the channel column may be less than 30 nm.

[0327] Enclosing the channel column 1712a is a PZT ferroelectric column 1714a that covers most of the length of the channel column 1712a but recedes at the ends, leaving a portion of the channel column 1712a uncovered. The ferroelectricAttorney Docket No. P25-059-SEC-W001column 1714a is substantially cylindrical and contains lead, zirconium and titanium as key elemental constituents.

[0328] The vertical structures 1708a and 1708b traverse through several horizontal gate-electrode layers 1706a, 1706b and 1706c that may be made of tungsten, which are positioned at fixed intervals to form an interconnected grid layout. These layers influence the electric field within the ferroelectric column 1714a during memory operations.

[0329] At the top of the memory structure 1700, a drain select layer 1702 (e.g., titanium nitride ) runs parallel to the horizontal gate-electrode layers 1706. Where the drain select layer 1702 intersects the vertical structures 1708a and 1708b, end dielectric columns 1718a and 1718b are visible. These end columns (e.g., made of HfO2 ) touch the ferroelectric column 1714a on one end and surround the uncovered portion of channel column 1712a, providing insulation.

[0330] Similarly at the bottom of the structure 1700, a source select layer 1704 (e.g., made of tungsten), also parallel to the electrode layers 1706, interfaces with the vertical structures. End dielectric columns 1716a and 1716b can be observed at these intersection points, enclosing the open ends of the channel columns 1712a and 1712b.

[0331] Within the hollow region of the pass-gate electrode columns 1718a and 1718b at the ends, a thin dielectric horizontal layers 1720a and 1720b may be placed near the bottom terminals (e.g., HfO2). These layers seal off the bottom open ends of the vertical hollow voids.

[0332] Fig. 18 illustrates a 3D memory structure 1800 that can be configured as a 3D NOR Vertical Transistor memory array. The 3D memory structure 1800 comprises a first vertical structure 1808a and an identical (or substantially identical) second vertical structure 1808b arranged adjacent to one another.

[0333] The first vertical structure 1808a includes a vertical plug column 1802a that provides an electrical connection to the lower portions of the 3D memory structure. The vertical plug column 1802a may have a uniform diameter along its entire length or may have a larger diameter on its lower end than on its upper end. The plug column 1802a can be fabricated as a solid column or as a hollow column in various embodiments.

[0334] Disposed adjacent to the vertical plug column 1802a is a source electrode column 1804a and a drain electrode column 1816a. The source electrode column 1804a and drain electrode column 1816a provide electrical connections to theAttorney Docket No. P25-059-SEC-W001source and drain nodes of the vertical transistors formed along the vertical structure 1808a. The source electrode column 1804a and drain electrode column 1816a may be comprised of various conducting materials including, but not limited to, tungsten, titanium nitride, tantalum nitride, nickel, molybdenum, platinum, palladium, cobalt, gold, aluminum, copper, hafnium, hafnium nitride, iridium, iridium oxide, ruthenium, ruthenium oxide, silicides, graphene, carbon nanotubes, doped polysilicon, indium tin oxide, silver, aluminum-doped zinc oxide, gallium, gallium arsenide, indium gallium zinc oxide, metal alloys such as AICu and TiW, and conducting polymers.

[0335] Surrounding the vertical plug column 1802a, source electrode column 1804a, and drain electrode column 1816a is a channel column 1812a that provides the semiconductor channel region for the vertical transistors along the first vertical structure 1808a. The channel column 1812a may be formed from materials including, but not limited to, indium gallium zinc oxide (IGZO), indium zinc oxide (IZO), zinc tin oxide (ZTO), aluminum zinc oxide (AZO), indium tungsten oxide (IWO), gallium zinc oxide (GZO), hafnium indium oxide (HIO), cadmium oxide (CdO), polysilicon, polygermanium, cadmium selenide (CdSe), copper indium gallium selenide (CIGS), crystalline silicon, crystalline germanium, gallium arsenide (GaAs), indium phosphide (InP), indium antimonide (InSb), silicon carbide (SiC), gallium nitride (GaN), zinc oxide (ZnO), pentacene, P3HT, polythiophene, PPV, graphene, carbon nanotubes (CNTs), methylammonium lead halides, cesium lead halides, lead sulfide (PbS), lead selenide (PbSe), cadmium selenide (CdSe), indium arsenide (InAs), and other semiconducting materials.

[0336] Surrounding the channel column 1812a is a ferroelectric column 1814a that provides the gate dielectric for the vertical transistors along the first vertical structure 1808a. The ferroelectric column 1814a may be comprised of ferroelectric materials including, but not limited to, perovskite oxides, lead zirconate titanate (PZT), barium titanate (BaTiO3), strontium titanate (SrTiO3), bismuth ferrite (BiFeO3), potassium niobate (KNbO3), lithium niobate (LiNbO3), lithium tantalate (LiTaO3), sodium bismuth titanate (NaO.5BiO.5TiG3), bismuth titanate (Bi4Ti3O12), bismuth zinc niobate (Bi(Znl / 2Til / 2)O3), bismuth lanthanum titanate (BiLaTiO3), bismuth nickel titanate (BiNiTiO3), PMN-PT, PLZT, neodymium-doped bismuth titanate (Bi4-xNdxTi3O12), hafnium-based oxides like hafnium oxide (HfO2) and doped hafnium oxide, tungsten bronze structure materials, barium strontium niobate (BSN), lead barium niobate (PBN), potassium tantalate niobate (KTN), bismuth titanateAttorney Docket No. P25-059-SEC-W001(Bi4Ti3O12). strontium bismuth tantalate (SBT), calcium bismuth niobate (CBN), organic ferroelectrics like PVDF, TrFE and P(VDF-TrFE) copolymers, aurivillius phase oxides, rare earth manganites like YMnO3, lanthanum-modified PLZT, nickel manganese oxide (NiMnO3), relaxor ferroelectrics like PMN, PST and PIN, multiferroic materials like TbMnO3, EuTiO3, SbSI, GeTe, SnTe, PZT thin films, SBT thin films, HfO2-based thin films, layered superlattices, and PbTiO3 / SrTiO3.

[0337] The 3D memory structure 1800 further comprises multiple horizontal gate electrode layers 1806 including layers 1806a, 1806b, 1806c etc. The horizontal gate electrode layers 1806 are disposed at regular intervals along the vertical structures 1808 and provide the gate electrodes for the vertical transistors. The gate electrode layers 1806 may be formed from materials such as tungsten, titanium nitride, tantalum nitride, nickel, molybdenum, platinum, palladium, cobalt, gold, aluminum, copper, hafnium, hafnium nitride, iridium, iridium oxide, ruthenium, ruthenium oxide, silicides, graphene, carbon nanotubes, doped polysilicon, indium tin oxide, silver, aluminum-doped zinc oxide, gallium, gallium arsenide, indium gallium zinc oxide, metal alloys such as AICu and TiW, and conducting polymers.

[0338] Each of the horizontal gate electrode layers 1806 may be surrounded by an oxide / nitride / oxide (ONO) stack 1810, such as 1810a surrounding gate electrode layer 1806a, to provide insulation between the gate electrodes.

[0339] The second vertical structure 1808b in the 3D memory structure 1800 is identically configured as the first vertical structure 1808a. The two vertical structures 1808a and 1808b are arranged horizontally adjacent to each other with a spacing that allows integration of the gate electrode layers 1806 and ONO stacks 1810. Together, the first and second vertical structures 1808a, 1808b along with the horizontal gate electrode layers 1806 can be configured as a 3D NOR memory architecture.

[0340] Fig. 19 illustrates an embodiment of a planar FeFET 1900. The FeFET 1900 comprises a substrate 1910 upon which various layers and components are formed. The substrate 1910 may be comprised of silicon or other suitable semiconductor materials. Disposed on top of the substrate 1910 is a layer of TiN 1912. The TiN layer 1912 may act as an electrode and can be deposited by sputtering or other suitable deposition techniques.

[0341] On top of the substrate 1910 and TiN layer 1912, a layer of HZO 1908 is disposed. HZO 1908 comprises hafnium, zirconium, and oxygen and can exhibit ferroelectric properties. The HZO 1908 may be deposited by ALD, CVD, PVD or otherAttorney Docket No. P25-059-SEC-W001suitable deposition methods and can have a thickness in the range of 5-50 nm. Acting as a ferroelectric layer, the HZO 1908 enables the non-volatile storage of data in the FeFET 1900.

[0342] Deposited conformally on top of the HZO 1908 is a layer of IWO 1906. IWO 1906 comprises indium, tungsten, and oxygen. It can be deposited by sputtering or other suitable techniques and may have a thickness in various ranges. The IWO 1906 layer serves as a control oxide layer in the FeFET 1900.

[0343] On top of the IWO 1906 layer, a drain contact 1904 and source contact 1914 are formed. The drain contact 1904 and source contact 1914 may comprise metals such as copper, aluminum, or alloys thereof and can be deposited by PVD, CVD or other suitable methods. The drain contact 1904 and source contact 1914 allow electrical connection to the FeFET 1900. They may have thicknesses in the range of 50-500 nm.

[0344] In operation, a voltage applied to the drain 1904, source 1 14, and TiN gate contact 1912 can control the ferroelectric polarization of the HZO 1908 layer. The polarization state can be used to store information in a non-volatile manner, enabling memory storage capabilities. The IWO 1906 layer helps improve the switching speed and endurance of the FeFET 1900. Overall, the layered structure shown in Fig. 19 enables a FeFET 1900 suitable for non-volatile memory applications.

[0345] Fig. 20 presents the transfer characteristics of a Ferroelectric FET (FeFET) device, illustrating the relationship between the gate voltage (V_GS) on the x-axis and the resulting drain current (I_D) on the y-axis. Fig. 20 may show the characteristics of a FeFET as disclosed herein. The x-axis spans from -IV to IV, while the y-axis, on a logarithmic scale, displays current values from 10A-12A / pm to 10A-4A / pm

[0346] Two distinct curves represent the drain current behavior under clockwise (CW) and counterclockwise (MW) polarization of the FeFET. The blue curve (CW) starts at approximately 10 -llA / pm at -IV and exhibits a steep increase around -I V, reaching just above 10 -5 A / pm at IV. This demonstrates the rapid increase in drain current exhibited by the FeFET under forward bias in the clockwise polarization state.

[0347] Conversely, the red curve (MW) starts at approximately 10A- 11 A / pm at -IV and increases more gradually as it approaches 0V. At around IV, it then follows closely with the blue curve past IV. This demonstrates comparable drain current behavior under reverse bias conditions regardless of the polarization state.Attorney Docket No. P25-059-SEC-W001

[0348] Notably, the separation between the red and blue curves spans several orders of magnitude in the negative voltage range near -IV. This substantial difference in off-state current highlights the non-volatile memory effect achievable with the FeFET depending on its polarization direction. This large memory window is explicitly called out in the green box labeled “Large Memory Window” at the top left.

[0349] Additional key details provided include the FeFET device dimensions, with a width / length ratio of lpm / 50nm specified. The drain voltage is also fixed at 0.05V. Specific points along the curves are annotated, such as “MW @5e-7A / pm =1 V” on the red MW curve denoting the IV memory window at 5xlOA-7A / pm drain current. Another point marked is “CW @-0.5V = lxlOA6” on the blue CW curve, highlighting the clockwise current value of lxlO-6A / pm at -0.5V gate voltage.

[0350] Fig. 20 thereby comprehensively depicts the bidirectional transfer characteristics of the FeFET device, highlighting the large memory window achievable through polarization switching and providing detailed voltage, current, and dimensional specifications to fully convey the measurement conditions and transistor performance. The paired curves effectively compare the clockwise and counterclockwise operation modes over the full gate voltage range.

[0351] Fig. 21 shows a block diagram illustrating ganging of modules in accordance with an embodiment of the present disclosure. Fig. 21 thus shows a memory system 2100 to manage data storage and retrieval using multiple memory modules 2104, 2106, 2108 and a shared read-address bus. The system 2100 includes a modules group 2102 comprising several memory modules, such as module 1 (2104), module 2 (2106), up to module N 2108. Each module is paired with its own read peripheral — read peripheral 1 (2110) for module 1, read peripheral 2 (2112) for module 2, and so on through read peripheral N 2114. These read peripherals 2120, 2112, 2114 may be used to enable independent and simultaneous read operations from their respective modules. As shown in Fig. 21, read peripherals can be ganged together. Ganged read modules may be read in parallel relative to other modules not in that particular ganged circuit.

[0352] The memory system 2100 may employ a single write peripheral 2111 that handles write operations across all modules in the modules group 2102. This write peripheral 2111 can receive write addresses, write data, and write clock signals, allowing it to write data to one or more modules based on the provided addresses. AllAttorney Docket No. P25-059-SEC-W001of the modules within the modules group 402 may be accessible by using a single address space, in some embodiments.

[0353] A feature of the system 2100 is the ganged read-address bus 2142, which facilitates coordinated read operations across multiple modules. This bus includes the read address [0:MSB-l] 2132, representing the lower bits of the read address shared among all modules, and the read address [MSB] 2136, the most significant bit that determines which module's data is accessed. The read clock 2134 synchronizes read operations, while the read data line 2138 carries the output data from the selected module. Output enable signals, such as output enable 1 (2122) for read peripheral 1 and output enable 2 (2130) for read peripheral 2, are controlled by the read address [MSB] 2136 to ensure that only the data from the selected module is placed on the read data bus at any given time, preventing collisions.

[0354] In some embodiments of the system 2100, the output enable signals for each read peripheral are derived from the full read address to control access to the shared read data bus 2138. Specifically, for read peripheral 1 (2110), the output enable 1 (2122) is directly connected to the read address |MSB| 2136 in some embodiments. This means that when the read address [MSB] is at a logic high level (T), output enable 1 is activated, allowing read peripheral 1 to place its data onto the read data bus.

[0355] For read peripheral 2 (2112), the output enable 2 (2130) is derived from the inverse of the read address [MSB] 2136. In this configuration, output enable 2 is activated when the read address [MSB] is at a logic low level ('0'). This inversion can be implemented using a NOT gate connected to the read address [MSB] line, ensuring that only one read peripheral is enabled at a time between read peripheral 1 and read peripheral 2.

[0356] For additional read peripherals, such as read peripheral N 2114, the output enable signals are derived using additional logic to interpret higher-order bits of the read address beyond the most significant bit. For instance, when more than two modules are present, multiple address bits can be utilized to uniquely select each module. A decoder circuit can be employed to generate the appropriate output enable signals based on the combination of the most significant bits in the read address. This decoding logic ensures that only the selected read peripheral’s output enable signal is activated, allowing its data to be placed on the shared read data bus 2138, while all other read peripherals remain inactive to prevent bus collision.Attorney Docket No. P25-059-SEC-W001

[0357] By leveraging additional address bits and decoding logic, the system 2100 can scale to support a larger number of modules with corresponding read peripherals. This approach maintains efficient control over module selection and data access, allowing the memory system to handle multiple modules using a shared readaddress bus architecture effectively. It ensures that read operations are properly coordinated, with only the desired module's data being accessible at any given time, thereby enhancing system flexibility and scalability.

[0358] The ganged read-address bus 2142 can provide a unified read address, split into two portions in a specific embodiment. That is, some bits may be common to all modules and one or more bits may select among the different modules. In the example shown in Fig. 21, the first portion of the read address is shared among all read peripherals, while the second portion acts as a selection mechanism, determining which specific read peripheral should be activated for a given read operation. This architecture allows for use of address lines and enables the system to access contiguous or overlapping memory spaces across multiple modules using a single address bus. The system 2100 can be configured with various enhancements, such as multiplexers for module selection, interlock mechanisms to prevent simultaneous read and write operations on the same module, and the ability to dynamically adjust the number of ganged read peripherals based on power consumption requirements or performance needs.

[0359] In an embodiment of the memory system 2100, a multiplexer (not shown in the figures is implemented to enhance module selection during read operations. This multiplexer is configured to receive the second portion of the read address, specifically the most significant bit(s) [MSB] 2136, and to generate enable signals for the outputenable ports of the read peripherals. The multiplexer decodes the MSB(s) of the read address to determine which specific read peripheral should be activated.

[0360] When a read operation is initiated, the full read address is supplied via the ganged read-address bus 2142. The first portion of the read address — comprising the least significant bits [0:MSB-l] 2132 — is distributed to the read-address ports 2116, 2124, etc. of all read peripherals 2110, 2112, 2114, specifying the address within each respective module 2104, 2106, 2108. Simultaneously, the second portion of the read address — the MSB(s) 2136 — is input to the multiplexer. The multiplexer deciphers this portion of the address and generates a corresponding enable signal that is provided to the output-enable port of the selected read peripheral.Attorney Docket No. P25-059-SEC-W001

[0361] For example, if the MSB indicates selection of the first module 2104, the multiplexer activates the output-enable signal 2122 for read peripheral 1 (2110) while keeping the output-enable signals for other read peripherals deactivated. This ensures that only read peripheral 1 places data onto the shared read data bus 2138. Similarly, if the MSB selects the second module 2106, the multiplexer activates the output-enable signal 2130 for read peripheral 2 (2112).

[0362] By utilizing the multiplexer in this manner, the memory system manages access to multiple modules using a shared read-address bus while preventing bus contention. The operational logic of the multiplexer allows for the system to scale beyond two modules. When additional modules are incorporated, the multiplexer can be designed with more select lines derived from additional MSBs of the read address, enabling precise selection among multiple read peripherals. This approach facilitates flexible memory configurations and simplifies the control logic required for module selection, enhancing the overall efficiency and scalability of the memory system 2100.

[0363] The benefits of employing a multiplexer in this context include reduced complexity in routing enable signals and minimizing the number of control lines required. It centralizes the decoding of the module selection logic, allowing for easier adjustments and expansions of the memory system. Additionally, it streamlines the process of enabling the appropriate read peripheral based on the read address provided, contributing to faster read operations and improved system performance.

[0364] The modules group 2102 comprises a plurality of memory modules, including module 1 (2104), module 2 (2106), up to module N 2108. These modules store data and can be accessed individually or collectively, allowing for scalable memory capacity and parallel processing. The modules group 2102 is connected to the write peripheral 2111 for write operations and to respective read peripherals 2110, 2112, 2114 for read operations. Each module in the group may have its own read and write address spaces. In some embodiments, modules may have overlapping read address spaces to facilitate specific data management strategies. The modules group 2102 enables flexible memory configurations to meet various application requirements for capacity, speed, and parallel access.

[0365] The write peripheral 2111 is configured to handle write operations for all modules in the modules group 2102. It receives write addresses, write data, and write clock signals to facilitate writing data to one or more modules based on the write address. The write peripheral 2111 may utilize a shared write logic system, potentiallyAttorney Docket No. P25-059-SEC-W001employing a shift register-based design. In some embodiments, the write peripheral 2111 can perform simultaneous or individual writes across multiple modules in the modules group 2102. The write peripheral 2111 may be designed to operate at higher voltages compared to other components to meet specific voltage requirements of the memory system. Additionally, the write peripheral 2111 can be configured to perform write operations at a slower speed relative to read operations handled by the read peripherals 2110, 2112, 2114.

[0366] The first module 2104 is one of the modules in the modules group 2102 of the memory system 2100. The first module 2104 may comprise a plurality of memory cells configured to store data. In some embodiments, the first module 2104 may include non-volatile memory cells such as ferroelectric field-effect transistors (FeFETs). The first module 2104 has a dedicated read peripheral 2110 associated with it to handle read operations. The read peripheral 2110 for the first module 2104 includes ports for receiving a read address 2116, outputting read data 2118, receiving a read clock 2120, and receiving an output enable signal 2122. The first module 2104 may share a common write peripheral 2111 with other modules in the modules group 2102 for write operations. In some embodiments, the first module 2104 may have a read address space that at least partially overlaps with the read address space of other modules, such as the second module 2106. The first module 2104 can be accessed for read operations via the ganged read-address bus 2142, which provides addressing signals to multiple modules simultaneously. The size and capacity of the first module 2104 may vary depending on the specific implementation requirements of the memory system 2100.

[0367] The second module 2106 is one of multiple modules in the modules group 2102 of the memory system 2100. It may have a dedicated read peripheral 2112 for reading data from the module. The second module 2106 can be accessed for write operations via the shared write peripheral 2111. In some embodiments, the second module 2106 may have overlapping read address space with the first module 2104 to facilitate specific data management strategies. The second module 2106 is connected to its read peripheral 2112 and the write peripheral 2111 to enable read and write operations. As part of the modules group 2102, it contributes to the overall memory capacity and functionality of the memory system 2100.

[0368] The Nth module 2108 represents an additional memory module in the memory system 2100 beyond module 1 (2104) and module 2 (2106). Like the other modules, the Nth module 2108 has its own dedicated read peripheral, specifically readAttorney Docket No. P25-059-SEC-W001peripheral N 2114. The Nth module 2108 can receive write operations from the shared write peripheral 2111 and can output read data via its dedicated read peripheral N 2114. The Nth module 2108 allows the memory system 2100 to be scaled up to include additional memory capacity as needed. In some embodiments, other modules may be interfaced with the ganged read-address bus 2142 in the same manner as the other modules, receiving read address signals and outputting read data under the control of its dedicated read peripheral. In some embodiments multiple ganged groups of modules may be implemented.

[0369] The first read peripheral 2110 is a component of the memory system 2100 that enables reading data from module 1 (2104). It receives inputs including a read address (0:M-l) 2116, read clock 1 2120, and output enable 1 (2122). The first read peripheral 2110 uses these inputs to access and retrieve data from module 1 (2104), outputting the read data on read data 1 2118. The read address (0:M-l) 2116 specifies which memory location to read from within module 1 (2104). The read clock 1 2120 provides timing for the read operation. The output enable 1 (2122) controls when data is output on read data 1 2118. The first read peripheral 2110 operates independently from other read peripherals, allowing concurrent read access to module 1 (2104) while other modules may be accessed simultaneously.

[0370] The second read peripheral 2112 is coupled to the second module 2106 and is configured to read data from the second module 2106. The second read peripheral 2112 includes a read address port to receive read addresses, a read data port to output read data, a read clock port to receive a read clock signal, and an output enable port. The read address port of the second read peripheral 2112 is coupled to the read address (0:M-l) lines 2124 of the ganged read-address bus 2142 to receive the lower bits of the read address. The output enable port is coupled to the read address [MSB] line 2136 to receive the most significant bit of the read address, which controls whether the second read peripheral 2112 is enabled to output data. The read clock port receives the read clock 2 signal 2128 to synchronize read operations. When enabled, the second read peripheral 2112 outputs read data from the second module 2106 on its read data 2 lines 2126.

[0371] The Nth read peripheral 2114 is configured to read data from module N 2108 of the modules group 2102. It can receive a read address (0:M-l). The Nth read peripheral 2114 may also receive a read clock signal from the ganged read-address bus to synchronize read operations. An output enable signal, derived from the mostAttorney Docket No. P25-059-SEC-W001significant bit of the read address, controls whether the Nth read peripheral 2114 outputs data onto the read data bus of the ganged read-address bus. This allows the Nth read peripheral 2114 to be selectively activated for reading from module N while avoiding bus contention with other read peripherals. The Nth read peripheral 2114 can operate concurrently with other read peripherals to enable parallel data access from multiple modules.

[0372] The read address (0:M-l) 2116 is a component within the memory system 2100 that provides addressing functionality for read operations on the first read peripheral 2110. It encompasses a range of address bits from 0 to M-l, where M represents the number of addressable locations within the associated module. This address input allows the system to specify which memory location should be accessed during a read operation, enabling precise data retrieval from the designated module.

[0373] The Nth read peripheral 2114 is configured to read data from module N 2108 of the modules group 2102. It includes a read address input, designated as read address (0:M-l) for read peripheral N 2144, which receives address signals specifying memory locations within module N. The read data output provides the data retrieved from these locations, and the read clock input synchronizes read operations, potentially receiving the shared read clock 2134 from the ganged read-address bus 2142, or it may receive a clock from a different circuit (e.g., another processing module, for example). The output enable input for read peripheral N can be controlled by the read address [MSB] 2136, its inverse, or another suitable control signal, ensuring that only one read peripheral outputs data onto the shared read data bus 2138 at any given time to prevent bus contention. In some embodiments, the Nth read peripheral 2114 may be individually accessed by another circuit using a dedicated read address bus, allowing independent operation from the ganged read-address bus. Alternatively, it may be ganged with other modules or groups of modules, enabling coordinated read operations across multiple modules.

[0374] The read address (0:M-l) 2116 interacts closely with other elements in the memory system to facilitate efficient read operations. It may receive address information from the read address [0:MSB-l] 2132 of the ganged read-address bus 2142, which provides a common addressing mechanism for two or more read peripherals. The address bits supplied by 2116 are utilized by the read peripheral 1 (2110) to locate and access the requested data within its corresponding module. This interaction enables the memory system 2100 to perform targeted read operations acrossAttorney Docket No. P25-059-SEC-W001different modules while maintaining a unified addressing scheme through the ganged read- address bus.

[0375] There are several potential variations and alternative embodiments for the read address (0:M-l) 2116. The bit width M can be adjusted to accommodate different module sizes and addressing requirements. In some implementations, the address input may incorporate error detection or correction codes to enhance reliability. Alternative embodiments may include address translation or remapping logic within the read peripheral to allow for flexible memory organization. The timing and synchronization of the address signals can also be varied, potentially using different clock domains or asynchronous designs. Additionally, the physical implementation of the address lines can be modified, such as using serial addressing protocols or multiplexed address / data buses, to optimize for different design constraints or system architectures.

[0376] The first read data 2118 is a data output component of read peripheral 1 within the memory system 2100. This can serve as the conduit for transmitting data that has been read from module 1 to external components or systems. The first read data 2118 can output digital information in various formats, such as parallel or serial data streams, depending on the specific implementation of the memory system 2100. It is designed to handle data transfers to accommodate read operations from the associated memory module.

[0377] The first read data 2118 can interact with other elements within the memory system 2100. It may receive data directly from module 1 via read peripheral 1, which processes the read requests initiated by the read address (0:M-l) 2116. The output of first read data 2118 is controlled by the output enable 1 signal 2122, which determines when data can be placed on the output bus. Additionally, the read clock 1 signal 2120 synchronizes the data output timing to ensure proper data capture by the receiving system. The first read data 2118 also interfaces with the read data of ganged read-address bus 2138, allowing its output to be combined with data from other read peripherals when multiple modules are accessed simultaneously.

[0378] In some configurations, the first read data 2118 may include error correction capabilities to enhance data integrity during transmission. Another variation could incorporate a small buffer to temporarily store read data, allowing for more flexible timing in data output. The width of the data path could be adjustable, ranging from a single bit to multiple bytes, to accommodate different system requirements.Attorney Docket No. P25-059-SEC-W001Some implementations may include multiple data outputs operating at different speeds or voltages to interface with various types of external systems. Additionally, the first read data 2118 could be designed with power-saving features, such as the ability to enter a low-power state when not in use, thus contributing to overall system energy efficiency.

[0379] The first read clock 2120 is a signal input to read peripheral 1 (2110) within the memory system 2100. This clock signal serves to synchronize read operations for module 1 (2104), ensuring that data retrieval occurs at the appropriate timing intervals. The first read clock 2120 operates in conjunction with other components such as the read address (0:M-l) 2116 and output enable 1 (2122) to facilitate coordinated data access from module 1.

[0380] In the context of the memory system 2100, the first read clock 2120 interacts closely with read peripheral 1 (2110) to govern the timing of read operations. It works in tandem with the read clock of the ganged read-address bus 2134, allowing for synchronized data retrieval across multiple read peripherals. The first read clock 2120 may be used to latch incoming read addresses, trigger the activation of sense amplifiers, and control the timing of data output from module 1 to the read data 1 bus 2118. This clock signal helps maintain proper timing relationships between the various components involved in the read process, potentially reducing the likelihood of data conflicts or timing-related errors.

[0381] Additional embodiments for the first read clock 2120 includes derivation of the first read clock 2120 from a global system clock and further divided or phase-shifted to meet specific timing requirements of read peripheral 1. It could be generated independently for each read peripheral, allowing for more flexible clock domain management. The clock signal may also be configurable in terms of frequency or phase, enabling dynamic adjustment based on system conditions or performance requirements. In some embodiments, the first read clock 2120 may incorporate features such as clock gating to reduce power consumption when read peripheral 1 is inactive. Additionally, advanced clock distribution techniques like H-tree or clock mesh networks may be employed to minimize clock skew across the memory system.

[0382] The first output enable 2122 is an input signal that controls the data output functionality of read peripheral 1. This signal determines whether read peripheral 1 is allowed to place its read data onto the read data of ganged read-address bus 2138. When activated, the first output enable 2122 permits read peripheral 1 toAttorney Docket No. P25-059-SEC-W001transmit the data retrieved from module 1 onto the shared read data bus. Conversely, when deactivated, it prevents read peripheral 1 from outputting data, effectively isolating it from the shared bus. The first output enable 2122 may be coupled to the most significant bit of the read address 2136.

[0383] The first output enable 2122 may interact with other components in the memory system to coordinate read operations. It receives its input from the read address [MSB] 2136 of the ganged read-address bus, which uses the most significant bit of the read address to select which read peripheral should be active. This enables the system to multiplex data from multiple read peripherals onto a single shared data bus. The first output enable 2122 also works in conjunction with the read clock 1 2120 and read address (0:M-l) 2116 to ensure that data is output at the correct time and from the correct memory location within module 1.

[0384] Various implementations and alternatives can be considered for the first output enable 2122. In some embodiments, it may be designed as an active-high signal, while in others it could be active-low. The timing and duration of its activation may be adjustable to accommodate different memory access patterns or system requirements. Additionally, the first output enable 2122 could be implemented with level-shifting circuitry to interface between different voltage domains if read peripheral 1 and the ganged read-address bus 2142 operate at different voltage levels. Further variations may include adding a delay element to fine-tune the timing of the output enable activation relative to other signals in the system.

[0385] The read address (0:M-l) 2124 is a component of the memory system 2100 that provides addressing functionality for read peripheral 2 (2112). This element represents the read address bus that carries address signals to specify which memory locations within the associated module should be accessed during read operations. The read address (0:M-l) 2124 can support a range of addresses from 0 to M-l, where M represents the maximum addressable space within the connected module.

[0386] The read address (0:M-l) 2124 interacts closely with other components in the memory system to facilitate efficient data retrieval. It receives address information from the read address [0:MSB-l] 2132 of the ganged read-address bus 2142, allowing for coordinated addressing across multiple read peripherals. The read address (0:M-l) 2124 works in conjunction with the read clock 2 2128 and output enable 2 (2130) to control the timing and activation of read operations for read peripheral 2 (2112). When an address is presented on the read address (0:M-l) 2124,Attorney Docket No. P25-059-SEC-W001and the output enable 2 (2130) is activated, the specified data can be retrieved from the associated module and output through the read data 22126.

[0387] Various implementations and alternatives of the read address (0:M-l) 2124 can be considered to enhance system flexibility and performance. For instance, the address width (M) can be adjusted to accommodate different module sizes or memory capacities. The read address (0:M-l) 2124 may incorporate pipelining or buffering techniques to improve address throughput. Additionally, error detection or correction mechanisms could be integrated into the address bus to enhance reliability. In some embodiments, the read address (0:M-l) 2124 might support burst mode addressing or other advanced addressing schemes to optimize data access patterns for specific applications.

[0388] The second read data 2126 represents the output data stream from read peripheral 2 in the memory system 2100. This data stream carries the information retrieved from module 2 based on the read address provided through the read address (0:M-l) for read peripheral 2 2124. The second read data 2126 is generated when the read clock 2 2128 and output enable 2 (2130) signals are activated, allowing the requested data to be transmitted from the memory module to the external system or processing unit requiring the information.

[0389] The second read data 2126 interacts with other components within the memory system 2100 to facilitate efficient data retrieval operations. It is controlled by the read address (0:M-l) for read peripheral 22124, which specifies the location of the desired data within module 2. The timing of the data output is controlled by the read clock 2 2128, ensuring synchronization with the system's overall timing requirements. Additionally, the output enable 2 (2130) signal gates the flow of data, preventing conflicts with other data streams on the shared read data of ganged read-address bus 2138. In some embodiments, the output enable 2 (2130) may be the inverse (e.g., a NOT-gate) of the most significant bit of the read address 2136. This interaction allows for coordinated and collision-free data retrieval across multiple read peripherals.

[0390] There are several potential variations and alternative embodiments for the second read data 2126. In some configurations, it may support multiple data widths, allowing for flexible data retrieval based on system requirements. Another variation could involve implementing error detection and correction mechanisms within the data stream to enhance reliability. Additionally, the second read data 2126 could be designed to support burst mode operations, enabling high-speed sequential data transfers. SomeAttorney Docket No. P25-059-SEC-W001embodiments may incorporate data compression techniques to optimize bandwidth utilization. Furthermore, the timing and synchronization of the second read data 2126 could be made adjustable to accommodate different system clock domains or to support dynamic frequency scaling for power optimization.

[0391] The second read clock 2128 is a timing signal input for read peripheral 2 within the memory system 2100. This clock signal provides the synchronization mechanism for read operations conducted through read peripheral 2, ensuring that data retrieval from the associated module occurs at the appropriate timing intervals. The second read clock 2128 coordinates the sampling of read data, enabling the read peripheral 2 to capture information from the memory module at precise moments, thereby maintaining data integrity during the read process.

[0392] The second read clock 2128 clocks the read peripheral 2 using the read address (0:M-l) 2124, read data 2 2126, and output enable 2 (2130) to output data. When a read operation is initiated, the second read clock 2128 synchronizes with the read address to ensure that the correct memory location is accessed at the right moment. It also coordinates with the output enable 2 signal to determine when valid data should be presented on the read data 2 output. Furthermore, the second read clock 2128 may be derived from or synchronized with the read clock 2134 of the ganged read-address bus, allowing for coordinated operations across multiple read peripherals within the memory system 2100. In some embodiments, the read clock 2128 may be directly coupled to the second read clock 2128 and the first read clock 2120.

[0393] The implementation of the second read clock 2128 can vary depending on system requirements and design considerations. In some embodiments, it may be a continuous clock signal, while in others, it could be gated or enabled only during active read operations to conserve power. The frequency of the second read clock 2128 can be adjustable, allowing for flexibility in read speeds to accommodate different operational modes or to match the capabilities of the connected memory module. Additionally, the clock signal may incorporate features such as phase adjustment or delay lines to fine-tune its timing relative to other signals in the read path, potentially enhancing the reliability of data capture in high-speed or variable-latency scenarios in some specific embodiments.

[0394] The second output enable 2130 is a control signal input for read peripheral 2 (2112) that regulates the output of read data from module 2 (2106) onto the read data 2138 of ganged read-address bus 2142. This signal determines whetherAttorney Docket No. P25-059-SEC-W001read peripheral 2 (2112) is allowed to place its data on the shared read data bus 2138. When activated, the second output enable 2130 permits read data 2 from read peripheral 2 to be transmitted to the read data of ganged read-address bus. When deactivated, it prevents read peripheral 2 from outputting data, effectively isolating its read path from the shared bus.

[0395] The second output enable 2130 can interact with other elements in the memory system to coordinate read operations. The second output enable 2130 may receive a control signal from the read address [MSB] of the ganged read-address bus, which uses the most significant bit (or the inverse) of the read address to select between multiple read peripherals. This can allow the system to access data from different modules using a single shared address bus. The second output enable 2130 may work in conjunction with other output enable signals, such as output enable 1 for read peripheral 1, to ensure that only one read peripheral at a time can place data on the shared read data 2138 of ganged read-address bus, preventing bus contention and data corruption.

[0396] In some embodiments, the second output enable 2130 may respond to multiple address bits rather than just the MSB, allowing for more granular control over module selection. The signal can potentially be augmented with additional logic to implement features like power gating, where entire read peripherals are disabled when not in use. Another variation could involve using a separate dedicated control line for the output enable signal instead of deriving it from the address bus. Additionally, the timing and synchronization of the second output enable 2130 may be adjusted to accommodate different system clock domains or to implement pipeline stages in the read path.

[0397] The read clock 2134 is part of the ganged read-address bus 2142 in the memory system 2100. It provides a shared clock signal to synchronize read operations across multiple read peripherals, such as read peripheral 1 (2110) and read peripheral 2 (2112). The read clock 2134 connects to the read clock 1 2120 input of read peripheral 1 and the read clock 2 2128 input of read peripheral 2. By supplying a common clock signal, the read clock 2134 enables coordinated timing for read accesses from different modules in the memory system. This shared clock arrangement can help facilitate parallel read operations from multiple modules using the ganged read-address bus architecture.Attorney Docket No. P25-059-SEC-W001

[0398] The read address [0:MSB-l] 2132 is a component of the ganged readaddress bus 2142 within the memory system 2100. This element represents the least significant bits of the read address, spanning from bit 0 up to but not including the most significant bit (MSB). The read address [0:MSB-l] 2132 is responsible for specifying the memory location within each module that is to be accessed during read operations. By sharing these address bits across multiple read peripherals, the system enables coordinated access to data stored in different modules.

[0399] In operation, the read address [0:MSB-l] 2132 may interact with other components of the memory system 2100. It is connected to the read address inputs 2116, 2124 of read peripherals 2110, 2112. allowing the same address to be simultaneously presented to two modules. This configuration enables parallel access to corresponding memory locations across different modules, potentially increasing the overall read bandwidth of the system. The read address [0:MSB-l] 2132 works in conjunction with the read address [MSB] 2136 to fully specify which module and location within that module should be read.

[0400] The number of bits the read address |0:MSB-l | 2132 encompasses can be adjusted based on the size and organization of the memory modules, allowing for scalability in the system design. In some implementations, the read address [0:MSB-l] 2132 can be further subdivided to enable more granular control over addressing within modules. Additionally, the system could be designed to allow dynamic reconfiguration of the address bit allocation, providing flexibility to adapt to different operational modes or memory configurations. Additional embodiments may also incorporate error detection or correction codes within the address bits to enhance the reliability of memory access operations.

[0401] The read address [MSB] 2136 is a component of the ganged readaddress bus in the memory system 2100. This element represents the most significant bit (MSB) of the read address, which is used to select between different read peripherals when accessing data from the modules group. The read address [MSB] 2136 is connected to the output enable inputs of the various read peripherals, such as output enable 1 (2122) for read peripheral 1 and output enable 2 (2130) for read peripheral 2. By controlling these output enable signals, the read address [MSB] 2136 determines which read peripheral is active and allowed to output data onto the shared read data bus 2138.Attorney Docket No. P25-059-SEC-W001

[0402] In operation, the read address [MSB] 2136 works in conjunction with other components of the ganged read-address bus to facilitate efficient data retrieval from multiple modules. When a read operation is initiated, the read address [0:MSB-l] 2132 specifies the address within a module, while the read address [MSB] 2136 selects which module's read peripheral should be activated. This arrangement allows the memory system to access data from different modules using a single address bus, potentially reducing the number of address lines required and simplifying the overall system design. The read address [MSB] 2136 enables the memory system to dynamically switch between different modules during read operations, providing flexibility in data access patterns.

[0403] There are several potential variations and alternative embodiments for the read address [MSB] 2136. In some implementations, additional address bits beyond the MSB can be used to select between a larger number of read peripherals, allowing for more modules to be addressed. Another variation might involve using the read address [MSB] 2136 in combination with additional control logic to implement more complex module selection schemes, such as interleaved addressing or multi-module reads. Additionally, the functionality of the read address [MSB] 2136 can be expanded to include error detection or correction features, enhancing the reliability of the memory system. Some embodiments may also incorporate programmable logic to allow dynamic reconfiguration of how the read address [MSB] 2136 maps to different read peripherals, providing adaptability to changing system requirements or different memory configurations.

[0404] The read data 2138 element in Fig. 21 represents the read data output of the ganged read-address bus within the memory system 2100. This data path carries the information retrieved from the selected memory module in response to a read operation initiated through the ganged read-address bus. The read data 2138 serves as a conduit for transferring the requested data from the activated read peripheral to external components or processing units that require access to the stored information. Its functionality is integral to the overall data retrieval process of the memory system, enabling the extraction of stored data for subsequent use or manipulation.

[0405] The Not read address[MSB] 2140 is an inverted version of the most significant bit (MSB) of the read address. It may be used in control logic to ensure correct selection of output enable signals for certain read peripherals, facilitating selection between modules during read operations. The Not read address[MSB] 2140Attorney Docket No. P25-059-SEC-W001is derived from the read address [MSB] 2136 of the ganged read-address bus. It is part of the ganged read-address bus 2142 and may be used to control output enable signals of certain read peripherals, such as read peripheral N.

[0406] The ganged read-address bus 2142 is coupled to multiple read peripherals and provides a shared read address to access data from different modules. It includes address lines that connect to the read address inputs of multiple read peripherals, allowing a single read address to be sent to multiple modules simultaneously. The ganged read-address bus 2142 enables efficient concurrent access to data across multiple modules by sharing address signals, while still allowing individual module selection through output enable signals. This configuration provides a flexible and scalable architecture for accessing data from multiple memory modules using a common addressing scheme.

[0407] Various alternatives and modifications can be devised by those skilled in the art without departing from the disclosure. Accordingly, the present disclosure is intended to embrace all such alternatives, modifications, and variances. Additionally, while several embodiments of the present disclosure have been shown in the drawings and / or discussed herein, it is not intended that the disclosure be limited thereto, as it is intended that the disclosure be as broad in scope as the art will allow and that the specification be read likewise. Therefore, the above description should not be construed as limiting, but merely as exemplifications of particular embodiments. And those skilled in the art will envision other modifications within the scope and spirit of the claims appended hereto. Other elements, steps, methods, and techniques that are insubstantially different from those described above and / or in the appended claims are also intended to be within the scope of the disclosure.

[0408] The embodiments shown in the drawings are presented only to demonstrate certain examples of the disclosure. And the drawings described are only illustrative and are non-limiting. In the drawings, for illustrative purposes, the size of some of the elements may be exaggerated and not drawn to a particular scale. Additionally, elements shown within the drawings that have the same numbers may be identical elements or may be similar elements, depending on the context.

[0409] Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. Where an indefinite or definite article is used when referring to a singular noun, e.g., "a," "an," or "the,” this includes a plural of that noun unless something otherwise is specifically stated. Hence, the termAttorney Docket No. P25-059-SEC-W001"comprising" should not be interpreted as being restricted to the items listed thereafter; it does not exclude other elements or steps, and so the scope of the expression "a device comprising items A and B" should not be limited to devices consisting only of components A and B. This expression signifies that, with respect to the present disclosure, the only relevant components of the device are A and B.

[0410] Furthermore, the terms "first," "second," "third," and the like, whether used in the description or in the claims, are provided for distinguishing between similar elements and not necessarily for describing a sequential or chronological order, ft is to be understood that the terms so used are interchangeable under appropriate circumstances (unless clearly disclosed otherwise) and that the embodiments of the disclosure described herein are capable of operation in other sequences and / or arrangements than are described or illustrated herein.

Claims

Attorney Docket No. P25-059-SEC-W001What is claimed is:

1. A memory system comprising :a modules group having a plurality of modules, the plurality of modules including a first module and a second module, the first module having a first read address space and a first write address space, the second module having a second read address space and a second write address space, wherein the first read address space at least partially overlaps with the second read address space;a write peripheral configured to receive a write address, write data, and a write clock signal, the write peripheral further configured to write the write data to one or more of the plurality of modules according to the write address;a plurality of read peripherals, the plurality of read peripherals including a first read peripheral and a second read peripheral, the first read peripheral coupled to the first module and the second read peripheral coupled to the second module, the first read peripheral having a first read-address port, a first read-clock port, and a first output-enable port, the second read peripheral having a second read-address port, a second read-clock port, and a second output-enable port; anda ganged read-address bus coupled to the first read peripheral and the second read peripheral, wherein the ganged read address bus is configured to provide a read address having a first portion of the read address and a second portion of the read address, wherein the first portion of the read address is coupled to the first readaddress port of the first read peripheral and to the second the second read-address port of the second read peripheral, wherein the second portion of the read address is coupled to the first output-enable port of the first read peripheral and to the second output-enable port of the second read peripheral whereby the memory system is configured to read from the first read peripheral and the second read peripheral using the ganged read-address bus.

2. The memory system according to claim 1, wherein a first memory space of the first module and a second memory space of the second module are accessible via the ganged read-address bus.

3. The memory system according to claim 2, wherein the first memory space and the second memory space are contiguous as accessed via the read address of the ganged read address bus.Attorney Docket No. P25-059-SEC-W0014. The memory system according to claim 1, wherein the second portion of the read address is a pre-determined bit of the provided read address.

5. The memory system according to claim 4, wherein the pre-determined bit of the provided read address is a most-significant bit of the read address.

6. The memory system according to claim 1 , wherein the ganged readaddress bus is formed by plurality of conductive paths.

7. The memory system according to claim 1, wherein the first portion of the read address comprises least significant bits corresponding to addresses within the first module and the second module, and the second portion of the read address comprises a most significant bit configured to select between the first read peripheral and the second read peripheral via their output-enable ports.

8. The memory system according to claim 1, wherein each read peripheral is configured to receive the first portion of the read address at its respective read-address port, the first portion corresponding to addresses within its associated module.

9. The memory system according to claim 1, wherein the write peripheral is configured to write the write data to multiple modules of the plurality of modules simultaneously based on the write address.

10. The memory system according to claim 1, wherein each module comprises a plurality of bit-cells, and each bit-cell comprises a ferroelectric fieldeffect transistor (FeFET).

11. The memory system according to claim 1, further comprising a multiplexer configured to select one of the plurality of modules based on a most significant bit of the read address.Attorney Docket No. P25-059-SEC-W00112. The memory system according to claim 1 , wherein the write peripheral is configured to perform write operations at a slower speed than read operations performed by the plurality of read peripherals.

13. The memory system according to claim 1, wherein the plurality of modules are configured in a three-dimensional arrangement.

14. The memory system according to claim 1 , wherein a third read peripheral is configured to allow read operations to be performed in parallel with read operations with the ganged read-address bus.

15. The memory system according to claim 1, wherein the write peripheral includes a shared write logic system utilizing a shift register-based design.

16. The memory system according to claim 1, further comprising an interlock mechanism configured to prevent simultaneous read and write operations to the same module.

17. The memory system according to claim 1, wherein the number of modules whose read peripherals are ganged via the ganged read-address bus is configurable.

18. The memory system according to claim 1, wherein the memory system is configured to dynamically adjust the number of modules whose read peripherals are ganged via the ganged read-address bus based on power consumption requirements.

19. The memory system according to claim 1 , wherein each read peripheral further comprises a local buffer configured to temporarily store read data before output.

20. The memory system according to claim 1 , wherein the memory system is configured to adjust read latency based on the number of modules whose read peripherals are ganged via the ganged read-address bus.Attorney Docket No. P25-059-SEC-W00121. The memory system according to claim 1, further comprising a multiplexer configured to receive the second portion of the read address and to generate an enable signal, wherein the enable signal is provided to a selected outputenable port of a respective read peripheral to read out data from a respective module of the respective read peripheral.

22. The memory system according to claim 1 , wherein each read peripheral is configured to receive a read clock signal, the read clock signal being synchronized across the plurality of read peripherals to coordinate read operations.

23. The memory system according to claim 1, wherein each read peripheral is configured to receive a read clock signal, the read clock signal being coupled to the first read-clock port and the second read-clock port.

24. The memory system according to claim 1, further comprising a logic circuit configured to invert a most significant bit of the read address prior to being sent to the second output-enable port of the second read peripheral.

25. The memory system according to claim 1, wherein the read data paths from the first and second read peripherals are configured to converge into a common data bus.

26. The memory system according to any one of claims 1 to 25, wherein each of the plurality of modules includes a respective write port and write peripheral.

27. The memory system according to any one of claims 1 to 26, wherein the plurality of modules are disposed on a chiplet that is face-to-face bonded on an application semiconductor.