Scaling bandwidth on high bandwidth memory devices and associated systems and methods
By increasing the number of TSV buses and adjusting the timing ratio, the synchronization problem of SiP and HBM devices when increasing bandwidth and data rate was solved, achieving higher data transmission efficiency and lower power consumption.
Patent Information
- Application Number
- CN202510611151.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-14
- Filing Date
- 2025-05-13
- Publication Date
- 2025-11-14
AI Technical Summary
When increasing bandwidth and data rate, existing SiP and HBM devices face the problem of asynchronous timing of memory array, TSV bus, and DQ bus, which leads to a decrease in timing tolerance and an increase in power consumption.
By increasing the number of TSV buses and adjusting the timing ratio tCCDL/tCCDS, the timing of the memory array is synchronized with the timing of the TSV buses, while keeping the DQ bus saturated. Multiple TSV paths are used for data transmission to match the total data rate, thus avoiding raising the TSV bus voltage.
It achieves increased bandwidth and data rate of HBM devices without changing memory array structure and power consumption, while maintaining memory array timing synchronization and avoiding timing tolerance decline and power consumption increase.
Smart Images

Figure CN120954459A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to vertically stacked semiconductor devices, and more specifically, to vertically stacked high-bandwidth memory devices for semiconductor packaging. Background Technology
[0002] Microelectronic devices, such as memory devices, microprocessors, and other electronic components, typically include one or more semiconductor dies mounted to a substrate and encased in a protective cover. These semiconductor dies contain functional features such as memory cells, processor circuitry, imaging devices, interconnect circuitry systems, etc. To meet the ongoing demand for size reduction, wafers, individual semiconductor dies, and / or active components are typically manufactured in batches, monolithically, and then stacked on a support substrate (e.g., a printed circuit board (PCB) or other suitable substrate). The stacked dies are then coupled to the support substrate (sometimes also called a package substrate) via through-substrate (silicon vias) between the dies and the support substrate. Summary of the Invention
[0003] According to one aspect of this disclosure, a system-in-package (SiP) device is provided. The SiP device includes: a substrate; a processing unit carried by the substrate; and a high-bandwidth memory (HBM) device carried by the substrate and electrically coupled to the processing unit, wherein the HBM device includes: an interface die including bus switching circuitry configured to select a TSV bus from a plurality of through-silicon via (TSV) buses, each TSV bus having a set of TSVs, and communicatively coupling a DQ bus having a set of DQ pins to the selected TSV bus; and one or more stacks carried by the interface die, each stack having one or more dies, wherein each die includes TSV bus selection circuitry configured to communicatively couple groups of the dies to the TSV bus selected by the bus switching circuitry of the interface die, the groups comprising one or more groups having memory arrays.
[0004] According to another aspect of this disclosure, a high-bandwidth memory (HBM) device is provided. The HBM device includes: an interface die operatively coupled to a host device, the interface die including bus switching circuitry configured to select a TSV bus from a plurality of through-silicon via (TSV) buses, each TSV bus having a set of TSVs, and communicatively coupling a DQ bus having a set of DQ pins to the selected TSV bus; and one or more stacks carried by the interface die, each stack having one or more dies, wherein each die includes TSV bus selection circuitry configured to communicatively couple groups of the dies to the TSV bus selected by the bus switching circuitry of the interface die, the groups comprising one or more groups having memory arrays.
[0005] According to another aspect of this disclosure, a method is provided. The method includes: selecting a TSV bus from a plurality of through-silicon via (TSV) buses, each TSV bus having a set of TSVs; communicatively coupling a DQ bus having a set of DQ pins to the selected TSV bus; and communicatively coupling groups of dies to the selected TSV bus, the groups comprising one or more groups having memory arrays, wherein the DQ bus corresponds to a pseudo-channel or channel of a high-bandwidth memory (HBM) device. Attached Figure Description
[0006] Figure 1 A partial schematic cross-sectional view of a system-in-package device consistent with this disclosure.
[0007] Figure 2 A simplified timing diagram of related technologies for data flow via TSV.
[0008] Figure 3 A block diagram of an embodiment of an HBM device consistent with this disclosure.
[0009] Figure 4A To be merged into Figure 3 A schematic block diagram of the bus switching circuit in the HBM device.
[0010] Figure 4B For can Figure 3 An example of a switch used in a bus switching circuit.
[0011] Figure 4C To be merged into Figure 3 A schematic block diagram of the TSV selection circuit in an HBM device.
[0012] Figure 5A A simplified timing diagram of the data flow through the TSV during a write operation, consistent with this disclosure.
[0013] Figure 5B A simplified timing diagram of the data flow through the TSV during a read operation, consistent with this disclosure.
[0014] Figure 6 A flowchart illustrating a method for communicatively coupling a DQ bus to a TSV bus, consistent with this disclosure.
[0015] The drawings are not necessarily drawn to scale. Furthermore, it should be understood that some drawings have been shown schematically and / or partially schematically. Similarly, for the purpose of illustrating some embodiments of the invention, some components and / or operations may be divided into different blocks or combined into a single block. Moreover, while the technology is permissible in various modifications and alternatives, specific embodiments have been shown in the drawings by way of example and are described in detail below. However, it is not intended to limit the technology to the specific embodiments described. Detailed Implementation
[0016] High data reliability, high-speed memory access, higher data bandwidth, lower power consumption, and reduced chip size are desirable characteristics for semiconductor memories. In recent years, vertically stacked memory devices have been introduced, often referred to as 2.5D (“2.5D”) memory devices when placed adjacent to a host device, or as 3D (“3D”) memory devices when stacked on top of a host device. Some 2.5D and 3D memory devices are formed by vertically stacking memory dies and using through-silicon (or through-substrate) vias (TSVs) to interconnect the dies. Memory dies can be “stacked” into groups, with each stack, designated by a stack ID (“SID”), containing one or more dies (e.g., four dies). The benefits of 2.5D and 3D memory devices include shorter interconnects (which reduce circuit delay and power consumption), a large number of vertical vias between layers (which allow for wide-bandwidth buses between functional blocks, such as memory dies, in different layers), and a significantly smaller footprint. Therefore, 2.5D and 3D memory devices contribute to higher memory access speeds, lower power consumption, and reduced chip size. Example 2.5D and / or 3D memory devices include hybrid memory cubes (HMCs) and high-bandwidth memory (HBMs). For example, HBM is a vertically stacked memory comprising dynamic random access memory (DRAM) dies and interface dies (which, for example, provide an interface between the DRAM dies of the HBM device and a host device). In the following description, the terms "stack" and "SID" are used interchangeably.
[0017] In a system-in-package (SiP) configuration, an HBM device can be integrated with a host device (e.g., a GPU, CPU, Tensor Processing Unit (TCU), and / or any other suitable processing unit) using a substrate (e.g., a silicon interposer, a substrate of organic material, a substrate of inorganic material, and / or any other suitable material that provides interconnection between the graphics processing unit (GPU) / computer processing unit (CPU) and the HBM device and / or provides mechanical support for the components of the SiP device). The HBM device and the host communicate through the substrate. Because the traffic between the HBM device and the host device resides within the SiP (e.g., using signals routed via the silicon interposer), higher bandwidth can be achieved between the HBM device and the host device than in conventional systems. In other words, the TSV interconnecting the DRAM dies within the HBM device and the silicon interposer integrating the HBM device with the host device enable (e.g., via a printed circuit board (PCB)) routing of a greater number of signals (e.g., a wider data bus) than is typically found between packaged memory devices and host devices. The high-bandwidth interface within a SiP (System-on-Package) enables large amounts of data to move rapidly between the host device (e.g., GPU / CPU / TCU) and the HBM (Hardware-Based Memory) device during operation. For example, a high-bandwidth channel can be approximately 1000 gigabits per second (GB / s, sometimes also called gigabits (Gb)). Therefore, once data is loaded into the HBM device, the SiP device can quickly complete computational operations. SiP devices are typically integrated with other electronics and / or the packaging substrate (e.g., PCB) of other SiP devices within the package system. It should be understood that such high-bandwidth data transfer between the host device and the memory of the HBM device can be advantageous in various high-performance computing applications, such as video rendering, high-resolution graphics applications, artificial intelligence and / or machine learning (AI / ML) computing systems, and other complex computing systems and / or various other computing applications.
[0018] However, market demand for SiP devices and / or HBM devices within them may present certain challenges. One such challenge is the requirement for SiP devices (and HBM devices within them) to continuously increase bandwidth and corresponding DQ pin data rates. Increased data rates mean that data paths within HBM devices operate under strict timing constraints. For example, the timing parameter t corresponding to two CLK cycles... CCDR Potential degradation is possible. Furthermore, higher bandwidth means faster operation of HBM devices (e.g., faster system clock frequency), which leads to increased power consumption. Therefore, it is necessary to increase the bandwidth on HBM devices while maintaining the same memory array timings, which will reduce t CCDR The CLK cycle is maintained at 2 CLK cycles, while keeping power consumption as low as possible.
[0019] As used herein, the terms “vertical,” “horizontal,” “upper,” “lower,” “top,” and “bottom” may refer to the relative orientation or position of a feature in a device oriented as shown in the figures. For example, “bottom” may refer to a feature positioned closer to the bottom of the page than another feature. However, these terms should be interpreted broadly to include devices with other orientations, such as inverted or tilted orientations, where top / bottom, above / below, above / below, up / down, and left / right may be interchanged depending on the orientation.
[0020] Furthermore, although this document is primarily discussed in the context of 2.5D HBM devices for use in SiP devices, those skilled in the art will understand that the scope of this disclosure is not limited thereto. For example, various components of the SiP devices described herein can also be implemented in 3D HBM devices and various other stacked semiconductor devices to help address problems associated with the high data rates discussed above. Therefore, the scope of this disclosure is not limited to any subset of the embodiments and is limited only to the limitations set forth in the appended claims.
[0021] Figure 1 This is a partial schematic cross-sectional view of a SiP device 100 according to an embodiment of the present disclosure. Figure 1 As described herein, the SiP device 100 includes a substrate 110 (e.g., a silicon interposer, another organic interposer, an inorganic interposer, and / or any other suitable substrate), and each is connected via a plurality of interconnect structures 140 (in... Figure 1 The host device 120 and HBM device 130 are integrated (e.g., carried and coupled to the upper surface 112 of the substrate 110) with the substrate 110. The interconnect structure 140 may be a solder structure (e.g., solder ball), a metal-to-metal joint, and / or any other suitable conductive structure that mechanically and electrically couples the substrate 110 to each of the host device 120 and HBM device 130. Furthermore, the host device 120 is coupled to the HBM device 130 via one or more communication channels 150 formed in the substrate 110. The communication channels 150 may include one or more routing paths formed into (or on) the substrate 110. Figure 1 (The two are illustrated in the image).
[0022] like Figure 1Further explanation is provided: the substrate 110 includes a plurality of external signal TSVs 116 and a plurality of external power supply TSVs 118 extending between the upper surface 112 and the lower surface 114 of the substrate 110. The external signal TSVs 116 can transmit signals (e.g., data, control signals, processing commands, etc.) between the host device 120 and / or the HBM device 130 and external components (e.g., a PCB integrated with the substrate 110, an external controller, etc.). The external power supply TSVs 118 provide power from an external power source to the host device 120 and / or the HBM device 130.
[0023] In the described environment, host device 120 may include various components, such as processing units (e.g., CPU / GPU / TCU, etc.), one or more registers, one or more cache memories, and / or various other components. For example, in the described environment, host device 120 includes host I / O circuitry 123 that can direct signals to and / or from the HBM device via a communication channel 150, which may include a DQ (data) signal 150a. Alternatively, host I / O circuitry 123 can direct signals to and / or from external components (e.g., controllers coupled to one or more external signals such as TSV 116, etc.) and / or from the external components.
[0024] HBM device 130 may include interface die 132 and one or more memory stacks 136 carried by interface die 132. Figure 1 The memory stack consists of four layers (described in the diagram). Each of the memory stacks 136 may contain one or more DRAM dies (…). Figure 1 (Not shown in the image). Each memory stack 136 may encompass the physical and / or logical arrangement of one or more dies and may be associated with a stack ID (SID). The HBM device 130 also includes one or more signal TSVs 138, each extending from the interface die 132 to the uppermost memory stack 136a. Figure 1 (The text describes two) one or more DQ signals TSV 137a, b ( Figure 1 The description includes two sets) and one or more power supplies TSV 139 ( Figure 1(One example is provided below). As further discussed below, DQ signals TSV 137a and 137b represent different paths (TSV0 and TSV1, respectively) from the DQ pin to the DRAM. One or more power supplies TSV 139 provide power (e.g., received from one or more external power supplies TSV 118) to each of the interface die 132 and the memory stack 136. Signals TSV 138, including TSVs for carrying control and address signals, and DQ signals TSV 137a, b, carrying the DQ signal, communicatively couple the corresponding memory die in each of the memory stack 136 to the HBM memory controller circuitry 133 in the interface die 132 (and various other circuits in the interface die 132). Furthermore, the HBM memory controller circuit 133 can direct DQ, control, and / or address signals to the host device 120 and / or external components (e.g., one or more external storage devices coupled to external signals such as TSV 116) and / or from the host device and / or the external components. In some embodiments of this disclosure, the HBM device 130 may include a bus switch circuit 135 between the HBM memory controller circuit 133 and the TSV 138. As further discussed below, the bus switch circuit 135 can receive DQ, control, and / or address signals from the HBM memory controller circuit 133 and select the TSV path to be used by the DQ signal when communicating between the HBM device 130 and the host device 120 (and / or external devices).
[0025] Additional details regarding HBM devices, SiP devices having HBM devices, and associated systems and methods are set forth below. For ease of reference, simplified assemblies of semiconductor packages (and their components) are described herein. However, it should be understood that semiconductor assemblies (and their components) may be moved to different spatial orientations and used in different spatial orientations without altering the structure and / or function of the disclosed embodiments of the invention. Additionally, embodiments of semiconductor packages (and their components) are sometimes described herein with reference to control, read, and / or write signals. However, it should be understood that other terms may be used to describe signals and / or other types of signals not discussed may be used in the embodiments without altering the structure and / or function of the disclosed embodiments of the invention.
[0026] Figure 2Timing diagram 200 illustrates the relevant SiP technology, showing data transfer using a set of TSVs (“TSV bus”) during a write operation. This timing diagram may correspond to a relevant HBM device with a data rate of 8 Gbps. For simplicity, a read timing diagram is not shown. As used herein, “TSV bus” may refer to one or more TSVs carrying DQ signals. For example, depending on the context, TSV bus may refer to all TSVs or a subset of TSVs in an HBM device (e.g., TSVs corresponding to channels, pseudo-channels, etc.). Figure 2 As seen, the frequency of the system clock CLK determines the frequency of the write clock WCK, which can be, for example, twice the system CLK frequency. The WCK signal uses, for example, Double Data Rate (DDR) to provide timing for data transfer. That is, data transfer occurs on both the rising and falling edges of the WCK clock.
[0027] The CLK signal determines the duration of timing parameters, such as column access timing parameters t, which can be set according to the standard configuration of the HBM device. CCDL t CCDS and t CCDR Timing parameter t CCDL Timing parameter t represents the read / write (RD / WR) command latency between different groups (BAs) within the same group (BG). CCDS For the RD / WR command delay between different BGs, and the timing parameter t CCDR This refers to the delay of RD commands between different SIDs. As seen in timing diagram 200, the timing parameter t... CCDL It is set to 4 CLK cycles, and the timing parameter t CCDS It is set to 2 CLKs. Timing parameters are part of the interface protocol between the host device and the HBM device, and the HBM device can provide the host device with timing requirements for scheduling memory operations. In other words, the HBM device can let the host device know the timing parameters (e.g., t) used for timing parameters. CCDL and t CCDS The CLK cycle setting. When communicating with an HBM device, the host device observes any limitations in the timing parameters. For example, based on t CCDL Timing parameters, the host device will not be in the same t CCDL Within a CLK cycle, read or write commands are scheduled to groups within the same group. That is, after sending a command (e.g., read, write, etc.) to a group within a group, the host device will wait for t to be scheduled to another read or write command in the same group. CCDL CLK cycle (e.g., 4 CLK cycles in a related SiP technology). Regarding the timing parameter t... CCDSAfter a read or write command is executed on a group within a group, the host device will wait for t to be executed before scheduling another read or write command to a group in a different group. CCDS CLK loop. When a memory command is scheduled to the HBM device, the host device will not violate the timing protocol. That is, the host device will wait for at least the number of loops specified by the timing parameters (e.g., some timing parameters specify a minimum number of loops between certain types of commands) before issuing a series of commands involving timing parameters. Those skilled in the art will understand the interface protocol between the host device and the HBM device, and therefore, for the sake of brevity, it will not be discussed further except as necessary to explain embodiments of this disclosure.
[0028] As seen in timing diagram 200, t CCDL The CLK cycle period is set to 4 CLK cycles, and the timing parameter t CCDS It is set to 2 CLKs, and t CCDS The CLK cycle period is set to 2 CLKs. Timing parameters are configured to ensure timing synchronization of the memory array on the die, timing via the TSV bus, and timing via the DQ bus to ensure proper operation of the HBM device. For example, in a related technology HBM device with a CLK frequency of 2 GHz and a bit rate of 8 gigabits per second (Gbps) (using a burst length of 8), t CCDL The CLK cycle period is set to 4 CLK cycles, and t CCDS The CLK cycle is set to 2 CLKs to synchronize data transfer between the HBM device and the host device, in order to keep the DQ bus saturated (e.g., for the DQ bus of PC0, channel 0). That is, as... Figure 2 As seen, to maintain an 8Gbps rate, the DQ bus corresponding to the channel or pseudo-channel can be used for write operations every two CLK cycles (e.g., a new set of 32 bytes of pseudo-channel data can be transferred on the DQ bus every two CLK cycles). Similarly, for read operations (not shown), the DQ bus can be used to receive a new 32 bytes of pseudo-channel data every two CLK cycles.
[0029] like Figure 2 As seen in t CCDL During a CLK cycle (4 CLK cycles), two BGs can be accessed, such as group 2 in BG3 and group 3 in BG7. Once a W1 write command is issued to group 2 in BG3, the host device (e.g., host device 120) will wait for t before issuing a W2 write command to group 3 in BG7. CCDS CLK cycle (2 CLK cycles). Depending on how the group is arranged in the HBM device, the BG can be in the same stack or in different stacks (also referred to as "SID" in this document). Figure 2 As seen, the two write commands for BG3 and BG7 cost t. CCDL CLK loop (4 CLK loops). Therefore, after scheduling the W1 write command to BG3, t CCDL The CLK cycle allows the host device to schedule another write command to a different group in BG3 if needed. The t command follows the first command. CCDL The host device will not issue another command to the same group until the CLK cycle is complete.
[0030] For illustrative purposes, it is assumed that BG3 and BG7 use the same TSV bus (the same group of TSVs) to communicate with the DQ bus. Furthermore, for clarity, the W1 and W2 data streams are identified by hash lines traveling in different directions. At time T0, based on a write command W1 to group 2 of BG3 with BL of 8, 32 bytes of data are transferred to the DQ bus (e.g., the DQ bus for PC0, CH0) using 2 CLK cycles (4 WCK cycles). At time T1, the W1 data is transferred to group 2 via the TSV bus communicatively coupled to BG3. Figure 2 As seen, the transmission cost t to group 2 of BG3 is... CCDS CLK cycle (2 CLK cycles). Still at time T1, based on the write command W2 to group 3 of BG7, after the W1 data transfer to the DQ bus is complete, 32 bytes of data are transferred to the DQ bus. At time T2, the W1 data is transferred via the TSV bus used for BG3. The W1 data transfer via the TSV bus takes t... CCDS During the CLK cycle (two CLK cycles), the TSV bus is freed up for another transmission. At time T2, W2 data is transmitted via the communication-coupled TSV bus of BG7. Figure 2 In the relevant technical systems, groups are accessed sequentially, and the HBM device uses 4 CLK cycles of t. CCDL CLK cycle period and t for 2 CLK cycles CCDS The CLK cycle is used to ensure the synchronization of memory array timing, TSV bus timing, and DQ bus timing, so that data is not lost and the DQ bus is saturated.
[0031] However, it is necessary to increase the communication bandwidth between the host device and the HBM device on, for example, communication channel 150 (e.g., from 8Gbps to greater than 8Gbps, such as 16Gbps, 24Gbps, 32Gbps or greater). To achieve this, for example, a duration t CCDLDuring this period, more BGs (e.g., per channel or per pseudo-channel) are opened for read / write operations, and the data rate at the DQ pin can be increased accordingly. However, a potential problem is that because the data path in an HBM device operates under strict timing tolerances, increasing the data rate at the DQ pin can lead to a decline in timing tolerances. That is, the increased data rate may mean that the memory array timing, TSV bus timing, and / or DQ bus timing are no longer synchronized. A solution could be t CCDS and t CCDR The CLK cycle period (e.g., setting it to 3 or 4 CLK cycles instead of 2 CLK cycles) ensures that data is not lost during transfers to / from the DQ bus, which is based on external requirements. CCDS Timing operation of CLK cycles (2 CLK cycles). However, by waiting for additional CLK cycles, data transfer in HBM devices may be less efficient because the DQ bus may no longer be saturated (e.g., gaps or bubbles may exist when there is no data to be processed).
[0032] Another potential problem is that the TSV bus must be able to handle increased data rates. One solution is to increase the TSV bus timing frequency to increase the data rate through the TSV bus, but this would mean increasing the clock voltage. If the clock voltage increases, the use of low-swing signaling may no longer be an option, as the TSV bus voltage may not have enough time to swing between low and high. Therefore, increasing the TSV bus timing frequency is not desirable, as power consumption in HBM devices will also increase.
[0033] Furthermore, the memory array timing is configured such that read / write operations on the BG require access to the TSV bus within a predetermined time period. For example, related HBM devices can... CCDL During the CLK cycle, read / write operations are performed at a data rate of 8Gbps on both BGs (see [link]). Figure 2 For each read / write operation, the memory array timing requires taking an appropriate TSV bus within two CLK cycles (1ns) before the TSV bus can be released for the next read / write operation. Therefore, the related technology in HBM devices... CCDL The CLK cycle period is set to 4 CLK cycles (2ns) to accommodate t CCDL Two BGs are opened during the CLK cycle. Therefore, in t of 4 CLK cycles CCDL The CLK cycle period (duration of 2ns) and t for 2 CLK cycles CCDS In the case of a CLK cycle period (duration of 1ns), the memory array timing is synchronized with the TSV bus timing and DQ bus timing in the related HBM device.
[0034] Compared to related HBM devices, embodiments of this disclosure achieve increased bandwidth. To increase bandwidth, the bandwidth at t can be increased. CCDL The number of BGs accessed during the CLK cycle (e.g., per channel or per pseudo-channel), and can be appropriately increased by t. CCDL The CLK cycle period is adjusted to accommodate the increased number of BGs. For example, to double the bandwidth, the CLK cycle period can be adjusted per ton. CCDL The number of BGs in the CLK cycle period is increased (e.g., per channel or per pseudo-channel) to 4 BGs, and t CCDL The CLK cycle can be set to 8 CLK cycles. Additionally, the CLK clock frequency can be doubled, resulting in a data rate of 16Gbps at the DQ pin. However, at higher clock frequencies, the memory array timing and TSV bus timing will become asynchronous. For example, if the data rate is doubled from 8Gbps to 16Gbps, where t CCDL The CLK cycle has a period of 4 CLK cycles and t CCDS The CLK cycle has a period of 2 CLK cycles, then t CCDL The duration will change from 2ns to 1ns and t CCDS The duration will change from 1 ns to 0.5 ns. As discussed above, when t CCDL The duration is 2ns and t CCDS The memory array timing is synchronized for a duration of 1 ns. The memory array may not be able to cycle through an increasing number of groups within less than 2 ns. A solution could be to... CCDS The duration of the CLK cycle is increased to 4 CLK cycles and / or the timing parameters in the memory array are appropriately modified to accommodate the higher frequency of the TSV bus. However, t DDCS The change of the CLK cycle period to 4 CLK cycles also means t CCDR The CLK cycle period would need to be changed to 4 CLK cycles, which is undesirable because the DQ bus will not saturate, as discussed above. Furthermore, redesigning the memory array architecture is also undesirable due to complexity, and therefore may be infeasible or cost-inefficient. Therefore, it is necessary to increase the bandwidth on the HBM device while maintaining the same memory array timing and... CCDR Maintain two CLK cycles to keep the DQ bus saturated. Additionally, power consumption on the HBM device needs to be kept as low as possible.
[0035] Allow t CCDLA potential option for maintaining a CLK cycle of 4 CLK cycles (1 ns duration) is to simultaneously open two memory groups for access. This option keeps the memory array timing synchronized and also accommodates increased data rates. However, such a design means that the two memory groups are fixedly paired and must be accessed as a single cell. This configuration effectively reduces the number of independently addressable memory groups and therefore reduces the flexibility of the HBM device's memory scheduler in selecting memory groups during read / write operations. Therefore, it is necessary to increase the bandwidth of the HBM device without changing the memory array architecture of the relevant technology HBM device (e.g., HBM devices conforming to the JEDEC standard, High Bandwidth Memory DRAM (HBM4) specification) and / or changing the number of addressable memory groups. Additionally, it is also necessary to increase the t CCDR Maintaining two CLK cycles keeps the DQ bus saturated and keeps power consumption on the HBM device as low as possible.
[0036] In embodiments of this disclosure, it is possible to t CCDL During the CLK cycle, three or more BGs are opened (e.g., per channel or per pseudo-channel) to increase the bandwidth of the HBM device. Additionally, t CCDL The CLK cycle period can be extended accordingly (e.g., to 8 CLK cycles, 12 CLK cycles, 16 CLK cycles, etc.) to accommodate a larger number of BGs, and t CCDS and t CCDR The CLK cycle period can be set to two CLK cycles to keep the DQ bus saturated. This helps synchronize TSV bus timing and memory array timing, instead of keeping the TSV bus timing at t as in existing technology devices. CCDS The CLK cycle, in exemplary embodiments of this disclosure, sets the TSV bus timing to correspond to t. CCDL / t CCDS The ratio (hereinafter referred to as "time ratio" or "time ratio t") CCDL / t CCDS The timing ratio tK is reduced to a CLK cycle period, which provides the memory array with more access time to the TSV bus. In some embodiments, the timing ratio tK is reduced to a CLK cycle period. CCDL / t CCDS The addition of this can be used to change the firmware and / or Basic Input / Output System (BIOS) of an HBM device. Timing ratio t CCDL / t CCDS The addition of indicates a change in the specification or interface between the HBM device and the host device.
[0037] Except for timing ratio t CCDL / t CCDSIn addition, the total data rate through the TSV must be kept the same as the total data rate of the DQ bus without incurring certain disadvantages (e.g., increased TSV bus voltage). That is, embodiments of this disclosure increase the number of available TSV data paths (e.g., per channel and / or pseudo-channels), allowing a larger amount of data to be transmitted via the TSV at any given time. By using multiple TSV data paths, the DQ signal on consecutive commands (read or write) can use individual TSV paths in a "pipeline" type arrangement. This provides more transmission time for the DQ signal between the DRAM and the DQ bus, and therefore, the data rate on a given TSV data path can be lower than the data rate of the DQ bus, while the data rate across all TSV paths matches the data rate of the DQ bus. Therefore, in embodiments of this disclosure, the data rate (and corresponding voltage) through individual TSVs or the TSV bus can be kept low enough to allow low-swing signaling while still maintaining the total data rate on the TSV equal to the total data rate of the DQ bus.
[0038] For example, in some embodiments, the HBM device may have a data rate of 16 Gbps, where the system clock CLK frequency is 4 GHz. The number of open BGs (e.g., per channel or per pseudo-channel) may be 4 to accommodate the increased bandwidth, and t CCDL The CLK cycle can be set to 8 CLK cycles (2ns) to accommodate 4 BGs. Additionally, in some embodiments, to maintain the total data rate through the TSV the same as the data rate through the DQ bus, additional TSV paths are added, for example, to each channel and / or pseudo-channel in the HBM device. Furthermore, in some embodiments, t CCDS The CLK cycle period is maintained at 2 CLK cycles (0.5ns) to keep the DQ bus saturated and synchronized with external communication, and can be set to a timing ratio t. CCDL / t CCDS The TSV bus timing can correspond to 4 CLK cycles (1ns). The 1ns TSV bus timing will be identical to that of a related technology HBM device operating at 8Gbps. Therefore, there is no need to change the memory array timing to accommodate the higher bandwidth of embodiments of this disclosure. Additional details of embodiments of this disclosure are discussed below.
[0039] In the following discussion, reference will be made to DQ pins, channels, pseudo-channels, and corresponding TSVs. Those skilled in the art will understand that, depending on the architecture of the HBM device, the number of TSVs per DQ pin can be in some relationship other than a one-to-one ratio. For example, based on a burst length (BL) of 8, there can be 8 TSVs per DQ pin. Depending on the design, other HBM devices may have other TSV / DQ pin ratios, such as 4 TSV / DQ pins, 1 TSV / DQ pin, etc. Therefore, although the following discussion focuses on the TSV bus and DQ pins, those skilled in the art will understand that more than one TSV can correspond to a DQ pin, even if not explicitly stated.
[0040] In some embodiments, a TSV bus comprising one or more TSVs may be associated with a DQ bus in an HBM device having a set of DQ pins. The DQ bus may correspond to, for example, a channel, a pseudo-channel, or some other grouping of DQ pins. In some embodiments, more than one TSV bus may be associated with a DQ bus (e.g., a channel, a pseudo-channel, etc.). Associating more than one TSV bus with each DQ bus (e.g., a channel, a pseudo-channel, etc.) provides more transmission paths for data, allowing for slower data rates across each TSV or TSV bus, while the data rate across all TSVs is equal to the data rate of the DQ bus. In some embodiments, N TSV buses may exist for each DQ bus (e.g., a channel, a pseudo-channel, etc.), where N is a positive integer greater than 1. For example, as further discussed below, in some embodiments, each pseudo-channel PC0 or PC1 may be associated with two TSV buses TSV0 and TSV1 (e.g., TSV0 and TSV1 for PC0, and TSV0 and TSV1 for PC1).
[0041] Figure 3 illustrate Figure 1 Block diagram of HBM device 130. Figure 3 The illustrated embodiment has a 4N architecture, wherein the HBM device 130 includes four stacked SID0 to SID3, which can be connected to... Figure 1 The stacks 136 are identical, and each of stacks SID0 to SID3 (labeled 302a to d respectively) may contain four DRAM dies DIE0 to DIE3 (dies DIE0 in each stack are labeled 310a to d respectively, and dies DIE1 to DIE3 in each stack are collectively labeled 312a to d respectively). However, other embodiments may have other arrangements with fewer or more stacks and / or dies. For example, in some embodiments, the number of stacks and / or dies may be 1, 2, or 3.
[0042] Each die 310a to d and 312a to d may have one or more channels providing independent data access to one or more groups of a memory array (not shown). For example, in Figure 3 In the embodiments, channels 0 and 1, and their corresponding pseudo-channels PC0 and PC1, are shown extending through stacks 302a to d. Dies 310a to d in each stack have groups BG0 320 and BG1 322 communicatively coupled to channel 0 (for clarity, only BG0 and BG1 in stacks 302a and die 310a are labeled), and groups BG2 324 and BG3 326 communicatively coupled to channel 1 (for clarity, only BG0 and BG1 in stacks 302a and die 310a are labeled). Each group 320, 322, 324, 326 may contain one or more memory groups (e.g., eight memory groups) each containing one or more memory arrays. Other channels 2 to 7 (not shown) have a similar configuration but are communicatively coupled to different groups in different dies. For example, other channels may be coupled to BG4 to BG15.
[0043] In some embodiments, each channel 0 to 7 may be split into two semi-independent pseudo-channels, for example, pseudo-channel PC0 corresponding to DQ bits 0 to 31 and pseudo-channel PC1 corresponding to DQ bits 32 to 64. Channels and / or pseudo-channels can provide independent access to corresponding BGs, where each BG may contain one or more groups. For example, if the die has 16 groups, then each BG may have four groups, and independent channels can provide access to said BG. The die may contain fewer than 16 groups, such as 4 groups, 8 groups, etc. In some embodiments, the die may contain more than 16 groups. Similarly, the number of BGs in the die may be less than or greater than four. Dividing memory devices into groups and groupings is known in this art, and therefore, for the sake of brevity, will not be discussed further. Furthermore, those skilled in the art will understand that HBM devices may have different arrangements regarding the number of dies, groups, groupings, channels, and / or pseudo-channels than in the disclosed embodiments, while still being consistent with this disclosure.
[0044] The following description focuses on pseudo-channel PC0 in SID0 302a and DIE0 310a. However, the description applies to pseudo-channel PC1, other stacks 302b to d, and other bare dies 310b to d and 312a to d, and therefore will not be repeated for the sake of brevity and clarity. Figure 3As seen, groups 320, 322, 324, and 326 are each split into two groups, each corresponding to a different pseudo-channel (PC0 or PC1). Groups 320 and 322 for PC0 are selectively and communicatively coupled to either the TSV0 bus or the TSV1 bus for PC0 in channel 0. The TSV selection circuit 332 (for clarity, only the TSV selection circuit for PC0 in stack 302a of die 310a is labeled) selects which bus (TSV0 or TSV1) groups 320 and 322 are communicatively coupled to based on an enable signal from bus switch circuit 135, as discussed below. During read or write operations on groups 320 or 322, the TSV selection circuit 332 ensures that the groups are in a state of read or write operation based on a timing ratio t. CCDL / t CCDS Within the CLK cycle, it can access the TSV bus (TSV0 bus or TSV1 bus). CCDL / t CCDS During a CLK cycle, another group cannot communicatively couple to the TSV bus (TSV0 or TSV1) that is in use. However, the TSV bus that is not currently in use can be accessed by another group.
[0045] BG selection circuit 334 (for clarity, only the BG selection circuit for PC0 in stack 302a of die 310a is labeled) selects which group (e.g., 320, 322) should be communicatively coupled to TSV selection circuit 332. In some embodiments, the determination of which BG should be communicatively coupled to which TSV bus (TSV0 or TSV1) may be performed in bus switching circuit 135 (and / or another circuit in the HBM device) based, for example, SID, BG, and / or BA information from read / write commands from HBM memory controller circuit 133. BG selection circuit 334 ensures that at any given time only one of group 320 or 322 is communicatively coupled to TSV selection circuit 332. BG selection circuit 334 also ensures that at t CCDL The same group is not accessed within a CLK cycle. The operation description for groups 320 and 322 corresponding to PC1, and other groups 324 and 326, will be similar to the operation description for groups 320 and 322 for PC0; therefore, for brevity, it will not be discussed further. Additionally, the group configurations in other dies 310b to d and 312a to d, and in other stacks 302b to d, are similar, and therefore will not be discussed further for brevity. Although Figure 3The embodiments shown depict two BGs first communicatively coupled to a BG selection circuit, which is then communicatively coupled to a TSV selection circuit. However, in other embodiments, based on this arrangement, each BG can be directly communicatively coupled to the TSV selection circuit without the intervention of the BG selection circuit. Those skilled in the art will understand that the grouping and group numbering and specific configurations may differ. Figure 3 The numbering and specific configurations shown in this document are not applicable to other group configurations.
[0046] In related technology systems, each channel (when a pseudo-channel is not used) or each pseudo-channel includes one TSV bus as needed. However, in exemplary embodiments of this disclosure, each channel (when a pseudo-channel is not used) or pseudo-channel has N number of TSV buses that can be selectively accessed, where N is an integer greater than 1 (e.g., 2, 3, 4, etc.). In some embodiments, N may be limited to an even number to simplify the design of sequential circuitry. That is, each channel or pseudo-channel may have an even number of TSV buses to be selected from. As further discussed below, when t CCDL When more BGs are opened during the CLK cycle to increase bandwidth, the additional TSV bus, along with the corresponding timing ratio t, CCDL / t CCDS The TSV bus timing can provide different data paths to help relax timing constraints on the TSV bus.
[0047] For the sake of brevity, the following describes an embodiment with pseudo-channels, where each pseudo-channel has two corresponding TSV buses. However, those skilled in the art will understand that the concepts discussed below also apply to embodiments where the channel is not split into pseudo-channels and / or more than two TSV buses are associated with pseudo-channels or channels.
[0048] like Figure 3 As seen, each channel comprises two pseudo-channels, PC0 and PC1, and each pseudo-channel PC0 and PC1 contains a TSV0 bus (represented by solid lines) and a TSV1 bus (represented by dashed lines). For clarity, only the TSV0 and TSV1 buses for each pseudo-channel used for channels 0 and 1 are shown, but those skilled in the art will understand that other pseudo-channels may also contain TSV0 and TSV1 buses. Figure 3As seen, the bus switch circuit 135, together with the HBM memory controller circuit 133, is located in the interface die 132. However, some or all of the functions of the bus switch circuit 135 may be incorporated into the stacked die, the HBM memory controller circuit 133, and / or another circuit. The HBM memory controller circuit 133 controls external access to the DQ bus and manages the DQ signals to and from the bus switch circuit 135 based on memory operations (e.g., read, write, etc.). The configuration and operation of the HBM memory controller circuit are known to those skilled in the art, and therefore will not be discussed further for the sake of brevity. The bus switch circuit 135 is communicatively coupled to the HBM memory controller circuit 133 to receive / transmit DQ signals for each pseudo-channel from / to the HBM memory controller circuit 133, and is selected and communicatively coupled to the appropriate TSV bus (TSV0 bus or TSV1 bus) based on the address, control, and / or data signals from the HBM memory controller circuit 133, according to the pseudo-channel corresponding to the read / write operation. Additionally, based on the address, control, and data signals from the HBM memory controller circuit 133, the bus switch circuit 135 sends an enable signal to the appropriate TSV selection circuit 332 in the die.
[0049] For example, Figure 4A To illustrate a block diagram of a portion of the bus switch circuit 135, the bus switch circuit selects and communicatively couples to the TSV bus for channel 0 and transmits an enable signal. For simplicity and clarity, Figure 4A Only pseudo-channels PC0 and PC1 for channel 0 are shown. However, those skilled in the art will understand that the appropriate TSV bus selection for other channels will be similar. In some embodiments, each path selection switch 402 may correspond to a pseudo-channel bus and may include multiple bit switches corresponding to individual DQ pins (see [link to documentation]). Figure 4B ).like Figure 4A As seen, path selection switch 402a communicatively couples DQ pins 0 to 31 of PC0 in channel 0 to either the TSV0 or TSV1 bus for PC0. Similarly, path selection switch 402b communicatively couples DQ pins 32 to 64 of PC1 in channel 0 to either the TSV0 or TSV1 bus for PC1.
[0050] In some embodiments, based on address, control, and / or data signals from HBM memory controller circuitry 133, path selection sequence circuitry 404 selects the appropriate TSV bus and transmits an enable signal to the appropriate path selection switch 402 and to the appropriate TSV selection circuitry 320 or 322 in the appropriate die. Path selection sequence circuitry 404 and / or another circuit may include one or more processors, memories, lookup tables, and / or other circuitry to determine the appropriate TSV bus, channel, pseudo-channel, stack, die, and / or TSV selection circuitry for selection based on address, control, and / or data information from HBM memory controller circuitry 133. As further discussed below, the enable signal may include TSV0 / RD select signals, TSV1 / RD select signals, TSV0 / WR select signals, and TSV1 / WR select signals. However, other embodiments may include more or fewer signals depending on the configuration of the HBM device. Based on the enable signal of path selection switch 402, a data path is selected between the DQ bus and the TSV0 bus, and the DQ bus and the TSV0 bus are communicatively coupled; or a data path is selected between the DQ bus and the TSV1 bus, and the DQ bus and the TSV1 bus are communicatively coupled; or no data path is selected.
[0051] Figure 4B Examples of individual bit switches 410 that may be included in a path selection switch 402 are shown. Each path selection switch 402 may contain multiple bit switches 410, where each bit switch 410 corresponds to a bit in an appropriate pseudo-channel. Figure 4B As seen, bit switch 410 may include one or more tri-state inverter circuits (or another suitable switching circuit) to communicatively couple the DQ pin to one or more suitable TSVs to provide a bidirectional data path. Bit switch 410 may receive an enable signal from path selection sequence circuit 404 and select an appropriate path between the appropriate TSV (TSV0 or TSV1) and the DQ pin. For example, if the TSV0 / RD select signal or the TSV0 / WR select signal is enabled, then a data path is selected between the DQ pin and the TSV on the TSV0 bus. If the TSV1 / RD select signal or the TSV1 / WR select signal is enabled, then a data path is selected between the DQ pin and the TSV on the TSV1 bus. If no signal is enabled, no data path is selected (e.g., no data is sent / received to / from the pseudo-channel). In some embodiments, instead of two signals (e.g., the TSV0 / RD select signal and the TSV0 / WR signal), only one signal for the TSV0 bus can be used. Similarly, only one signal for the TSV1 bus can be used.
[0052] In operation, when the HBM memory controller circuitry 133 sends data to be written to the memory bank via a pseudo-channel, for example based on a command from the host device, the path selection switch 402 for the pseudo-channel selects either the TSV0 bus or the TSV1 bus based on an enable signal and communicatively couples the DQ bus to the appropriate TSV bus. Similarly, when data is received from the memory bank based on, for example, a command from the host device, the path selection switch 402 selects the appropriate TSV bus (e.g., TSV0 or TSV1) based on an enable signal and communicatively couples it to the DQ bus. In some embodiments, the enable signal may, for example, be hardwired to each path selection switch 402. In other embodiments, the enable signal includes switch identification information and is communicated via a bus to some or all of the path selection switches 402.
[0053] In some embodiments, the path selection sequence circuit 404 transmits the enable signal in an alternating ("ping-pong") pattern between the selected TSV0 bus and the selected TSV0 bus. The alternating sequence may be based on t CCDS The CLK cycle period can be, for example, two CLK cycles. For instance, the bus switch circuit 135 can alternatively be used, for example, in t... CCDL During the CLK cycle, every t CCDS The CLK cycle period is selected between the first TSV bus and the second TSV bus. In other embodiments, the alternation sequence may be based on consecutive commands (e.g., read or write commands). For example, a bus switching circuit may alternatively select between the first TSV bus and the second TSV bus during consecutive read or write commands. In other embodiments, the TSV bus selection may be determined based on other criteria, such as whether the TSV bus (e.g., the default TSV bus) is busy before selecting another bus.
[0054] As discussed above, bus switch circuit 135 (and / or another circuit) transmits enable signals (e.g., TSV0 / RD select, TSV0 / WR select, TSV1 / RD select, and TSV1 / WR select) to the appropriate TSV select circuit 332 in the appropriate die. Based on the enable signal, each TSV select circuit 332 can direct data to the appropriate pseudo-channel. For example, Figure 4C Explain the corresponding pseudo-channel PC0 used for channel 0. Figure 3 The TSV selection circuit 332 and TSV selection circuit 450 are described in the text. TSV selection circuits used for other pseudo-channels are similar. For example... Figure 4C As seen, the TSV selection circuit 450 may include a set of driver circuits 460a, 460b to, where applicable, drive data from 320 or 322 on the PC0 bus in die 310a (see [link]). Figure 3The data is transmitted to the TSV0 PC0 bus or the TSV1 PC0 bus. The TSV selection circuit 450 may also include a set of input buffer circuits 465a, 465b to receive data from the TSV0 PC0 bus or the TSV1 PC0 bus, as applicable, and to transmit the data to the PC0 bus corresponding to groups 320 and 322 in die 310a.
[0055] To enable the drivers 460a, 460b and input buffers 465a, 465b as described above, the TSV selection circuit 450 may receive four enable signals from, for example, the bus switch circuit 135: a TSV0 / RD select signal, a TSV0 / WR select signal, a TSV1 / RD select signal, and a TSV1 / RD select signal. The TSV0 / RD select signal, when enabled, activates driver 460a to communicatively couple the PC0 bus to the TSV0 PC0 bus during a read operation. Similarly, the TSV1 / RD select signal, when enabled, activates driver 460b to communicatively couple the PC0 bus to the TSV1 PC0 bus during a read operation. The TSV0 / WR select signal, when enabled, activates input buffer 465a to communicatively couple the PC0 bus to the TSV0 PC0 bus during a write operation. Similarly, the TSV1 / WR select signal, when enabled, activates input buffer 465b to communicatively couple the PC0 bus to the TSV1 PC0 bus during a write operation.
[0056] Figure 5A A simplified timing diagram 500 illustrating a write operation consistent with this disclosure is provided. The timing diagram may correspond to an HBM device with a data rate of 16 Gbps. As shown in the figure, t... CCDS Write commands, separated by CLK cycles (two CKL cycles), alternate between the TSV buses (TSV0 and TSV1). For example, write commands W1 and W3 correspond to TSV0, and write commands W2 and W4 correspond to TSV1. Figure 5A As seen, the data used for each command can be found in the corresponding timing ratio t. CCDL / t CCDS The memory access frequency corresponds to the TSV bus for each CLK cycle, and the timing ratio can be, for example, 4 CLK cycles. As discussed above, with a TSV bus timing of 4 CLK cycles (1ns), there is no need to change the memory array timing. Furthermore, in t... CCDL Within a CLK cycle (e.g., 8 CLK cycles), there are four consecutive write commands that open four group groups. As discussed above, compared to related technology HBM devices, in 8 CLK cycles (2ns), t CCDLIn the case of a CLK cycle, the amount of data that the memory array can cyclically pass through groups and transfer within the same time period is doubled. For clarity, in Figure 5A In this context, different hash lines and crosshairs are used to identify different W# data streams.
[0057] The time from T0 to T4 corresponds to t CCDL The CLK cycle period, in this embodiment, is 8 CLK cycles. Figure 5A As seen, during the duration t CCDL During a CLK cycle, four batches (BGs) can be opened (e.g., per channel or per pseudo-channel) for write operations, allowing more bandwidth than related technology devices that only open two BGs. In the following embodiment, the group is written to correspond to PC0 of channel 0, and therefore, the TSV bus corresponds to PC0 of channel 0.
[0058] At time T0, based on a write command W1 to group 2 of BG0 in SID0 with BL of 8, 32 bytes of data are transferred from, for example, host device 120 to the DQ bus via HBM memory controller circuit 133 using 2 CLK cycles (4 WCK cycles). The 32 bytes of W1 may correspond to pseudo-channel PC0 (e.g., based on PC bit information in the address signal). At time T1, based on information from, for example, host device, HBM memory controller 133, and / or bus switch circuit 135, the TSV0 / WR select signal from path selection sequence circuit 404 goes high (and the TSV1 / WR select signal goes low) to select the TSV0 bus corresponding to BG0 in SID0, and the W1 data is transferred to group 2 via the TSV0 bus. Figure 5A As seen in the diagram, once the transmission begins, group 2 can be accessed at t. CCDL / t CCDS The memory accesses the corresponding TSV0 bus for each CLK cycle (which in this case is 8 / 2 = 4 CLK cycles). In this embodiment, 4 CLK cycles correspond to 1 ns. Therefore, the timing of the memory array in Group 2 can remain the same as that of related art HBM devices at a data rate of 8 Gbps.
[0059] Still at time T1, based on the write command W2 for group 3 of BG0 in SID1, 32 bytes of data are transferred to the DQ bus immediately after the data transfer to the DQ bus for write command W1 is completed. The 32 bytes of W2 can correspond to pseudo-channel PC0. At time T2, while the W1 data is still being transferred via the TSV0 bus for BG0 in SID0, the TSV1 / WR select signal goes high (and the TSV0 / WR signal goes low) to select the TSV1 bus corresponding to BG0 in SID1, and the W2 data is transferred to group 3 via the TSV1 bus. Similar to the W1 write operation, once the transfer begins, group 3 can be transferred at time T2. CCDL / t CCDS The CLK cycle (4 CLK cycles) takes the corresponding TSV1 bus from memory.
[0060] Still at time T2, based on the write command W3 to group 1 of BG1 in SID2, 32 bytes of data are transferred to the DQ bus immediately after the data transfer to the DQ bus for write command W2 is completed. The 32 bytes of W3 can correspond to pseudo-channel PC0. At time T3, group 2 of BG0 in SID0 has been transferred and the TSV0 bus has been released. Still at time T3, while W2 data is still being transferred via the TSV1 bus for BG0 in SID1, the TSV0 / WR select signal goes high (and the TSV1 / WR signal goes low) to select the TSV0 bus for BG1 in SID2, and W3 data is transferred to group 1 via the TSV0 bus. Similar to other write operations, once the transfer begins, group 1 can be accessed at time T2. CCDL / t CCDS The CLK cycle (4 CLK cycles) takes the corresponding TSV1 bus from memory.
[0061] Still at time T3, based on the write command W4 for group 2 of BG1 in SID3, 32 bytes of data are transferred to the DQ bus immediately after the data transfer to the DQ bus for write command W3 is complete. The 32 bytes of W4 can correspond to pseudo-channel PC0. At time T4, group 3 of BG0 in SID1 has been transferred and the TSV1 bus has been released. Still at time T4, while the W3 data is still being transferred via the TSV0 bus for BG1 in SID2, the TSV1 / WR select signal goes high (and the TSV0 / WR signal goes low) to select the TSV1 bus for BG1 in SID3, and the W4 data is transferred to group 2 via the TSV1 bus. Similar to other write operations, once the transfer begins, group 2 can be transferred at time T4. CCDL / t CCDSThe CLK cycle (4 CLK cycles) retrieves the corresponding TSV1 bus from memory. At time T5, the transfer of W3 data to group 1 of BG1 in SID2 is completed, and the TSV0 bus is released. At time T6, the transfer of W3 data to group 2 of BG1 in SID3 is completed, and the TSV1 bus is released.
[0062] Figure 5B A simplified timing diagram 550 illustrating a read operation consistent with embodiments of this disclosure is provided. The timing diagram may correspond to an HBM device with a data rate of 16 Gbps. As shown in the figure, t CCDS Read commands, separated by CLK cycles (two CKL cycles), alternate between the TSV buses (TSV0 and TSV1). For example, read commands R1 and R3 correspond to TSV0, and read commands R2 and R4 correspond to TSV1. Figure 5B As seen, the data used for each command can be found in the corresponding timing ratio t. CCDL / t CCDS The memory access frequency corresponds to the TSV bus for each CLK cycle, and the timing ratio can be, for example, 4 CLK cycles. As discussed above, with a TSV bus timing of 4 CLK cycles (1ns), there is no need to change the memory array timing. Furthermore, in t... CCDL Within a CLK cycle (e.g., 8 CLK cycles), there are four consecutive read commands that open four groups. As discussed above, compared to related art HBM devices, in 8 CLK cycles (2ns), t CCDL In the case of a CLK cycle, the amount of data that the memory array can cyclically pass through groups and transfer within the same time period is doubled. For clarity, in Figure 5B In this context, different hash lines and crosshairs are used to identify different R# data streams.
[0063] The time from T0 to T4 corresponds to t CCDL The CLK cycle period, in this embodiment, is 8 CLK cycles. Figure 5B As seen in t CCDL During the CLK cycle, four BGs (e.g., per channel or per pseudo channel) can be opened for read operations, which allows for more bandwidth than related technology devices that only open two BGs.
[0064] At time T0, based on information from, for example, the host device, HBM memory controller 133, and / or bus switch circuit 135, the TSV0 / RD select signal from the path selection sequence circuit 404 goes high (and the TSV1 / RD select signal goes low) to select the TSV0 bus corresponding to PC0. Still at T0, based on read command R1, 32 bytes of data (BL is 8) are read from group 2 of BG0 in SID0 corresponding to PC0 (e.g., based on the PC bit information in the address signal) for transmission via the TSV0 bus. Figure 5B As seen in the diagram, once the transmission begins, group 2 can be accessed at t. CCDL / t CCDS The memory uses the corresponding TSV0 bus for each CLK cycle (which in this case is 8 / 2 = 4 CLK cycles). Therefore, the timing of the memory array in Group 2 can remain the same as that of related technology HBM devices at a data rate of 8 Gbps.
[0065] At time T1, while data from group 2 is still being transmitted to the TSV0 bus, the TSV1 / RD select signal goes high (and the TSV0 / RD select signal goes low) to select the TSV1 bus corresponding to PC0. Still at T1, based on read command R2, 32 bytes of data are read from group 3 of BG3 in SID1 corresponding to PC0 for transmission via the TSV1 bus. Figure 5B As seen in the diagram, once the transmission begins, group 3 can be accessed at t. CCDL / t CCDS The CLK cycle (4 CLK cycles) takes the corresponding TSV1 bus from memory.
[0066] At time T2, the R1 read transfer via the TSV0 bus from group 2 of BG0 in SID0 is completed, and the TSV0 bus is released. Additionally, the R1 read data is then... CCDS The CLK cycle (two CLK cycles) is internally available on the DQ bus for transmission to, for example, host device 120 via HBM memory controller circuitry 133. Still at T2, while data from group 3 is still being transmitted via the TSV1 bus, the TSV0 / RD select signal goes high (and the TSV1 / RD select signal goes low) to select the TSV0 bus corresponding to PC0. Based on read command R3, 32 bytes of data are read from group 1 (BG1 in SID2 corresponding to PC0) for transmission via the TSV0 bus. Once transmission begins, group 1 can be accessed at t CCDL / t CCDS The CLK cycle (4 CLK cycles) retrieves the corresponding TSV0 bus from memory.
[0067] At time T3, the R2 read transfer via the TSV1 bus from group 3 of BG0 in SID1 is completed, and the TSV1 bus is released. Additionally, the R2 read data is then processed at time t. CCDS The CLK cycle (two CLK cycles) is internally available on the DQ bus for transmission to, for example, host device 120 via HBM memory controller circuitry 133. Still at T3, while data from group 1 is still being transmitted via the TSV0 bus, the TSV1 / RD select signal goes high (and the TSV0 / RD select signal goes low) to select the TSV1 bus corresponding to PC0. Based on read command R4, 32 bytes of data are read from group 2, corresponding to BG1 in SID3 of PC0, for transmission via the TSV1 bus. Once transmission begins, group 2 can be accessed at t CCDL / t CCDS The CLK cycle (4 CLK cycles) takes the corresponding TSV1 bus from memory.
[0068] At time T4, the R3 read transfer via the TSV0 bus from group 1 of BG1 in SID2 is completed, and the TSV0 bus is released. Additionally, the R3 read data is then processed at time t. CCDS The CLK cycle (two CLK cycles) is internally available on the DQ bus for transfer to, for example, the host device 120 via the HBM memory controller circuit 133. At time T5, the R4 read transfer via the TSV1 bus from group 2 of BG1 in SID3 is completed, and the TSV1 bus is released. This allows the R4 read data to be transferred at time T5. CCDS The CLK cycle (2 CLK cycles) is available on the DQ bus to be transmitted to, for example, host device 120 via HBM memory controller circuit 133.
[0069] like Figure 5A and 5B As seen in the diagram, because there is more than one TSV bus, therefore in t CCDL Groups opened during a CLK cycle can be accessed in an interleaved, overlapping mode. Therefore, in the exemplary embodiments of this disclosure, bandwidth can be increased while keeping the DQ bus saturated during read / write operations, and simultaneously at a t equal to two CLK cycles. CCDS CLK cycle operation. Additionally, as... Figure 5B As observed, the command latency for read operations between different stack SIDs remains within 2 CLK cycles. CCDR The CLK cycle further ensures that the DQ bus remains saturated.
[0070] Figure 6The illustration shows a flowchart 600 illustrating method steps performed by one or more processors and / or hardwired circuitry systems in an HBM device (e.g., HBM memory controller circuitry 133, bus switch circuitry 135, and / or TSV selection circuitry TSV SEL). In step 610, the HBM device selects a TSV bus from a plurality of TSV buses, each TSV bus having a set of TSVs. For example, as discussed above, based on information from HBM memory controller circuitry 133, bus switch circuitry 135 selects the TSV0 bus or TSV1 bus of the appropriate pseudo-channel (e.g., PC0 or PC1) to perform a read / write operation.
[0071] In step 620, the HBM device communicatively couples a DQ bus with a set of DQ pins to a selected TSV bus. For example, as discussed above, in addition to selecting an appropriate TSV bus, the HBM memory controller circuit 133 communicatively couples a DQ bus with an appropriate channel to the selected TSV bus.
[0072] In step 630, the HBM device communicatively couples groups of dies to a selected TSV bus, wherein the groups comprise one or more groups having memory arrays. For example, as discussed above, based on an enable signal from path selection sequence circuit 404, corresponding TSV selection circuits 320 or 322 for appropriate DRAM dies and pseudo-channels communicatively couple groups of dies containing memory arrays to a selected TSV bus (TSV0 or TSV1).
[0073] As described above, embodiments of this disclosure provide increased bandwidth compared to related HBM devices while ensuring synchronization of DRAM memory array timing, TSV bus timing, and DQ bus timing. For example, in some embodiments, the data rate at the DQ pin is increased while maintaining the same memory array as related HBM devices. Furthermore, by relaxing the frequency cycling timing in the TSV bus, embodiments of this disclosure can perform low-voltage switching in the TSV to maintain low power consumption. Moreover, compared to related HBM devices, embodiments of this disclosure increase the bandwidth available for DRAM memory array timing, TSV bus timing, and DQ bus timing. CCDL The number of groups opened during the CLK cycle, while maintaining the 4N architecture and the same number of groups.
[0074] Additionally, it should be understood that specific embodiments of the technology have been described herein for illustrative purposes, but well-known structures and functions have not been shown or described in detail to avoid unnecessarily obscuring the description of embodiments of the technology. In the event of any conflict between any material incorporated herein by reference and this disclosure, this disclosure shall prevail. Where the context permits, singular or plural terms may also include plural or singular terms, respectively. Furthermore, unless the word “or” is expressly limited to meaning only a single item exclusive to another item in a list referring to two or more items, the use of “or” in such a list shall be understood to include (a) any single item in the list, (b) all items in the list, or (c) any combination of items in the list. Furthermore, as used herein, the phrase “and / or” in “A and / or B” means only A, only B, and both A and B. Furthermore, the terms “comprising,” “including,” “having,” and “with” are used throughout to mean at least one or more of the described features, such that no larger number of identical features and / or other features of additional types are excluded. Furthermore, the terms “roughly,” “approximately,” and “about” are used here to mean at least 10% of a given value or limit. By way of example only, an approximate ratio means within ten percent of a given ratio.
[0075] The foregoing figures illustrate several embodiments of the disclosed technology. A computing device on which the described technology can be implemented may include one or more central processing units, memory, input devices (e.g., keyboard and pointing devices), output devices (e.g., display devices), storage devices (e.g., disk drives), and network devices (e.g., network interfaces). The memory and storage devices are computer-readable storage media capable of storing at least a portion of the instructions for implementing the described technology. Additionally, data structures and message structures may be stored or transmitted via data transmission media (e.g., signals on a communication link). Therefore, computer-readable media may include computer-readable storage media (e.g., “non-transitory” media) and computer-readable transmission media.
[0076] It should also be understood that various modifications can be made without departing from this disclosure or the present technology. For example, the dies in an HBM device can be arranged in any other suitable order (e.g., one or more non-volatile memory dies are positioned between interface dies and volatile memory dies; volatile memory dies are at the bottom of a die stack; etc.). Furthermore, those skilled in the art will understand that the various components of the present technology can be further divided into sub-components, or that the various components and functions of the present technology can be combined and integrated. Additionally, certain aspects of the present technology described in the context of a particular embodiment can be combined or eliminated in other embodiments. For example, although it is discussed herein as using non-volatile memory dies (e.g., NAND dies and / or NOR dies) to extend the memory of an HBM device, it should be understood that alternative memory can be used to extend the dies (e.g., larger capacity DRAM dies and / or any other suitable memory components). While such embodiments may forgo certain benefits (e.g., non-volatile memory), such embodiments can still provide additional benefits (e.g., reduced throughput through bottlenecks, allowing relatively fast execution of many complex computational operations, etc.).
[0077] Furthermore, while advantages associated with certain embodiments of the present technology have been described in the context of those embodiments, other embodiments may also exhibit such advantages, and not all embodiments necessarily exhibit such advantages to fall within the scope of the present technology. Therefore, this disclosure and related technologies may cover other embodiments not expressly shown or described herein.
Claims
1. A system-in-package (SiP) device, comprising: Substrate; The processing unit is supported by the substrate; and A high-bandwidth memory (HBM) device, which is carried on the substrate and electrically coupled to the processing unit, wherein the HBM device includes: An interface die includes a bus switching circuit configured to select a TSV bus from a plurality of through-silicon via (TSV) buses, each TSV bus having a set of TSVs, and communicatively coupling a DQ bus with a set of DQ pins to the selected TSV bus. and One or more stacks, each stack carrying one or more dies, wherein each die includes a TSV bus selection circuit configured to communicatively couple a group of dies to a TSV bus selected by the bus switching circuit of the interface die, the group comprising one or more groups having memory arrays.
2. The SiP device according to claim 1, wherein the DQ bus corresponds to a pseudo channel or channel of the HBM device.
3. The SiP device of claim 1, wherein the bus switching circuit selects between a first TSV bus and a second TSV bus during a continuous read command or a continuous write command.
4. The SiP device according to claim 1, wherein the bus switching circuit is in t CCDL Every t during the clock CLK cycle CCDS The CLK cycle period is selected between the first TSV bus and the second TSV bus. Where t CCDL This corresponds to the delay between commands associated with different groups within the same group, and Where t CCDS This corresponds to the delay between commands associated with different groups within different groupings.
5. The SiP device according to claim 1, wherein the data rate at the DQ bus is greater than 8 gigabits per second (Gbps).
6. The SiP device according to claim 1, in, During a read or write operation on a group within the group, each TSV selection circuit can access the TSV bus selected by the bus switching circuit of the interface die within X CLK cycle periods, where X is the number of t. CCDL / t CCDS The ratio, Where t CCDL This corresponds to the delay between commands associated with different groups within the same group, and Where t CCDS This corresponds to the delay between different commands associated with groups in different groups.
7. The SiP device according to claim 6, wherein at t CCDL During the CLK cycle, more than two groups are accessed in an interleaved overlapping mode.
8. A high-bandwidth memory (HBM) device, comprising: An interface die, operatively coupled to a host device, includes bus switching circuitry configured to select a TSV bus from a plurality of through-silicon via (TSV) buses, each TSV bus having a set of TSVs, and communicatively coupling a DQ bus with a set of DQ pins to the selected TSV bus. and One or more stacks, each stack carrying one or more dies, wherein each die includes a TSV bus selection circuit configured to communicatively couple a group of dies to a TSV bus selected by the bus switching circuit of the interface die, the group comprising one or more groups having memory arrays.
9. The HBM device according to claim 8, wherein the DQ bus corresponds to a pseudo channel or channel of the HBM device.
10. The HBM device of claim 8, wherein the bus switching circuit selects between a first TSV bus and a second TSV bus during a continuous read command or a continuous write command.
11. The HBM device according to claim 8, wherein the bus switching circuit is at t CCDL Every t during the clock CLK cycle CCDS The CLK cycle period is selected between the first TSV bus and the second TSV bus. Where t CCDL This corresponds to the delay between commands associated with different groups within the same group, and Where t CCDS This corresponds to the delay between commands associated with different groups within different groupings.
12. The HBM device of claim 8, wherein the data rate at the DQ bus is greater than 8 gigabits per second (Gbps).
13. The HBM device according to claim 8, in, During a read or write operation on a group within the group, the corresponding TSV selection circuit can access the TSV bus selected by the bus switching circuit of the interface die during a read or write operation within X CLK cycle periods, where X is the number of t. CCDL / t CCDS The ratio, Where t CCDL This corresponds to the delay between commands associated with different groups within the same group, and Where t CCDS This corresponds to the delay between commands associated with different groups within different groupings.
14. The HBM device according to claim 13, wherein at t CCDL During the CLK cycle, more than two groups are accessed in an interleaved overlapping mode.
15. A method comprising: Select a TSV bus from multiple through-silicon via (TSV) buses, each TSV bus having a set of TSVs. A DQ bus with a set of DQ pins is communicatively coupled to a selected TSV bus; and Groups of bare dies are communicatively coupled to the selected TSV bus, the groups comprising one or more groups having memory arrays. The DQ bus here corresponds to a pseudo channel or channel of a high-bandwidth memory (HBM) device.
16. The method of claim 15, wherein the selection further comprises selecting between a first TSV bus and a second TSV bus during a continuous read command or a continuous write command.
17. The method of claim 15, further comprising: In t CCDL Every t during the clock CLK cycle CCDS The CLK cycle period is selected between the first TSV bus and the second TSV bus. Where t CCDL This corresponds to the delay between commands associated with different groups within the same group, and Where t CCDS This corresponds to the delay between commands associated with different groups within different groupings.
18. The method of claim 15, further comprising: The HBM device is operated such that the data rate at the DQ bus is greater than 8 gigabits per second (Gbps).
19. The method according to claim 15, Access to the selected TSV bus is controlled during a read or write operation within X CLK cycle periods, where X is the number of CLK cycles. CCDL / t CCDS The ratio, Where t CCDL This corresponds to the delay between commands associated with different groups within the same group, where t CCDS This corresponds to the delay between commands associated with different groups within different groupings.
20. The method of claim 19, further comprising: In t CCDL During the duration, more than two groups are accessed in an interleaved overlapping pattern.