Vertically stacked high bandwidth storage devices and associated systems and methods
By integrating high-bandwidth memory and high-bandwidth storage devices in system-level packaging, the bottleneck problem between data transmission and storage capacity is solved, efficient data processing and storage is achieved, and the processing speed and reliability of the system are improved.
Patent Information
- Application Number
- CN202510054109.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-16
- Filing Date
- 2025-01-14
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, high bandwidth storage devices have bottlenecks between data transmission and storage capacity, especially when processing large amounts of data, the contradiction between high bandwidth communication of volatile memory and low bandwidth communication of nonvolatile memory leads to limited processing speed and efficiency.
System-level packaging (SiP) technology is adopted to integrate high-bandwidth memory (HBM) and high-bandwidth storage (HBS) devices on the same substrate, and electrically couple it through the SiP bus to realize high-bandwidth communication of volatile and nonvolatile memory. The HBS device provides non-volatile storage and communicates with the HBM device and processing unit through a high-bandwidth path, reducing data transmission bottlenecks.
It improves data transmission speed and storage capacity, reduces processing time, reduces power consumption, and quickly saves and restores data state when powered down or powered on, enhancing the processing speed and reliability of the system.
Smart Images

Figure CN120343926A_ABST
Abstract
Description
Technical Field
[0001] The present technology generally relates to vertically stacked semiconductor devices, and more particularly, to vertically stacked high bandwidth memory devices for semiconductor packaging. Background Art
[0002] Microelectronic devices (such as memory devices, microprocessors, and other electronic devices) typically include one or more semiconductor dies mounted to a substrate and enclosed in a protective overcoat. The semiconductor die includes functional features such as memory cells, processor circuitry, imager devices, interconnect circuitry, and the like. To meet the continuing demand for ever-decreasing size, wafers, individual semiconductor dies, and / or active components are typically manufactured in batches, singulated, and then stacked on a support substrate (e.g., a printed circuit board (PCB) or other suitable substrate). The stacked dies can then be coupled to the support substrate (sometimes also referred to as a package substrate) by bond wires in a shingled stacked die (e.g., dies stacked with a certain offset for each die) and / or by through-substrate vias (TSVs) between the die and the support substrate. Summary of the Invention
[0003] In one aspect, the present disclosure relates to a system-in-package (SiP) device including: a base substrate; a processing unit carried by the base substrate; a high bandwidth memory (HBM) device carried by the base substrate and electrically coupled to the processing unit via a SiP bus, wherein the HBM device includes: a first interface die; one or more volatile memory dies carried by the first interface die; and an HBM bus electrically coupled to the first interface die and each of the one or more volatile memory dies; and a high bandwidth storage (HBS) device carried by the base substrate and electrically coupled to the HBM device via the SiP bus, wherein the HBS device includes: a second interface die; one or more non-volatile memory dies carried by the second interface die; and an HBS bus electrically coupled to the second interface die and each of the one or more non-volatile memory dies.
[0004] In another aspect, the present disclosure relates to a method including: generating a request for a subset of a dataset stored in a plurality of non-volatile memory dies in a high bandwidth storage (HBS) device; writing a copy of the subset to a plurality of volatile memory dies in a high bandwidth memory (HBM) device; reading the subset from the plurality of volatile memory dies into a processing unit; processing the subset at the processing unit; and writing a result of processing the subset to the plurality of volatile memory dies.
[0005] In yet another aspect, the present disclosure relates to a method that includes: updating a current state of a high-bandwidth memory (HBM) device during a processing operation at a processing unit communicatively coupled to the HBM device; receiving a power-down or idle request; and in response to the power-down or idle request, controlling the HBM device to write the current state of the HBM device from the HBM device to a high-bandwidth storage (HBS) device via a system-in-package (SiP) bus, where the HBS device includes one or more non-volatile memory dies. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 is a schematic diagram illustrating an environment incorporating a high-bandwidth memory architecture.
[0007] Figure 2 is a schematic diagram illustrating an environment incorporating a high-bandwidth memory architecture according to some embodiments of the present technology.
[0008] Figure 3 is a partial cross-sectional schematic diagram of a system-in-package configured according to some embodiments of the present technology.
[0009] Figure 4 is a partial exploded schematic diagram of a high-bandwidth storage device configured according to some embodiments of the present technology.
[0010] Figure 5 is a flowchart of a process for an operating system-level packaging device according to some embodiments of the present technology.
[0011] Figure 6 is a flowchart of a process for operating a high-bandwidth storage device according to some embodiments of the present technology.
[0012] Figure 7 is a partial cross-sectional schematic diagram of a high-bandwidth storage device configured according to further embodiments of the present technology.
[0013] Figure 8A AND 8B is a flowchart of a process for powering down and powering up a system-in-package device using a high-bandwidth storage device according to some embodiments of the present technology.
[0014] The figures are not necessarily drawn to scale. Additionally, it will be understood that several of the figures have been drawn schematically and / or partially schematically. Similarly, for purposes of discussing some embodiments of the present technology, some components / or operations may be separated into different blocks or combined into a single block. Further, although the present technology may have various modifications and alternative forms, specific embodiments have been shown by way of example in the figures and are described in detail below. However, it is not intended to limit the technology to the specific embodiments described. DETAILED DESCRIPTION
[0015] High data reliability, high-speed memory access, low power consumption, and reduced chip size are characteristics required of semiconductor memories. In recent years, three-dimensional (3D) memory devices have been introduced. Some 3D memory devices are formed by vertically stacking memory dies and interconnecting the dies using through-silicon (or through-substrate) vias (TSVs). Benefits of 3D memory devices include: shorter interconnects (which reduce signal delay and power consumption); a larger number of vertical vias between layers (which allow wide-bandwidth buses between functional blocks (such as memory dies) in different layers); and a relatively small footprint. Thus, 3D memory devices contribute to higher memory access speeds, lower power consumption, and reduced chip size. Example 3D memory devices include Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM). For example, HBM is a type of vertically stacked memory that includes dynamic random access memory (DRAM) dies and interface dies (which, for example, provide an interface between the DRAM dies of the HBM device and a host device).
[0016] In a system-in-package (SiP) configuration, an HBM device can be integrated with a host device (such as a graphics processing unit (GPU) and / or a central processing unit (CPU)) using a base substrate (such as a silicon interposer, a substrate of organic material, a substrate of inorganic material, and / or any other suitable material that provides interconnect between the GPU / CPU and the HBM device and / or provides mechanical support for components of the SiP device), and the HBM device communicates with the host through the base substrate. Because the traffic between the HBM device and the host device resides within the SiP (such as using signals routed through a silicon interposer), a higher bandwidth can be achieved between the HBM device and the host device than in a conventional system. In other words, the TSVs that interconnect the DRAM dies within the HBM device and the silicon interposer that integrates the HBM device and the host device enable routing of a greater number of signals (such as a wider data bus) than the number of signals typically found between a packaged memory device and a host device (such as through a printed circuit board (PCB)). The high-bandwidth interface within the SiP enables large amounts of data to move quickly between the host device (such as a GPU / CPU) and the HBM device during operation. For example, the high-bandwidth channel can be approximately 1000 gigabytes per second (GB / s, sometimes also referred to as gigabits (Gb)). It should be appreciated that this high-bandwidth data transfer between the GPU / CPU and the memory of the HBM device can be beneficial in various high-performance computing applications, such as video rendering, high-resolution graphics applications, artificial intelligence and / or machine learning (AI / ML) computing systems, and other complex computing systems and / or various other computing applications.
[0017] Figure 1is a schematic diagram illustrating an environment 100 incorporating a high - bandwidth memory architecture. As Figure 1 described, the environment 100 includes a SiP device 110 that has integrated with a silicon interposer 112 (or any other suitable substrate) one or more processing devices 120 (one is described in Figure 1 , and sometimes also referred to herein as one or more “hosts”) and one or more HBM devices 130 (one is described in Figure 1 ). The environment 100 further includes a storage device 140 coupled to the SiP device 110. The processing device 120 may include one or more CPUs and / or one or more GPUs, referred to as CPU / GPU 122, each of which may include registers 124 and a first - level cache 126. The first - level cache 126 (also referred to herein as the “L1 cache”) is communicatively coupled via a first communication path 152 to a second - level cache 128 (also referred to herein as the “L2 cache”). In the illustrated embodiment, the L2 cache 128 is incorporated into the processing device 120. However, it should be understood that the L2 cache 128 may be separately integrated into the SiP device 110 from the processing device 120. By way of example only, the processing device 120 may be carried by a substrate (e.g., an interposer that is itself carried by a package substrate) adjacent to the L2 cache 128 and communicate with the L2 cache 128 via one or more signal lines (or other suitable signal routing lines) therein. The L2 cache 128 may be shared by one or more processing devices 120 (and the CPU / GPU 122 therein). During operation of the SiP device 110, the CPU / GPU 122 may use the registers 124 and the L1 cache 126 to complete processing operations and attempt to retrieve data from the larger L2 cache 128 whenever a cache miss occurs in the L1 cache 126. Thus, the multi - level cache can help accelerate the average time taken for the processing device 120 to access data, thereby accelerating the overall processing rate.
[0018] As Figure 1 further described, the L2 cache 128 is communicatively coupled to the HBM device 130 via a second communication channel 154. As illustrated, the processing device 120 (and the L2 cache 128 therein) and the HBM device 130 are carried by the silicon interposer 112 and electrically coupled (e.g., integrated by the silicon interposer 112). The second communication channel 154 is provided by the silicon interposer 112 (e.g., the silicon interposer includes and routes interface signals forming the second communication channel, such as via one or more redistribution layers (RDL)). As Figure 1Additionally, the L2 cache 128 is also communicatively coupled to the storage device 140 via a third communication channel 156. As illustrated, the storage device 140 is outside the SiP device 110 and utilizes signal routing components not included within the silicon interposer 112 (e.g., between the packaged SiP device 110 and the packaged storage device 140). For example, the third communication channel 156 can be a peripheral bus for connecting components on a motherboard or a PCB, such as a Peripheral Component Interconnect Express (PCIe) bus. Thus, during operation of the SiP device 110, the processing device 120 can read data from and / or write data to the HBM device 130 and / or the storage device 140 via the L2 cache 128.
[0019] In the illustrated environment 100, the HBM device 130 includes one or more stacked volatile memory dies 132 (e.g., DRAM dies) coupled to the second communication channel 154, Figure 1 as schematically illustrated by one. As explained above, the HBM device 130 can be positioned on the silicon interposer 112, on which the processing device 120 is also positioned. Thus, the second communication channel 154 can provide a high-bandwidth (e.g., approximately 1000 GB / s) channel through the silicon interposer 112. Additionally, as explained above, each HBM device 130 can provide a high-bandwidth channel (not shown) between the volatile memory dies 132 therein. Thus, data can be communicated at high speed between the processing device 120 and the HBM device 130 (and the volatile memory dies 132 therein), which can be advantageous for data-intensive processing operations. Although the HBM devices 130 of the SiP device 110 provide relatively high-bandwidth communication, their integration on the silicon interposer 112 has certain drawbacks. For example, each HBM device 130 can provide a limited amount of storage (e.g., approximately 16 GB per device), and the total storage provided by all HBM devices 130 may not be sufficient to maintain the working data set for the operations to be performed by the SiP device 110. Additionally or alternatively, the HBM devices 130 consist of volatile memory (e.g., each requires power to maintain the stored data, and the data is lost once the HBM device loses power and / or experiences an unexpected power outage).
[0020] In contrast to the characteristics of the HBM device 130, the storage device 140 can provide a large amount of storage (e.g., on the order of terabytes and / or tens of terabytes). The larger capacity of the storage device 140 is generally sufficient to maintain the working data set for complex operations to be performed by the SiP device 110. Additionally, the storage device 140 is generally non-volatile (e.g., consists of NAND-based storage devices, such as NAND flash memory, as Figure 1As described in the (description), and thus the stored data is retained even after a power-off. However, as discussed above, the storage device 140 is positioned external to the SiP device 110 (e.g., not placed on the silicon interposer 112), but is coupled to the SiP device 110 via a communication channel (e.g., PCIe) routed on a motherboard, a system board, or other form of PCB. Thus, the third communication channel 156 may have a relatively low bandwidth (e.g., about 8 GB / s), significantly lower than the bandwidth of the second communication channel 154. Therefore, when data moves between the storage device 140 and the SiP device 110, the low bandwidth of the third communication channel 156 becomes a bottleneck for processing operations (e.g., graphics rendering, AI / ML processes, and the like) involving large amounts of data that are not suitable for the storage capacity of the HBM device 130. Additionally, the relatively low bandwidth of the third communication channel 156 becomes a bottleneck for power-off / power-on operations that require data to move between the storage device 140 and the SiP device 110.
[0021] Disclosed herein are high bandwidth storage (HBS) devices, as well as associated systems and methods, that address the disadvantages discussed above. The HBS device may include an interface die and one or more non-volatile memory dies (e.g., NAND dies, NOR dies, PCM dies, FeRAM dies, MRAM dies, and / or any other suitable die). The HBS device may also include one or more TSVs that are electrically coupled to the interface die and the one or more non-volatile memory dies to establish a communication path therebetween. As described herein, the TSVs may provide a wide communication path (e.g., about 1024 I / Os) between the interface die and the non-volatile memory dies, enabling high bandwidth therebetween. Additionally, the HBS device may be integrated into a SiP device that includes one or more of the HBS devices, one or more HBM devices, and one or more host devices (e.g., a processing device including a CPU and / or a GPU). In such embodiments, the HBS device, the HBM device, and the host device of the SiP device are placed on and / or integrated with a silicon interposer that includes high bandwidth communication channels between the HBS device, the HBM device, and / or the host device. Thus, the HBS devices disclosed herein may significantly expand the storage available within the SiP device, thereby reducing the frequency at which large data sets must be communicated through the bottlenecks described above.
[0022] Advantageously, a large data set (e.g., from an external storage component) may be loaded into the HBS device via a low bandwidth communication path (e.g., PCIe) during an initialization phase. Then, during processing, portions of the large data set may be transferred between the HBS device and the HBM device via a high bandwidth communication path (e.g., a SiP bus) based on the portions of the large data set (e.g., a working data set) that are being processed at one time. In this instance, the HBM device may provide a similar... as referenced aboveFigure 1 The functionality of the HBM device 130 being discussed. That is, for example, the HBM device can provide DRAM-based storage of a working data set, which can be accessed by a host device via a high-bandwidth interface. Once the first portion of the data set has been processed, the results can be saved to the HBS device, and the second portion of the data set can be loaded from the data set in the HBS device into the HBM device via a high-bandwidth communication path. Then, the process can be repeated for the first, second, etc. portions of the data set to use the data set in any number of computations at the host device without the need to load the data set via a low-bandwidth communication path. In a specific non-limiting example, the data set can include training data for an artificial intelligence and / or machine learning (AI / ML) model, which needs to be accessed and / or processed hundreds, thousands, tens of thousands, or more times to train the AI / ML model. In this example, the HBS device can significantly reduce the processing time by requiring the data set to be communicated only once via a low-bandwidth channel during an initialization phase, and then providing high-bandwidth transfer of the data set (or portions thereof) between the HBS device, the HBM device, and / or the host device during a processing phase (e.g., reducing the processing time by hundreds, thousands, tens of thousands, or more seconds).
[0023] Additionally or alternatively, the HBS device can provide non-volatile storage for the data stored in the HBM device (e.g., the HBS device operates as a non-volatile HBM device on a SiP). In this embodiment, the HBS device can save data from the HBM device and / or the host device and / or restore data to the HBM device and / or the host device in response to various events (such as power-down and / or power-up, power-off, errors during processing, and / or the like). For example, in response to a power-down or idle request, data from the HBM device and / or any cache can be stored in the HBS device to store the current state of the SiP device. Since the HBS device can be available via a high-bandwidth communication path, requests can be satisfied much faster (e.g., on the order of tens of milliseconds instead of seconds) than if the data were communicated to an external storage component. Similarly, in response to receiving a power-up or wake-up request, the data can be moved back to the HBM device and / or the cache via a high-bandwidth communication path. Thus, the saved state of the SiP can be restored in tens of milliseconds instead of the seconds required when data must be loaded from a separate storage component, and the power-up request can be answered.
[0024] Additional details regarding the HBS device, the SiP device having the HBS device, and associated systems and methods are set forth below. For ease of reference, semiconductor packages (and their components) are sometimes described herein with reference to front and back, top and bottom, up and down, upward and downward, and / or horizontal planes, x-y planes, vertical or z-directions of spatial orientation relative to the embodiments shown in the figures. However, it should be understood that the semiconductor assemblies (and their components) can be moved to different spatial orientations and used in different spatial orientations without changing the structure and / or function of the disclosed embodiments of the present technology. Additionally, signals within semiconductor packages (and their components) are sometimes described herein with reference to downstream and upstream, forward and backward, and / or read and write relative to the embodiments shown in the figures. However, it should be understood that signal flow can be described in various other terms without changing the structure and / or function of the disclosed embodiments of the present technology.
[0025] Furthermore, although the memory device architectures disclosed herein are primarily discussed in the context of expanding memory capacity to improve artificial intelligence and machine learning models and / or creating non-volatile memory in dynamic random access memory (DRAM) components, those skilled in the art will understand that the scope of the present technology is not limited thereto. For example, the systems and methods disclosed herein can also be deployed to expand the available high-bandwidth memory for various other applications (such as video rendering, decryption systems, and the like) that process large amounts of data.
[0026] Figure 2 is a schematic diagram illustrating an environment 200 incorporating an HBM architecture as well as an HBS architecture in accordance with some embodiments of the present technology. Similar to the environment 100 discussed above, the environment 200 includes one or more processing devices 220 ( Figure 2 one of which is illustrated in Figure 2 ), one or more HBM devices 230 ( Figure 2The SiP device 210 described in. In addition, the processing device 220 and the HBM device 230 are each integrated on an interposer 212 (e.g., a silicon interposer, another organic interposer, an inorganic interposer, and / or any other suitable substrate) that may include one or more signal routing lines. The processing device 220 is driven by a CPU / GPU 222 that includes a register 224 and an L1 cache 226. The L1 cache 226 is communicatively coupled to an L2 cache 228 via a first communication channel 252. The L2 cache 228 is communicatively coupled to a stack of one or more volatile memory dies 232 (e.g., DRAM dies) in the HBM device 230 via a second communication channel 254, and is coupled to a storage device 240 via a third communication channel 256. Further, the second communication channel 254 may have a relatively high bandwidth (e.g., about 1000 GB / s), while the third communication channel 256 may have a relatively low bandwidth (e.g., about 8 GB / s).
[0027] In Figure 2 the embodiment described in, the SiP device 210 further includes one or more HBS devices 260 ( Figure 2 described in one), each HBS device 260 including a stack of one or more storage dies 262 (e.g., NAND dies, NOR dies, or other suitable non-volatile memory dies). The storage dies 262 may provide a relatively large storage capacity (e.g., on the order of hundreds of gigabytes and / or terabytes), as well as non-volatile storage within the SiP device 210. In addition, as discussed in more detail below, the HBS device 260 (and the storage dies 262 therein) may be coupled to the HBM device 230 and / or the processing device 220 via a fourth communication channel 258. The fourth communication channel 258 may have a relatively high bandwidth (e.g., about 1000 GB / s) that is generally similar (or equivalent) to the second communication channel 254. Thus, the HBS device 260 provides the SiP device 210 with high-bandwidth access to a large amount of non-volatile storage without having to access the storage device 240 via the third communication channel 256. Although Figure 2 the embodiment is described in which the HBS device 260 is coupled to the HBM device 230 via the fourth communication channel 258, in an embodiment, the fourth communication channel 258 may alternatively or additionally couple the HBS device 260 to the processing device 220, and / or an additional communication channel (not shown) may couple the HBS device 260 to the processing device 220.
[0028] The combination of volatile memory (e.g., via the HBM device 230) and non-volatile memory (e.g., via the HBS device 260) within the SiP device 210 can provide various advantages. For example, volatile memory such as DRAM typically provides relatively faster access (e.g., read and write) than non-volatile memory such as NAND, but has a lower density (e.g., storage capacity within the die footprint). In contrast, non-volatile memory such as NAND typically provides high storage density, but may be accessed relatively slowly and may incur certain overheads (e.g., wear leveling). Thus, the volatile memory die 232 can provide low-latency fast communication, making data quickly available to the processing device 220 of the SiP device 210 as needed. The non-volatile memory die 262 can provide a relatively large memory capacity, and the non-volatile memory die 262 is "closer" to the processing device 220 (e.g., the SiP device 210 can access it via a high-bandwidth bus such as the fourth communication channel 258, the second communication channel 254, and / or other communication channels not shown) compared to the storage device 240 (e.g., which can be accessed via a slower channel such as PCIe). Additionally, the non-volatile memory die 262 can provide a non-volatile memory capacity that is closer to the processing device 220 and / or the volatile memory die 232 compared to the storage device 240 and / or other non-volatile memory capacities.
[0029] Thus, for example, a relatively large data set can be transferred from the storage device 240 to the non-volatile memory die 262 in the HBS device 260 to initiate a processing operation (e.g., to run an AI / ML algorithm). For example, the entire data set required for the AI / ML operation can be copied from the storage device 240 to the HBS device 260. Then, a subset of the data set can be quickly transferred from the HBS device 260 to the HBM device 230 via the high bandwidth of the fourth communication channel 258, and then quickly transferred to the processing device 220 via the high bandwidth of the second communication channel 254 (sometimes also referred to herein as the "high-bandwidth communication path"). When the processing device 220 finishes processing the subset, a new subset can be quickly written from the HBS device 260 into the HBM device 230 without retrieving data from the storage device 240 in a situation where there is an accompanying bottleneck in the third communication channel 256 (sometimes also referred to herein as the "low-bandwidth communication path"). Additionally, the processing operation can be iteratively performed (e.g., hundreds, thousands, tens of thousands, or more iterations typically used for AI / ML algorithms) without repeatedly communicating large data sets through the bottleneck. Thus, including the HBS device 260 can increase the processing speed of the SiP device 210, thereby increasing the functionality of the environment 200. Further, since communicating data through a high-bandwidth channel is more efficient than communicating data through a low-bandwidth channel, including the HBS device 260 can reduce the total power consumption of the environment 200 and / or reduce the heat generated by the environment 200.
[0030] Additionally or alternatively, the non-volatile memory die 262 in the HBS device 260 may save a copy of the data being processed and / or the overall state of the SiP device 210 in the non-volatile component. Thus, for example, there is no need to write the state of the HBM device 230 between the volatile memory die 232 and the storage device 240 to power down and / or power up. Instead, the state may be written to the non-volatile memory die 262 in the HBS device 260. Therefore, the power down operation (sometimes also referred to herein as "hibernation operation" and / or "idle operation") can be completed almost immediately (e.g., by saving the copy using the high bandwidth of the fourth communication channel 258). Similarly, the power up operation (sometimes also referred to herein as "wake up operation") can write the state back from the non-volatile memory die 262 in the HBS device 260 to the volatile memory die 232 in the HBM device 230 via the fourth communication channel 258, rather than writing back from the storage device 240 via the third communication channel 256. Therefore, the power down and / or power up operations can be accelerated from several seconds to well under one second (e.g., tens of milliseconds). Additionally or alternatively, the HBS device can prevent power outages and / or other processing errors in the environment 200. For example, since the HBM device 230 can save the current state of the SiP device 210 (e.g., the current state of the HBM device 230 and / or the processing device 220) to the HBS device 260 within a few milliseconds, the HBM device 230 can save the current state of the SiP device 210 to the HBS device 260 after a predetermined period (e.g., every ten seconds, every minute, every five minutes, every thirty minutes, every hour, every two hours, every twelve hours, every day, and / or any other suitable period) and / or after various processing milestones, without significantly delaying the processing at the SiP device 210. Thus, the power outage and / or other errors can be reverted to the last saved state before the power outage and / or error, resulting in less processing time and / or less data loss (e.g., resuming half of the processing operations without having to start over).
[0031] The environment 200 can be configured to perform any of a wide variety of suitable computing, processing, storage, sensing, imaging, and / or other functions. For example, representative examples of systems that include the environment 200 (and / or its components, such as the SiP device 210) include, but are not limited to, computers and / or other data processors such as desktop computers, laptop computers, Internet appliances, handheld devices (such as palmtop computers, wearable computers, cellular or mobile phones, automotive electronics, personal digital assistants, music players, etc.), tablet computers, multiprocessor systems, processor-based or programmable consumer electronics, network computers, and minicomputers. Additional representative examples of systems that include the environment 200 (and / or its components) include lights, cameras, vehicles, and the like. With respect to these and other examples, the environment 200 can be housed in a single unit or distributed, for example, across multiple interconnected units via a communication network, at various locations on a motherboard, and the like. Additionally, the components of the environment 200 (and / or any of its components) can be coupled to a variety of other local and / or remote memory storage devices, processing devices, computer-readable storage media, and the like. Additional details regarding the architecture of the environment 200, the SiP device 210, the HBM device 230, and their operating procedures are presented below with reference to Figures 3 to 8B present additional details regarding the architecture of the environment 200, the SiP device 210, the HBM device 230, and their operating procedures.
[0032] Figure 3 is a partial cross-sectional schematic view of a SiP device 300 configured in accordance with some embodiments of the present technology. As Figure 3 illustrated, the SiP device 300 includes a substrate 310 (e.g., a silicon interposer, another suitable organic substrate, an inorganic substrate, and / or any other suitable material) and a processing unit 320, an HBM device 330, and an HBS device 350 each integrated with the upper surface 312 of the substrate 310. For example, as discussed in more detail below, each of the processing unit 320, the HBM device 330, and the HBS device 350 is communicatively coupled via a SiP bus 340 formed in the substrate 310.
[0033] In the illustrated embodiment, the processing unit 320 is illustrated as a single component. However, as discussed above, the processing unit 320 can include CPU / GPU components, registers, L1 caches, L2 caches, and / or various other suitable components integrated into a single package.
[0034] The HBM device 330 includes a stack of semiconductor dies. The stack of semiconductor dies in the HBM device 330 includes an interface die 332, one or more volatile memory dies 334 ([[]] Figure 3 illustrated as six) carried by the interface die 332, and one or more through-substrate vias 338 (“TSV 338,” Figure 3FIG. schematically illustrates six). The TSVs 338 (which are sometimes also referred to herein as part of the "HBM bus" (or forming the "HBM bus")) extend from the interface die 332 through each of the volatile memory dies 334. The TSVs 338 allow each die to communicate data within the HBM device 330 (e.g., between the volatile memory dies 334 (e.g., DRAM dies) and the interface die 332) at a relatively high rate (e.g., about 100 GB / s, 1000 GB / s, or higher).
[0035] In addition, the processing unit 320 is coupled to the HBM device 330 through a first portion 342 of the SiP bus 340, and the first portion 342 includes one or more routing lines formed into (or on) the base substrate 310 ( Figure 3 FIG. schematically illustrates two). In various embodiments, the routing lines of the first portion 342 may include one or more metallization layers and / or interconnect metallization layers and / or vias of traces formed in one or more RDL layers of the base substrate 310. In addition, it will be understood that the processing unit 320 and the HBM device 330 may each be coupled to the routing lines of the first portion 342 via interconnects (e.g., solder balls, micro-bumps, pillars (e.g., copper pillars), and / or any other suitable components), metal-to-metal bonding, and / or any other suitable conductive bonding. Further, the signal routing lines of the first portion 342 and the TSVs 338 of the HBM device 330 allow the dies in the HBM device 330 and the processing unit 320 to communicate data at a relatively high rate (e.g., about 100 GB / s, 1000 GB / s, or higher).
[0036] As Figure 3 FIG. further illustrates, the HBS device 350 also includes a semiconductor die stack. In the illustrated embodiment, the semiconductor die stack in the HBS device 350 includes an interface die 352, one or more volatile memory dies 354 carried by the interface die 352 ( Figure 3 FIG. illustrates one), one or more non-volatile memory dies 356 carried by the interface die 352 ( Figure 3 FIG. illustrates five), and one or more TSVs 358 ( Figure 3 FIG. schematically illustrates six). It should be understood that in some embodiments, the HBS device 350 does not include the volatile memory dies 354 (e.g., includes only the interface die 352 and the non-volatile memory dies 356). Omitting the volatile memory dies 354 can help simplify the construction of the die stack by limiting the number of different types of semiconductor dies stacked together.
[0037] The TSV 358 (the part that is sometimes also referred to herein as the "HBM bus" (or forms the "HBM bus") extends through each of the volatile memory die 354 and the non-volatile memory die 356 from the interface die 332. Similar to the TSV 338 discussed above, the TSV 358 allows each die to communicate data within the HBS device 350 at a relatively high rate (e.g., about 100 GB / s, 1000 GB / s or higher) (e.g., between the volatile memory die 354 (e.g., DRAM die) and the non-volatile memory die 356 (e.g., NAND die, NOR die and / or any other suitable die), between the volatile memory die 354 and the interface die 352, and / or between the non-volatile memory die 358 and the interface die 352).
[0038] In addition, the HBM device 330 is coupled to the HBS device 350 via a second portion 344 of the SiP bus 340, and the second portion 344 includes one or more routing lines formed in (or on) the base substrate 310 ( Figure 3 two are schematically illustrated). As discussed above, the routing lines may include one or more metallization layers and / or interconnect metallization layers and / or vias of traces formed in one or more RDL layers of the base substrate 310. In addition, it should be understood that the HBM device 330 and the HBS device 350 may each be coupled to the routing lines of the second portion 344 via interconnects (e.g., solder balls, micro-bumps, pillars (e.g., copper pillars) and / or any other suitable components), metal-to-metal bonding and / or any other suitable conductive bonding. Furthermore, the signal routing lines of the second portion 344, the TSV 338 of the HBM device 330, and the TSV 358 of the HBS device 350 allow the dies in the HBM device 330 and the dies in the HBS device 350 to communicate data at a relatively high rate (e.g., about 100 GB / s, 1000 GB / s or higher).
[0039] In some embodiments, the interface die 352 includes a controller for the volatile memory die 354 and / or the non-volatile memory die 356, allowing the interface die 352 to control the HBS device 350 in response to various read and write requests. By locating the controller within the interface die 352, the HBS device 350 can reduce the number of signals that must be communicated over the SiP bus 340. However, in various other embodiments, the controller for the volatile memory die 354 and / or the non-volatile memory die 356 is located within the processing unit 320 and / or the HBM device 330. The non-volatile memory die 356 provides relatively large non-volatile storage within the SiP device 300 (e.g., on the order of hundreds of gigabytes, terabytes, and / or the like). Thus, relatively large data sets and / or the like can be stored entirely within the SiP device 300, reducing the need to retrieve data from an external storage device.
[0040] For example, as discussed in more detail below, during operation of the SiP device 300, the processing unit 320 can send a request for a subset of a large data set to the HBM device 330 over a first portion 342 of the SiP bus 340. The HBM device 330 can check whether the subset is stored in the volatile memory die 334 and, if not, forward the request and / or generate a new request for the data over a second portion 344 of the SiP bus 340 to the HBS device 350. The HBS device 350 can then write a copy of the data subset to the HBM device 330 over the second portion 344 of the SiP bus 340, allowing the HBM device 330 to send the data subset to the processing unit 320 over the first portion 342 of the SiP bus 340 for processing. Once the subset has been processed (and / or at various times during processing), the processing unit 320 can write the result of the processing to the HBM device 330 over the first portion 342 of the SiP bus 340. In turn, the HBM device 330 can write the result of the processing to the HBS device 350 over the second portion 344 of the SiP bus 340. The processing unit 320 can then send a request for another subset of the data set to the HBM device 330, and so on. In some embodiments, the process can be repeated any number of times as needed (e.g., when iteratively training a machine learning model with a data set). Thus, when the data set is available in the HBS device 350, the SiP device 300 is capable of completing any number of iterations of a processing operation without communicating (e.g., via a PCI bus) with an external storage component, thereby avoiding (or reducing through) the bottleneck discussed in more detail above and increasing the overall processing speed of the SiP device 300.
[0041] In some embodiments, the volatile memory die 354 serves as a buffer for the HBS device 350 to improve the response speed of the HBS device 350. For example, as discussed in more detail below, the HBS device 350 may receive a first request instructing the interface die 352 to load a subset of data from the non-volatile memory die 356 into the volatile memory die 354 for an upcoming request (e.g., when the processing unit 320 knows which data it will need next), and then receive a second request instructing the interface die 352 to send the data from the volatile memory die 354 to the HBM device 330 and / or the processing unit 320. By loading the subset of data into the volatile memory die 354 in response to the first request, the HBS device 350 can help reduce the response time to the second request, thereby further improving the overall processing speed of the SiP device 300.
[0042] In some embodiments, the processing unit 320 may communicate directly with the HBS device 350 to retrieve a subset of data. For example, in the embodiment illustrated in Figure 3 , the processing unit 320 is directly coupled to the HBS device 350 through the third portion 346 of the SiP bus 340, and the third portion 346 includes one or more routing lines 344 formed in (or on) the substrate 310 ( Figure 3 one is schematically illustrated in). The direct coupling between the processing unit 320 and the HBS device 350 may allow a new subset of data to be directly loaded into the processing unit 320 at the start of a new operation (e.g., avoiding the buffering time associated with loading the subset into the HBM device 330 and then loading the subset into the processing unit 320). Additionally or alternatively, the direct coupling between the processing unit 320 and the HBS device 350 may allow the processing unit 320 to periodically save the state of the processing unit 320 directly to the HBS device 350 to create a non-volatile backup of the current state (e.g., after a predetermined amount of time, after a processing milestone, and / or the like).
[0043] However, as illustrated in Figure 3 , the third portion 346 of the SiP bus 340 may have a longer distance than either the first portion 342 or the second portion 344. The longer distance may in turn result in a lower bandwidth in the third portion 346 than in the first portion 342 and the second portion 344, or higher manufacturing costs and / or operating power requirements to create and / or convey data on a longer routing line having the same bandwidth as the first portion 342 and the second portion 344. In some embodiments, the reduction in bandwidth associated with the longer routing line in the third portion 346 causes larger data subsets to be conveyed more quickly through the first portion 342 and the second portion 344, with the HBM device 330 serving as a buffer.
[0044] As Figure 3It is further illustrated in that the SiP device 300 further includes interconnects 362 that extend from the upper surface 312 of the base substrate 310 to the lower surface 314 of the base substrate 310. The interconnects 362 can provide external connections for the processing unit 320, the HBM device 330, and the HBS device 350. For example, the interconnects 362 can couple any one of the processing unit 320, the HBM device 330, and / or the HBS device 350 to an external component (e.g., a PCI bus coupled to an external storage device, an external controller, and / or the like). Additionally or alternatively, the interconnects 362 can couple any one of the processing unit 320, the HBM device 330, and / or the HBS device 350 to a power source. Additionally or alternatively, the interconnects 362 can couple any one of the processing unit 320, the HBM device 330, and / or the HBS device 350 to test pins on the lower surface 314 of the base substrate 310 (e.g., to allow evaluation of the processing unit 320, the HBM device 330, and / or the HBS device 350 after assembling the SiP device 300).
[0045] Figure 4 is a partially exploded schematic view of an HBS device 400 configured according to some embodiments of the present technology. For example, the HBS device 400 can be used as the HBS device 350 discussed above with reference to Figure 3 In the illustrated embodiment, the HBS device 400 is a die stack that includes an interface die 410, one or more volatile memory dies 420 ( Figure 4 one is illustrated in), and one or more non-volatile memory dies 430 ( Figure 4 four are illustrated in). Additionally, the HBS device 400 includes a shared HBM bus 440 that communicatively couples the interface die 410, the volatile memory dies 420, and the non-volatile memory dies 430.
[0046] The interface die 410 can be an external component between the shared HBM bus 440 and the shared HBM bus 440 (e.g., Figure 3A physical layer ("PHY") that establishes an electrical connection between the routing lines in the second part 344 of the SiP bus 340. Additionally or alternatively, the interface die 410 may include one or more active components, such as a static random access memory (SRAM) cache, a memory, and / or a storage controller, and / or any other suitable components. The volatile memory die 420 may be a DRAM memory die that provides low-latency memory access to the HBS device 400 (e.g., serves as a buffer die for the HBS device 400). However, it should be understood that in some embodiments, the HBS device 400 does not include any volatile memory dies. The non-volatile memory die 430 (sometimes referred to herein as a "storage die", "memory expansion", "memory expansion die", and the like) may provide a non-volatile storage device (e.g., a NAND flash memory device) for the HBS device 400. Additionally, the non-volatile memory die 430 may provide a significant expansion of the available memory for the SiP device integrating the HBS device 400 (e.g., two times, three times, four times, five times, ten times, one hundred times, or any other suitable increase).
[0047] In a specific non-limiting example, each non-volatile memory die 430 may provide 64 GB of memory storage. Thus, Figure 4 The four non-volatile memory dies 430 illustrated in provide a total memory capacity of 256 GB. In another specific non-limiting example, each non-volatile memory die 430 may provide 128 GB of memory. Thus, Figure 4 The four non-volatile memory dies 430 illustrated in provide a total memory capacity of 512 GB. In yet another specific non-limiting example, each non-volatile memory die 430 may provide 256 GB of memory. Thus, Figure 4 The four non-volatile memory dies 430 illustrated in provide a total memory capacity of 1024 GB. In each of these examples, the SiP device incorporating the HBS device 400 (e.g., Figure 3 the SiP device 300) may reduce (or avoid) the latency of loading memory from an external storage component (and via a low-bandwidth communication channel) by loading data into the non-volatile memory die 430 once and then accessing the data via the high-bandwidth communication channels in the SiP bus discussed above.
[0048] As Figure 4 further illustrated in, the shared HBM bus 440 may include a plurality of TSVs 442 that extend from the interface die 410 through the volatile memory die 420 and through each non-volatile memory die 430 ( Figure 4Four are illustrated, but any suitable number of TSVs is possible). Each of the TSVs 442 can support independent bidirectional read / write operations to communicate data between dies in the HBS device 400 (e.g., between the interface die 410 and the volatile memory die 420, between the non-volatile memory die 430 and the volatile memory die 420, between the interface die 410 and the non-volatile memory die 430, and / or the like). Since the TSVs 442 establish a shared HBM bus 440 between each of the dies in the HBS device 400, the shared HBM bus 440 can reduce (or minimize) the footprint required to establish a high-bandwidth communication route through the HBS device 400. Thus, the shared HBM bus 440 can reduce (or minimize) the overall footprint of the HBS device 400.
[0049] Figure 5 is a flowchart of a process 500 for operating a SiP device according to some embodiments of the present technology. The process 500 can be accomplished by a controller (e.g., a package controller) that communicates with and / or is onboard the SiP device (e.g., Figure 3 the processing unit 320, within Figure 3 the HBS device 350 of
[0050] The process 500 begins at block 502, where data is written to an HBS device (e.g., Figure 3 the HBS device 350 of Figure 2The third communication channel 256) is written from an external storage component into the HBS device. In some such embodiments, the HBS device is large enough to store the entire data set for complex computing operations (e.g., image and / or video rendering, AI / ML algorithms, and / or the like). In such embodiments, the data set must pass through the bottleneck to be loaded into the SiP device (e.g., via a PCI bus) only once. Thereafter, the entire data set is available at a single location for any suitable number of iterations of the computing operation via the high-bandwidth communication channel. In some embodiments, the SiP device includes multiple HBS devices, and the data written into each HBS device is a partition of a larger data set. For example, for a corresponding number of HBS devices in the SiP and / or a corresponding number of SiP devices having one or more HBS devices, the larger data set can be split into two, three, four, and / or any other suitable number of parts. Additionally or alternatively, the data set can be split according to external requirements (e.g., according to the desired batch size of the data in the AI / ML process, to maximize resource utilization during the computing process, and the like). In a specific non-limiting example, the SiP device can include four HBS devices similar to the HBM device discussed above with respect to Figure 3 and 4 Each HBS device has a stack of non-volatile memory dies providing 128 GB of memory. In this example, a data set with 512 GB of data can be divided into four 128 GB partitions, each partition can be loaded into the corresponding HBS device to be accessed during the AI / ML process. In some embodiments, for example when the HBS device stores (and / or only stores) the primary copy of the data used by the electronic device including the SiP device, data is written from another suitable external component (e.g., a bus component coupled to another electronic device, a data capture device, an input / output device, and / or the like).
[0051] In some embodiments, the write operation at block 502 includes determining the role of one or more non-volatile memory dies in the HBS device. For example, a first subset of non-volatile memory dies can be assigned as core dies, a second subset of non-volatile memory dies can be assigned as spare dies, and a third subset of non-volatile memory dies can be assigned as error correction code (ECC) dies. Additional details regarding the functionality of the subsets are discussed below with respect to Figure 7 discussed.
[0052] Because the write operation at block 502 requires data to move from an external storage component and / or another external device into the HBS device, the write operation can require the data to move through a relatively low-bandwidth bus (e.g., as referenced above with respect to Figure 1The described bottleneck is approximately 8 GB / s). Thus, the write operation may take several seconds to complete. However, as discussed in more detail below, the data is then available via a high-bandwidth communication path within the SiP device, allowing the data to be used any number of times without passing through the bottleneck again.
[0053] At block 504, process 500 includes receiving (or generating) a request for a subset of data in the HBS device. For example, the request may be received from a CPU / GPU and / or any other suitable controller in the SiP device. Additionally or alternatively, the request may be generated by a controller in the HBS device (e.g., by Figure 4 interface die 410) when an external component is expected to need the data and / or based on a previous request from an external component. In some embodiments, receiving the request causes the HBS device (e.g., via a controller in the interface die) to check whether the requested subset of data is stored in a volatile memory die in the HBS device. When the requested subset is found in the volatile memory die, process 500 may proceed to block 506 (e.g., when the subset was written to the volatile memory die in anticipation of the request), otherwise process 500 must retrieve the data from one or more non-volatile memory dies in the HBS device.
[0054] At block 506, process 500 includes writing (or causing to be written) a copy of the subset of data from the HBS device to one or more HBM devices in the SiP device (e.g., Figure 3 HBM device 330) and / or directly to a processing device in the SiP device. The write operation may use a portion of the SiP bus between the HBS device and the HBM device (e.g., second portion 344) and / or a portion of the SiP bus between the HBS device and the processing device (e.g., third portion 346) to write the requested subset via a high-bandwidth communication path. Thus, the write operation at block 506 can be performed in a time frame of approximately a few tens of microseconds, such that the subset is almost immediately available. Once stored in the HBM device, the subset of data can be used by the controller and / or processing unit for typical purposes via the high-bandwidth communication path (e.g., via Figure 3 first portion 342 of SiP bus 340).
[0055] At block 508, process 500 includes reading a subset of data in the HBM device. The read operation may move a copy of the subset (and / or a portion of the subset) to the processing unit (e.g., Figure 3 processing unit 320) via another portion of the SiP bus (e.g., Figure 3 first portion 342 of SiP bus 340). At block 510, process 500 includes (e.g., at Figure 3A subset of the data to be processed is read at the processing unit 320). And at block 512, process 500 may write the result of the processing at block 510 to the HBM device via a high-bandwidth communication path. Since the read / write operations at blocks 508 and 512 can use the high-bandwidth communication path to convey data, the subset of data can be available for processing within a few tens of microseconds, and / or the result of the processing is saved within a few tens of microseconds, such that the processing at block 510 is typically the limiting factor for the speed of process 500 from block 508 to 512. After writing the result of the processing to the HBM device at block 512, process 500 may return to block 508 to repeat the process from block 508 to 512 any suitable number of times (e.g., when the processing at block 510 is part of an AI / ML algorithm that iteratively processes subsets of data), and / or may return to block 504 to receive (or generate) a request for a second subset of data in the HBS device and write the second subset of data to the HBM device for processing.
[0056] Additionally or alternatively, at block 514, process 500 includes writing the result of the processing to the HBS device. In some embodiments, the write at block 514 writes the result of the processing directly from the processing unit to the HBS device (e.g., via Figure 3 the third part 346 of the SiP bus 340). In some such embodiments, the write at block 514 may occur simultaneously (or substantially simultaneously) with the write at block 512. Additionally or alternatively, the write at block 514 may be performed instead of the write at block 512. In some embodiments, the write at block 514 writes the result of the processing from the HBM device to the HBS device (e.g., via Figure 3 the second part 344 of the SiP bus 340). After writing the result of the processing to the HBS device at block 514, process 500 may return to block 508 to repeat the process from block 508 to 512 any suitable number of times (e.g., when the processing at block 510 is part of an AI / ML algorithm that iteratively processes subsets of data, when the write at block 514 saves intermediate results of the processing during a long processing operation, and / or the like), and / or may return to block 504 to receive (or generate) a request for a second subset of data in the HBS device and write the second subset of data to the HBM device for processing.
[0057] In various specific non-limiting examples, process 500 may be part of an AI / ML algorithm, a video rendering process, a high-resolution graphics rendering process, various complex computer simulations, and / or any other suitable computing application. In such embodiments, the CPU / GPU will typically call and / or reference each subset of data more than once. Thus, as referenced above Figures 2 to 4The discussed SiP architecture allows process 500 to avoid reading data multiple times from the storage component (and through a low bandwidth communication channel). Instead, the data is written to the HBS device once and then written to the HBM device and read any suitable number of times. Although the initial write operation is bottlenecked by the low bandwidth communication path from the storage component, each subsequent access to a subset of the data (and / or sequential access to each subset) uses the high bandwidth path. Thus, each subsequent use of the data may take tens of microseconds instead of one second or more, potentially increasing the speed of the processing operation by several orders of magnitude.
[0058] Figure 6 is a flowchart of a process 600 for operating a high bandwidth storage device according to some embodiments of the present technology. Process 600 may be implemented by a storage controller within an interface die of the HBS device (e.g., Figure 3 interface die 352 in HBS device 350 of Figure 3 and / or another suitable controller in the SiP device (e.g., by
[0059] the controller in processing unit 320 of Figure 3 ). Figure 3 Process 600 begins at block 602 by receiving (or generating) a first request for a subset of data in the HBS device. The first request may be received from, for example, a CPU / GPU in the processing unit of the SiP device and / or any other suitable controller when it is anticipated that an external component will need the data in the future (e.g., when the CPU / GPU will need the data in the future). By way of example only, the first request may be received 10 cycles, 100 cycles, 1000 cycles, and / or any other suitable number of cycles earlier than the anticipated need for the data. The first request allows the HBS device to check whether the requested subset of data is available in the DRAM die of the HBS device (e.g.,
[0060] volatile memory die 354 of Figure 5writing at the frame 506 so that a subset of data is available for processing at the processing unit.
[0061] Figure 7 is a partial cross-sectional view of an HBS device 700 configured according to a further embodiment of the present technology. As Figure 7 illustrated, the HBS device 700 is generally similar to the HBS device 350 discussed above with reference to Figure 3 For example, the HBM device 700 may include an interface die 710, and one or more volatile memory dies 720 ([[]] Figure 7 one is illustrated therein) and one or more non-volatile memory dies 730 ([[]] Figure 7 seven are illustrated therein) carried by the interface die 710. In addition, each die in the HBS device is communicatively coupled via TSVs 742 in the HBS bus 740.
[0062] However, in the illustrated embodiment, the non-volatile memory dies 730 are divided into three groups to support the operation of the HBS device 700 and / or the SiP device as a whole. More specifically, the non-volatile memory dies 730 may include one or more core dies 732 ([[]] Figure 7 three are illustrated therein), one or more spare dies 734 ([[]] Figure 7 three are illustrated therein) and / or one or more ECC dies 736. The core dies 732 can be used to implement any of the functions discussed above (e.g., storing a data set, writing a subset of data to the processing unit of the HBM device and / or the SiP device, and / or the like). The spare dies 734 can be redundant dies that store a secondary copy and / or backup of the data stored in the core dies 732, and / or can replace any one of the core dies 732 in the event of a failure (e.g., when one or more arrays in the core dies 732 fail, when the interconnect with one or more of the core dies 732 is broken, and / or the like). The ECC dies 736 can operate similarly to the ECC dies used in DRAM memory devices. For example, the ECC dies 736 can store one or more ECC codes generated based on the actual data stored in the core dies 732. During a read of the core dies 732, a controller (e.g., the memory controller in the interface die 710) reads both the data from the core dies 732 and the corresponding ECC codes from the ECC dies 736, regenerates an ECC code from the data, and compares the regenerated code with the read code. If there is a match, then no error has occurred in the actual data during the storage and / or read process. If there is a mismatch, then the ECC codes can allow the controller (or another suitable component) to correct various errors in the data before writing the data outside the HBS device.
[0063] As discussed above, the partitioning of the non-volatile memory die 730 can be controlled by a memory controller (e.g., the memory controller in the interface die 710) before and / or when the HBS device 700 receives data. Thus, in various other embodiments, other partitioning of the non-volatile memory die is possible. For example, since the ECC die 736 provides a layer of protection against errors in the data, the HBS device 700 can include six core dies 732, one ECC die 736, and no spare dies 734. In another example, the HBS device 700 can include an equal number of core dies 732 and spare dies 734 and no ECC die 736. In such embodiments, the HBS device 700 may lack an additional layer of protection against errors in the data, but may respond more quickly to read requests (e.g., because there is no need to complete the check of the ECC code).
[0064] Figure 8A and 8B are flowcharts of processes 800, 820 for powering down and powering up a SiP device using an HBS device, respectively, according to some embodiments of the present technology. Processes 800, 820 can be completed by a controller (e.g., a package controller) communicating with the SiP device and / or a controller (e.g., Figure 3 the processing unit 320 of Figure 4 and / or a controller on the interface die 410 of
[0065] Figure 8A Process 800 of Figure 5 and 6 begins at block 802 by processing at least a portion of the data in the HBM device in the SiP device. The processing at block 802 can be generally similar (or identical) to the processing discussed above with reference to
[0066] For example, the processing at block 802 can include reading a portion of the data in the HBM device into the processing unit and processing the read portion of the data.
[0066] At block 804, process 800 includes updating the current state of the HBM device with the result of the processing at block 802 and / or the current state of the processing unit. The current state in the HBM device allows the result of the processing and / or the previous state of the processing unit to be quickly recalled as needed during processing (e.g., when an error occurs after saving, for further processing, and / or the like). In a specific non-limiting example, the update at block 804 can save the state of the processing unit during a relatively long processing operation such that, in response to an error in the processing, the processing unit can return to the saved checkpoint instead of restarting completely.
[0067] Process 800 may complete boxes 802 and 804 (collectively box 806) any number of times during the operation of the SiP to support typical processing in a semiconductor device. During the processing and updates at box 806, read / write operations may use a portion of the high-bandwidth SiP bus between the processing unit and the HBM device (e.g., Figure 3 the first portion 342) to rapidly communicate data back and forth between the HBM device and the processing component, allowing the read / write operations to not impose significant time constraints on the processing. Additionally, in some embodiments, process 800 includes periodically writing to the HBS device in the SiP at box 806 to save the results of various processing operations, save the current state of the SiP device, the current state of the processing device, the current state of the HBM device, and / or any relevant information. Because the HBS device is also coupled to the SiP bus (e.g., coupled to Figure 3 the second portion 344 and / or the third portion 346), the periodic saving can prevent power outages, other power losses, and / or various other error sources without requiring a significant time investment and / or suspension in the processing operations. In a specific non-limiting example, a processing operation may be expected to take multiple hours to complete. In this example, process 800 may save the results of the processing operation and / or the current state of the SiP device every minute, every ten minutes, and / or at any other suitable interval to reduce the likelihood that the processing operation will have to start completely over (e.g., when there is a power outage, in response to a severe error, and / or the like).
[0068] At box 808, process 800 includes receiving a power-down request (sometimes also referred to herein as an idle request). The power-down request may be received in response to an input from a user and / or another component of the system using the SiP device (e.g., to save power when the battery power of an electronic device is low and / or in response to a power outage).
[0069] At box 810, process 800 includes writing the state of the processing device, the HBM device, and / or any other suitable components of the SiP device to the HBS device. Because the HBS device is coupled to the SiP bus, the write operation can be completed in a few tens of microseconds (e.g., as compared to one second or more required to write data to a conventional storage device (e.g., Figure 1 the storage device 140)). Thus, the SiP device can comply with the power-down request in a few tens of microseconds, allowing the semiconductor device to save power, reduce data loss during a power outage, and / or otherwise quickly shut down when requested.
[0070] Correspondingly, Figure 8BThe process 820 can start at block 822 by receiving a power-on request (sometimes also referred to herein as a wake-up request). The power-on request can be received in response to an input from a user and / or another component of a system using the SiP device (e.g., another controller in a semiconductor device). And at block 824, the process 820 can read / write the previous state of the SiP device from the HBS device to the HBM device, a processing device, and / or any other suitable component of the SiP device. Similar to the discussion above, since the HBS device is coupled to the SiP bus, the SiP device (and the corresponding semiconductor device) can respond to the power-on request within tens of microseconds (e.g., rather than one second or more required to read / write from a conventional storage component). Thus, the SiP device (and the corresponding semiconductor device) can be prepared for computing activities significantly faster than conventional devices.
[0071] From the foregoing, it will be appreciated that specific embodiments of the present technology have been described herein for purposes of illustration, but well-known structures and functions have not been shown or described in detail to avoid unnecessarily obscuring the description of embodiments of the present technology. To the extent that any material incorporated by reference herein conflicts with the present disclosure, the present disclosure controls. Where context permits, singular or plural terms may also respectively include plural or singular terms. Additionally, unless the word "or" is explicitly limited to meaning a single item excluded from a list of two or more items only, the use of "or" in a list herein will be interpreted to include (a) any single item in the list, (b) all items in the list, or (c) any combination of items in the list. Further, as used herein, the phrase "and / or" in "A and / or B" means A alone, B alone, and both A and B. Additionally, the terms "comprising," "including," "having," and "owning" are always used to mean at least the stated feature, such that any greater number of the same features and / or additional types of other features are not excluded. Further, the terms "substantially," "about," and "approximately" are used herein to mean within at least 10% of a given value or limit. By way of pure example, an approximate ratio means within ten percent of a given ratio.
[0072] Several embodiments of the disclosed technology have been described above with reference to the figures. A computing device on which the described technology can be implemented can include one or more central processing units, a memory, an input device (e.g., a keyboard and a pointing device), an output device (e.g., a display device), a storage device (e.g., a disk drive), and a network device (e.g., a network interface). The memory and the storage device are computer-readable storage media that can store at least portions of instructions for implementing the described technology. Additionally, data structures and message structures can be stored or transmitted via a data transmission medium (e.g., a signal on a communication link). Thus, computer-readable media can include computer-readable storage media (e.g., "non-transitory" media) and computer-readable transmission media.
[0073] It will also be understood that various modifications can be made without departing from the present disclosure or the present technology. For example, the dies in the HBM device can be arranged in any other suitable order (e.g., the non-volatile memory die is positioned between the interface die and the volatile memory die; the volatile memory die is at the bottom of the die stack; and the like). In addition, those skilled in the art will understand that the various components of the present technology can be further divided into sub-components, or the various components and functions of the present technology can be combined and integrated. Additionally, the specific aspects of the technology described in the context of a particular embodiment can also be combined or eliminated in other embodiments. For example, although the use of non-volatile memory dies (e.g., NAND dies and / or NOR dies) to expand the memory of an HBM device is discussed herein, it should be understood that alternative memory expansion dies (e.g., larger capacity DRAM dies and / or any other suitable memory component) can be used. Although such embodiments may forego certain benefits (e.g., non-volatile storage), such embodiments may still provide additional benefits (e.g., reducing traffic through the bottleneck, allowing many complex computational operations to be performed relatively quickly, etc.).
[0074] In addition, although the advantages associated with the specific embodiments of the present inventive technology have been described in the context of those embodiments, other embodiments may also exhibit such advantages, and not all embodiments must exhibit such advantages to fall within the scope of the present inventive technology. Accordingly, the present disclosure and the associated technology may cover other embodiments not explicitly shown or described herein.
Claims
1. A system-in-package (SiP) device, comprising: A base substrate; A processing unit carried by the base substrate; A high-bandwidth memory (HBM) device carried by the base substrate and electrically coupled to the processing unit via a SiP bus, wherein the HBM device comprises: A first interface die; One or more volatile memory dies carried by the first interface die; and An HBM bus electrically coupled to the first interface die and each of the one or more volatile memory dies; and A high-bandwidth storage (HBS) device carried by the base substrate and electrically coupled to the HBM device via the SiP bus, wherein the HBS device comprises: A second interface die; One or more non-volatile memory dies carried by the second interface die; and An HBS bus electrically coupled to the second interface die and each of the one or more non-volatile memory dies.
2. The SiP device according to claim 1, wherein the second interface die includes a storage controller configured to: Receive a request for a subset of a data set stored in the one or more non-volatile memory dies; Read the subset of the data set from the non-volatile memory dies; and And Send a copy of the subset of the data set to the one or more volatile memory dies in the HBM device.
3. The SiP device according to claim 1, wherein the HBS device further includes a volatile memory die carried by the second interface die and electrically coupled to the HBS bus.
4. The SiP device according to claim 3, wherein the second interface die includes a storage controller configured to: Receive a first request for a subset of a data set stored in the one or more non-volatile memory dies; Write a copy of the subset of the data set from the one or more non-volatile memory dies to the volatile memory die carried in the HBS device; Receive a second request for the subset of the data set; and And Send a copy of the subset of the data set from the volatile memory die carried in the HBS device to the one or more volatile memory dies in the HBM device.
5. The SiP device according to claim 1, wherein the SiP bus includes: A first portion electrically coupled between the processing unit and the HBM device, wherein the first portion has a first bandwidth; and And A second portion electrically coupled between the HBM device and the HBS device, wherein the second portion has a second bandwidth substantially equal to the first bandwidth.
6. The SiP device according to claim 1, wherein the HBS device is directly electrically coupled to the processing unit via the SiP bus.
7. The SiP device according to claim 1, wherein the one or more non-volatile memory dies in the HBS device are configured to provide a non-volatile copy of data stored in the one or more volatile memory dies in the HBM device that is accessible by the one or more volatile memory dies via the SiP bus in response to a power-on request.
8. A method, comprising: Generating a request for a subset of a data set stored in a plurality of non-volatile memory dies in a high-bandwidth storage device (HBS); Writing a copy of the subset to a plurality of volatile memory dies in a high-bandwidth memory (HBM) device; Reading the subset from the plurality of volatile memory dies into a processing unit; Processing the subset at the processing unit; And Writing the result of processing the subset to the plurality of volatile memory dies.
9. The method according to claim 8, wherein the request for the subset of the data set is a second request, the copy of the subset of the data set is a second copy, and the method further comprises: Generating a first request for the subset of the data set; And Writing a first copy of the subset to one or more volatile memory dies in the HBS device before generating the second request, wherein writing the second copy of the subset to the plurality of volatile memory dies in the HBM device in response to the second request comprises writing the second copy from one or more volatile memory dies in the HBS device to the plurality of volatile memory dies in the HBM device.
10. The method according to claim 8, further comprising writing the result of processing the subset to the plurality of non-volatile memory dies.
11. The method according to claim 8, wherein the processing unit is communicatively coupled to the plurality of volatile memory dies in the HBM device through a portion of a system-in-package (SiP) bus, and wherein reading the subset from the plurality of volatile memory dies into the processing unit comprises conveying the subset through the portion of the SiP bus.
12. The method according to claim 8, wherein the plurality of non-volatile memory dies in the HBS device are coupled to the plurality of volatile memory dies in the HBM device through a portion of a system-in-package (SiP) bus, and wherein writing the copy of the subset to the plurality of volatile memory dies in the HBM device comprises conveying the subset through the portion of the SiP bus.
13. The method according to claim 8, further comprising: Receiving a power-off or idle request; And In response to the power-off or idle request, writing the current state of the plurality of volatile memory dies in the HBM device to the plurality of non-volatile memory dies in the HBS device through a system-in-package (SiP) bus.
14. The method according to claim 8, further comprising, after processing the subset for a predetermined period of time, writing the current states of the processing unit and the plurality of volatile memory dies in the HBM device to the plurality of non-volatile memory dies in the HBS device via a system-in-package SiP bus.
15. The method according to claim 8, wherein writing the data set to the plurality of non-volatile memory dies in the HBS device comprises: assigning a first subset of the plurality of non-volatile memory dies in the HBS device to store a primary copy of the data set; and assigning a second subset of the plurality of non-volatile memory dies in the HBS device to store a secondary copy of the data set.
16. The method according to claim 8, wherein writing the data set to the plurality of non-volatile memory dies in the HBS device comprises: assigning a first subset of the plurality of non-volatile memory dies in the HBS device to store a primary copy of the data set; and assigning a second subset of the plurality of non-volatile memory dies in the HBS device to serve as error correction code dies.
17. A method, comprising: updating a current state of the HBM device during a processing operation at a processing unit communicatively coupled to a high-bandwidth memory HBM device; receiving a power-down or idle request; and in response to the power-down or idle request, controlling the HBM device to write the current state of the HBM device from the HBM device to a high-bandwidth storage HBS device via a system-in-package SiP bus, wherein the HBS device includes one or more non-volatile memory dies.
18. The method according to claim 17, further comprising, in response to the power-down or idle request, writing the current state of the processing unit to the HBS device via the HBM device and the SiP bus.
19. The method according to claim 17, further comprising: receiving a power-up or wake-up request; and in response to the power-up or wake-up request, controlling the HBS device to write the current state of the HBM device from the HBS device back to the HBM device via the SiP bus.
20. The method according to claim 17, wherein writing the current state of the HBM device from the HBM device to the HBS device takes less than 100 milliseconds.